Which LLM is best for my Mac, and which will actually run?
The best LLM for your Mac is the largest good model that fits the memory its GPU may use, and that number is not total memory. Apple publishes it: Metal reports it as the recommended maximum working set size, and on Apple silicon it comes to roughly three quarters of unified memory, so a Mac with 48 GB has about 37 GB available to a model, and a model whose weights are larger than the budget will either refuse to load or fall back to swapping and become unusable. The practical rule is to keep the file on disk under about two thirds of your available budget, because context and key-value cache grow on top of the weights as the conversation gets longer. Grux OS reads the actual hardware, computes that budget, and rates each model in its catalogue as a good fit, tight, or not worth downloading.
Grux OS 3.0.0 · last checked 2026-09-30 · generated from the shipping release
Working the budget out by hand
Take your unified memory. system_profiler SPHardwareDataType | grep Memory.
Multiply by about 0.75. That is roughly what the GPU may take, and macOS keeps the rest. The exact number for your machine is what Metal returns for recommendedMaxWorkingSetSize, so it is a value you can read rather than a rule of thumb you have to trust.
Compare against the model file on disk, not the parameter count. A 30B model at 4-bit is about 19 GB; the same model at 8-bit is about double that.
Leave headroom for context. A long conversation can add several GB on top of the weights.
Unified memory
Roughly available to a model
Comfortable model size on disk
16 GB
about 12 GB
up to about 8 GB
24 GB
about 18 GB
up to about 12 GB
36 GB
about 27 GB
up to about 18 GB
48 GB
about 37 GB
up to about 24 GB
64 GB
about 48 GB
up to about 32 GB
128 GB
about 96 GB
up to about 64 GB
What Grux OS does with that
The Local Models surface reads your hardware profile rather than asking you for it, budgets the memory, and scores each model in a curated catalogue against the result. It returns one recommendation instead of a list, and it will tell you what not to bother with, which is the more useful half.
On an M4 Pro with 48 GB it reports about 37 GB available, marks Qwen 3 Coder 30B a good fit at 19.0 GB on disk, and marks a 35B at 24.0 GB as tight. That is the whole feature: the answer before the download rather than after it. Those disk sizes are the ones Ollama publishes per model and quantisation, so you can check any of them without installing Grux OS.
Why the parameter count is the wrong number
A model is sold by parameters and downloaded by bytes, and the ratio between them is the quantisation. The same 30B model can be a 19 GB file, or three times that at higher precision. Sizing a download by parameter count is how a Mac that could have run the model ends up swapping on one that never fitted.
Where these numbers come from
Nothing on this page is a Grux OS measurement. The memory arithmetic is Apple's and the model sizes are Ollama's, so every figure above can be checked against its own source without taking this page's word for it.
The GPU memory budget. Metal's recommendedMaxWorkingSetSize is the API that reports how much memory the GPU may take, and it is where the three quarters figure comes from.
Unified memory itself. Apple's Apple silicon documentation covers the shared memory architecture that makes the GPU share a budget with everything else on the machine.
Model file sizes. The Ollama model library publishes the size on disk of every model at every quantisation, which is the number to size a download by.
What Grux OS adds. The catalogue and the fit scoring live in the public source tree, so the rating logic is readable rather than asserted.
Questions
What is the best LLM for a Mac mini or an M1 Mac?
It depends on memory, not on the model of Mac. With 8 GB, about 6 GB is available to the GPU, so stay with small models of about 3B parameters. With 16 GB, about 12 GB is available, enough for a well quantised 7B or 8B model.
How much memory does a local LLM need on a Mac?
Roughly the size of the model file on disk, plus several GB for context, and all of it has to fit inside about three quarters of your unified memory.
Can a 16 GB Mac run a local LLM?
Yes, comfortably up to about an 8 GB model file, which in practice means a well quantised 7B to 14B model.
Does Grux OS download models for me?
No. Ollama does that. Local Models tells you which one is worth pulling for the hardware you actually have.