Outmake Download
Help

Models on this Mac

Getting a model, what runs locally, memory limits, and why the first answer can be slow.

A model on This Mac needs no account and nothing leaves your computer. That privacy has a cost: it has to fit in memory, and it has to load before it can answer.

Getting a model

Open Models (View › Models) and the Discover tab. Each card shows how well a model fits this Mac before you get it. Press Download to fetch it; a download can be paused and picked up again, and it carries on the next time Outmake opens if you quit partway through.

What runs on this Mac

Two things run locally, with nothing sent anywhere:

  • A model you downloaded in Models, which runs in Outmake's own engine.
  • Apple's built-in model, under This Mac › Apple Intelligence, which runs through Apple's Foundation Models framework.

A model served by Ollama or LM Studio on this Mac, or by a server at another address, is different again: OpenCode is what talks to those, not Claude or ChatGPT.

Memory limits

Every model card says how it fits: Runs great, Runs well, Tight fit, or Too big, worked out from this Mac's memory and how much everything else on it is using. Settings › Models › Run by Outmake has three more controls:

  • Keep Loaded: how long a model stays in memory after its last reply.
  • Models Loaded at Once: on a Mac with room for more than one, how many can be ready at the same time.
  • Memory for Models: how much of this Mac's memory models may use.

Why the first answer can take a minute

A model on this Mac isn't kept running all the time. The first message to it loads the model into memory, which for a large model can take a while, and the composer says Loading … on this Mac while it does. Once loaded, it has to read your whole message before it can start answering, shown as Reading the chat… with a percentage. After that first message the model stays loaded (for as long as Keep Loaded says), so the next reply starts much faster.

Still stuck? Write to hello@outmake.app.