AI† where you want it.
If you want it.
Every model in SQLly runs on your machine, on your terms - or in the cloud if you decide so. Nothing fires until you flip the switch. Below: exactly what runs where, and how fast your specific chip should feel. Honest status note: this stack is rolling out now - each card below is badged with where it actually stands.
Everything ships in the off position
By default, all AI† features are off. No background downloads, no silent inference, no tokens leaving your laptop. You opt in feature by feature - and you pick the engine each one runs on. Go on, flip it.
Four switches, zero surprises
Each capability is independent. Turn on what earns its keep, leave the rest dark.
Inline completion, on-device
Fill-in-the-middle autocomplete driven entirely on device by Qwen. It sees the code before and after your cursor, so completions land in context instead of guessing forward blindly.
SELECT o.id, o.total▌ ⟶ , o.created_at FROM orders o WHERE …
Tuned for low latency on a dedicated code model - not a chat model wearing a completion hat.
Qwen2.5-Coder · localAgent chat - your call
Run the agent fully on device for total privacy, or point it at the cloud when you want maximum horsepower. Same interface, your choice per request.
On-device performance estimates for every Apple Silicon chip are in the table below - so you know what “local” feels like before you commit.
on-device or cloud · per your callLocal vector DB
A local vector database backs fast AI† operations - semantic schema search, similarity lookups, and retrieval - without a single embedding round-trip to anyone’s server.
embeddings stay on disk, on your machineEntity cards & smart cache
SQLly builds local entity cards for all DDL so the model has dense, pre-chewed context the instant you ask. Pair that with advanced cache invalidation that refreshes only what actually changed.
- Entity cards for every object - tables, views, sprocs, functions.
- Surgical invalidation - stale entries die, hot paths stay warm.
How fast will your Mac feel?
Practical generation-speed estimates for local models running on Apple Silicon. Numbers are generation tokens/sec - not prompt ingestion.
Why Mac-only numbers? SQLly runs on macOS, Windows, and Linux from one codebase - but the pig daily-drives an M4 Max MacBook Pro with 128GB of RAM, so Apple Silicon is what he can measure honestly. We test on Windows and Linux too; share your real-world token-generation numbers and we’ll publish them right here.
Real performance bends to inference runtime, quantization, context size, settings, unified-memory pressure, thermals, and model packaging. Treat these as honest ballparks, not guarantees.
Cloud reference (the “fast cloud feel”): Artificial Analysis pegs GPT-5.5 non-reasoning API output around 58 tokens/sec, with a low/non-reasoning band of roughly 58–100 t/s. Perceived speed in a chat UI varies with routing, streaming, reasoning mode, and load - so this is a feel reference, not a benchmark target.
| Model | Practical RAM | Notes |
|---|---|---|
| Qwen3-4B-Instruct-2507 Q4_K_M | 8GB min / 16GB rec | Small default local SQL assistant model. Ollama package ~2.5GB. |
| gpt-oss-20b | 16GB min / 24GB rec | Stronger reasoning model. OpenAI lists 21B total params, 3.6B active params/token, 128K context. |
| Gemma 4 E2B | 8GB min / 16GB rec | Small Google-family model. Google lists Q4 inference memory ~2.9GB; Ollama package can be larger. |
model RAM assumptions
| Chip | Qwen3-4B Q4_K_M8GB min / 16GB rec | gpt-oss-20b16GB min / 24GB rec | Gemma 4 E2B8GB min / 16GB rec | How to identify chip / bin |
|---|---|---|---|---|
| M1 | 18–28 | 5–10 | 16–28 | Shows as Apple M1. Base M1 is the low-bandwidth tier. |
| M1 Pro | 38–52 | 14–24 | 32–48 | Shows as Apple M1 Pro. |
| M1 Max | 50–70 | 24–38 | 45–65 | Shows as Apple M1 Max. |
| M1 Ultra | 60–85 | 32–50 | 55–80 | Shows as Apple M1 Ultra. Usually Mac Studio. |
| M2 | 25–35 | 8–14 | 22–34 | Shows as Apple M2. |
| M2 Pro | 40–55 | 16–28 | 35–50 | Shows as Apple M2 Pro. |
| M2 Max | 55–75 | 28–44 | 50–70 | Shows as Apple M2 Max. |
| M2 Ultra | 70–100 | 38–60 | 60–90 | Shows as Apple M2 Ultra. Usually Mac Studio / Mac Pro. |
| M3 | 28–40 | 9–16 | 24–38 | Shows as Apple M3. |
| M3 Pro | 35–50 | 14–22 | 32–48 | Shows as Apple M3 Pro. |
| M3 Max, lower bin | 55–75 | 24–38 | 50–70 | Apple M3 Max - 14-core CPU / 30-core GPU. |
| M3 Max, higher bin | 65–90 | 32–48 | 60–85 | Apple M3 Max - 16-core CPU / 40-core GPU. |
| M3 Ultra | 80–115 | 45–70 | 70–105 | Shows as Apple M3 Ultra. Usually Mac Studio. |
| M4 | 35–50 | 10–18 | 30–45 | Shows as Apple M4. |
| M4 Pro | 55–70 | 22–35 | 45–65 | Shows as Apple M4 Pro. |
| M4 Max, lower bin | 70–95 | 32–48 | 60–85 | Apple M4 Max - 14-core CPU / 32-core GPU. |
| M4 Max, higher bin | 85–115 | 40–60 | 75–100 | Apple M4 Max - 16-core CPU / 40-core GPU. |
| M5 | 45–60 | 14–24 | 35–55 | Shows as Apple M5. |
| M5 Pro | 65–85 | 26–42 | 55–75 | Shows as Apple M5 Pro. |
| M5 Max, lower bin | 80–105 | 36–55 | 70–95 | Apple M5 Max - 18-core CPU / 32-core GPU. |
| M5 Max, higher bin | 95–125 | 45–70 | 85–115 | Apple M5 Max - 18-core CPU / 40-core GPU. |
estimated generation tokens/sec by Apple Silicon chip
Which model, and when
Sensible defaults, an escape hatch for hard problems, and a separate lane for autocomplete. This is the planned default lineup - it lands together with the local runtime integration above.
Qwen3-4B - the default
Small, fast, and good enough for schema search and SQL drafting. Reasonable on M1 through M5. This is the planned out-of-the-box default.
Qwen3-4B-Instruct-2507 Q4_K_M
gpt-oss-20b - the heavy hitter
The optional “better local reasoning” model: harder SQL generation, complex schema reasoning, stored-proc explanation, multi-step debugging, code review, query rewrites. Slower, hungrier for RAM.
gpt-oss-20b
Gemma 4 E2B - the alternative
For when you want Google-family support, multimodal ecosystem alignment, and a smaller reasoning option in the middle of the pack.
Gemma 4 E2B
FIM stays in its own lane
Do not draft inline completions with a chat/reasoning model. Autocomplete runs a dedicated FIM/code model - far lower latency.
Qwen2.5-Coder-1.5B Base Qwen2.5-Coder-3B Base
Or just pick a vibe
Don’t want to think about model names? Pick a mode and SQLly sorts the rest:
Fast → Qwen3-4B Balanced → Gemma 4 E2B Advanced → gpt-oss-20b Cloud → GPT-5.5 or other hosted model
The installer shows download size, required RAM, and a live estimate for your chip before anything lands on disk.
Sources & references
These estimates combine published model specs and package sizes with practical Apple Silicon memory-bandwidth scaling assumptions.
Useful references
- Qwen3-4B-Instruct-2507 Q4_K_M on Ollama - ollama.com/library/qwen3:4b-instruct-2507-q4_K_M
- OpenAI gpt-oss model specs - openai.com/index/introducing-gpt-oss
- Gemma 4 model memory requirements - ai.google.dev/gemma/docs/core
- Gemma 4 Ollama package listings - ollama.com/library/gemma4
- GPT-5.5 OpenAI announcement - openai.com/index/introducing-gpt-5-5
- GPT-5.5 non-reasoning output speed reference - artificialanalysis.ai/models/gpt-5-5-non-reasoning
Bring the intelligence home to the pig
Off by default. On by choice. On-device by design. The only AI† that asks permission first.
⬇ Download SQLly - it’s free