// silicon, not the cloud

AI† where you want it.
If you want it.

Every model in SQLly runs on your machine, on your terms - or in the cloud if you decide so. Nothing fires until you flip the switch. Below: exactly what runs where, and how fast your specific chip should feel. Honest status note: this stack is rolling out now - each card below is badged with where it actually stands.

Everything ships in the off position

By default, all AI† features are off. No background downloads, no silent inference, no tokens leaving your laptop. You opt in feature by feature - and you pick the engine each one runs on. Go on, flip it.

// the controls

Four switches, zero surprises

Each capability is independent. Turn on what earns its keep, leave the rest dark.

⌨️
off by default FIM · autocomplete ◐ in progress

Inline completion, on-device

Fill-in-the-middle autocomplete driven entirely on device by Qwen. It sees the code before and after your cursor, so completions land in context instead of guessing forward blindly.

SELECT o.id, o.total▌
      ⟶ , o.created_at  FROM orders o WHERE …

Tuned for low latency on a dedicated code model - not a chat model wearing a completion hat.

Qwen2.5-Coder · local
💬
off by default agent chat ◐ in progress

Agent chat - your call

Run the agent fully on device for total privacy, or point it at the cloud when you want maximum horsepower. Same interface, your choice per request.

On-device performance estimates for every Apple Silicon chip are in the table below - so you know what “local” feels like before you commit.

on-device or cloud · per your call
🧮
off by default advanced · on-device ○ planned

Local vector DB

A local vector database backs fast AI† operations - semantic schema search, similarity lookups, and retrieval - without a single embedding round-trip to anyone’s server.

embeddings stay on disk, on your machine
🗃️
off by default advanced · on-device ◐ in progress

Entity cards & smart cache

SQLly builds local entity cards for all DDL so the model has dense, pre-chewed context the instant you ask. Pair that with advanced cache invalidation that refreshes only what actually changed.

  • Entity cards for every object - tables, views, sprocs, functions.
  • Surgical invalidation - stale entries die, hot paths stay warm.
// the numbers, on Mac

How fast will your Mac feel?

Practical generation-speed estimates for local models running on Apple Silicon. Numbers are generation tokens/sec - not prompt ingestion.

Why Mac-only numbers? SQLly runs on macOS, Windows, and Linux from one codebase - but the pig daily-drives an M4 Max MacBook Pro with 128GB of RAM, so Apple Silicon is what he can measure honestly. We test on Windows and Linux too; share your real-world token-generation numbers and we’ll publish them right here.

Real performance bends to inference runtime, quantization, context size, settings, unified-memory pressure, thermals, and model packaging. Treat these as honest ballparks, not guarantees.

Cloud reference (the “fast cloud feel”): Artificial Analysis pegs GPT-5.5 non-reasoning API output around 58 tokens/sec, with a low/non-reasoning band of roughly 58–100 t/s. Perceived speed in a chat UI varies with routing, streaming, reasoning mode, and load - so this is a feel reference, not a benchmark target.

Model Practical RAM Notes
Qwen3-4B-Instruct-2507 Q4_K_M 8GB min / 16GB rec Small default local SQL assistant model. Ollama package ~2.5GB.
gpt-oss-20b 16GB min / 24GB rec Stronger reasoning model. OpenAI lists 21B total params, 3.6B active params/token, 128K context.
Gemma 4 E2B 8GB min / 16GB rec Small Google-family model. Google lists Q4 inference memory ~2.9GB; Ollama package can be larger.

model RAM assumptions

Chip Qwen3-4B Q4_K_M8GB min / 16GB rec gpt-oss-20b16GB min / 24GB rec Gemma 4 E2B8GB min / 16GB rec How to identify chip / bin
M118–285–1016–28Shows as Apple M1. Base M1 is the low-bandwidth tier.
M1 Pro38–5214–2432–48Shows as Apple M1 Pro.
M1 Max50–7024–3845–65Shows as Apple M1 Max.
M1 Ultra60–8532–5055–80Shows as Apple M1 Ultra. Usually Mac Studio.
M225–358–1422–34Shows as Apple M2.
M2 Pro40–5516–2835–50Shows as Apple M2 Pro.
M2 Max55–7528–4450–70Shows as Apple M2 Max.
M2 Ultra70–10038–6060–90Shows as Apple M2 Ultra. Usually Mac Studio / Mac Pro.
M328–409–1624–38Shows as Apple M3.
M3 Pro35–5014–2232–48Shows as Apple M3 Pro.
M3 Max, lower bin55–7524–3850–70Apple M3 Max - 14-core CPU / 30-core GPU.
M3 Max, higher bin65–9032–4860–85Apple M3 Max - 16-core CPU / 40-core GPU.
M3 Ultra80–11545–7070–105Shows as Apple M3 Ultra. Usually Mac Studio.
M435–5010–1830–45Shows as Apple M4.
M4 Pro55–7022–3545–65Shows as Apple M4 Pro.
M4 Max, lower bin70–9532–4860–85Apple M4 Max - 14-core CPU / 32-core GPU.
M4 Max, higher bin85–11540–6075–100Apple M4 Max - 16-core CPU / 40-core GPU.
M545–6014–2435–55Shows as Apple M5.
M5 Pro65–8526–4255–75Shows as Apple M5 Pro.
M5 Max, lower bin80–10536–5570–95Apple M5 Max - 18-core CPU / 32-core GPU.
M5 Max, higher bin95–12545–7085–115Apple M5 Max - 18-core CPU / 40-core GPU.

estimated generation tokens/sec by Apple Silicon chip

// pick your fighter

Which model, and when

Sensible defaults, an escape hatch for hard problems, and a separate lane for autocomplete. This is the planned default lineup - it lands together with the local runtime integration above.

⚡
default · fast

Qwen3-4B - the default

Small, fast, and good enough for schema search and SQL drafting. Reasonable on M1 through M5. This is the planned out-of-the-box default.

Qwen3-4B-Instruct-2507 Q4_K_M
🧠
advanced · local reasoning

gpt-oss-20b - the heavy hitter

The optional “better local reasoning” model: harder SQL generation, complex schema reasoning, stored-proc explanation, multi-step debugging, code review, query rewrites. Slower, hungrier for RAM.

gpt-oss-20b
🔷
balanced · google-family

Gemma 4 E2B - the alternative

For when you want Google-family support, multimodal ecosystem alignment, and a smaller reasoning option in the middle of the pack.

Gemma 4 E2B
⌨️
keep autocomplete separate

FIM stays in its own lane

Do not draft inline completions with a chat/reasoning model. Autocomplete runs a dedicated FIM/code model - far lower latency.

Qwen2.5-Coder-1.5B Base
Qwen2.5-Coder-3B Base
🎚️
performance mode ○ planned

Or just pick a vibe

Don’t want to think about model names? Pick a mode and SQLly sorts the rest:

Fast      → Qwen3-4B
Balanced  → Gemma 4 E2B
Advanced  → gpt-oss-20b
Cloud     → GPT-5.5 or other hosted model

The installer shows download size, required RAM, and a live estimate for your chip before anything lands on disk.

// show your work

Sources & references

These estimates combine published model specs and package sizes with practical Apple Silicon memory-bandwidth scaling assumptions.

Useful references

// your laptop, your rules

Bring the intelligence home to the pig

Off by default. On by choice. On-device by design. The only AI† that asks permission first.

⬇ Download SQLly - it’s free
† “AI” is a registered trademark of people who can’t be bothered to say machine learning.