What’s actually running under the hood
No marketing gloss past this point. This is the real architecture - the crates, the data structures, the design decisions - for the SQLly GPUI client, which is the one under active development and the one every screenshot on this site comes from.
A GPU-rendered native UI, not a webview
The window you're looking at isn't Electron and isn't a browser engine wearing a native-looking skin.
Every pixel here - text layout, the completion popup, the syntax highlighting - is drawn by GPUI's own GPU-accelerated renderer.
The same rendering approach as Zed
SQLly's client is built on GPUI, the Rust UI framework written for the Zed editor. It's a retained/immediate hybrid GPU-accelerated toolkit - not a DOM, not AppKit views underneath, just Rust drawing directly to the GPU.
The result grid is its own crate
Rather than hand-rolling grid virtualization inline, the result grid consumes a dedicated published crate - sqlly-datatable - handling virtualization, selection, sorting, and filtering as a standalone component.
macOS, Windows, and Linux from one tree
The GPUI client targets all three desktop platforms from the same Rust codebase, with per-OS packaging: an .app bundle on macOS (Apple Silicon only), a .desktop entry on Linux, and a native .exe on Windows.
Try sqlly-datatable - the actual result grid, as WebAssembly
The result grid is its own open-source crate, and its sample app compiles to WebAssembly on GPUI’s web backend - drawn by the same GPU renderer the desktop app uses, no screenshots, the real component running in a WebGPU canvas. We now host that live build in the docs, where it loads only when you open it and unloads the moment you leave, so the WebAssembly payload never sits idle in the background.
Needs a WebGPU-capable browser (Chrome/Edge 113+, or recent Safari/Firefox). Runs entirely on your machine - nothing is sent anywhere.
Two lexers, dialect-aware parsers, and a span-faithful tree
SQL isn't one language. T-SQL, Postgres, MySQL, Oracle, SQLite, ClickHouse and DuckDB disagree on quoting, comments, and half their syntax - and Redis isn't SQL at all - so the front of the pipeline treats dialect as a first-class parameter, not an afterthought.
A dialect-table-driven lexer
Tokenization rules - quoting styles, comment syntax, operator sets - are driven by per-dialect tables rather than hard-coded branches, so adding or correcting a dialect doesn't mean forking the lexer.
Recovering parsers, per dialect
Each supported dialect gets its own recovering parser that stays span-faithful to the source text - meaning the parsed tree keeps exact source positions, so IntelliSense and diagnostics can point at precisely the right character even when the query around it is malformed.
A second, faster scanner - on purpose
A separate single-pass tokenizer powers the filesystem-overlay object model (see below). It's intentionally simpler and faster than the full lexer/parser stack, because it runs over every file in a schema-as-code repo, not just the query you're actively editing.
A real T-SQL Language Server
On top of the parser stack sits a genuine LSP implementation over JSON-RPC/stdio - hover, diagnostics, semantic tokens, and completion - so the same intelligence that powers the editor is reachable by any LSP-speaking client.
Two collectors, one atomic snapshot
This is the part that makes the "local files can outrank the live database" trick in the DBA guide actually work - and stay fast.
// conceptual shape, not literal source generation += 1; let db_task = spawn(collect_from_database(generation)); let fs_task = spawn(collect_from_filesystem(generation)); let (db, fs) = join(db_task, fs_task).await; let merged = merge_by_provenance(db, fs); // newer mtime/hash wins snapshot.swap(merged); // atomic - readers never see a half-built model
Database and filesystem, collected concurrently
sqlly-modelengine runs a database collector and a filesystem collector in parallel, each tagged with a generation counter, then fans them back in through a barrier merge before anything downstream sees the result.
Freshness is measured, not assumed
The merge compares modification timestamps and content hashes between what the live database reports and what's on disk, so a local file only overrides the database when it's genuinely newer - not just present.
Readers never see a torn model
The merged model swaps in as one atomic snapshot per generation. IntelliSense and validation always read a complete, internally-consistent model - either the previous generation or the new one, never a partial rebuild in progress.
Confidence is tracked as data, not a guess
Where the model can't be fully validated - a rough sketch of DDL, a migration mid-edit - the engine tags that region with explicit Provenance and Fidelity metadata and runs completion against it through a distinct two-scope validation path, rather than silently upgrading a guess to a fact.
Local model inference, with cloud as an opt-in fallback
The AI story is genuinely layered - from fully local inference to a cloud provider you explicitly choose. Here's what's real today in the GPUI client, plainly labeled.
Fill-in-the-middle autocomplete
Code-style FIM completion is designed to run against a small local model (Qwen2.5-Coder class) for inline suggestions as you type. The prompt rendering, model/token catalog, and preferences are fully ported and unit-tested; the streaming generation transport is the piece still being wired.
An assistant panel with real architecture
The AI panel is a ported, structured chat surface - not a stub - for asking questions about your schema and getting generated SQL back. The panel and its state are real; the provider plumbing behind it is being connected now.
Bring your own local model via Ollama
If you're already running models through Ollama, SQLly wires straight into it for schema Q&A, embeddings, and search. The health/install state machine is ported and drives the UI; the process-management layer that finds and launches ollama is being finished.
An MLX sidecar for Apple Silicon
Local inference on Apple Silicon runs through a small external MLX sidecar process that SQLly talks to over HTTP - keeping the heavy ML runtime out of the main app process. The value types and health reporting are in; launching and managing the sidecar is in flight.
On-device vector search (LanceDB)
The plan is a fully local vector store for schema/embedding search via LanceDB. Today that crate is a deliberate stub - every call returns a clear, honest error rather than silently no-opping - while the rest of the pipeline is finished around it.
Cloud providers, opt-in and explicit
When you choose to reach past on-device models, the cloud provider picker (OpenRouter, OpenAI, Claude) and its preferences pane are in place - nothing leaves your machine unless you point it there yourself. The HTTP transport behind the picker is being wired up.
The rest of the Rust workspace
The engine is split into focused crates rather than one monolith - most reachable headlessly through sqlly-cli for scripting or testing without the UI at all.
| Crate | What it owns |
|---|---|
| sqlly-core | Connection profiles, command/event models, sessions, app preferences, crash-recovery snapshots. |
| sqlly-engines | The native drivers behind one transport trait: SQL Server over TDS (Tiberius-based), PostgreSQL, MySQL/MariaDB, SQLite, libSQL/Turso over Hrana, DuckDB in-process, ClickHouse over HTTP, Oracle through Instant Client, and Redis over RESP. |
| sqlly-query | Query execution/streaming, IntelliSense completion plumbing, query history, tokenizer glue. |
| sqlly-metadata | Object Explorer and catalog metadata models. |
| sqlly-scripting | Identifier quoting, script generation, and the browsable script template catalog. |
| sqlly-plans | Execution-plan (ShowPlan XML) parsing and plan warnings. |
| sqlly-admin | Backup/restore, activity monitor, table editing, query store, extended events, import/export, schema compare, and designer SQL generation. |
| sqlly-security | SQL and Entra authentication providers, secret-store abstractions, security-catalog SQL. |
| sqlly-catalog | Object "cards," hybrid full-text + vector search over schema, and AI cache invalidation. |
| sqlly-testkit | SQL Server container and fixture/golden-file test helpers used across the workspace's own test suite. |
A real script template catalog
The templating system uses the same <name,type,default> placeholder syntax SSMS templates use - deliberately familiar rather than a novel DSL you'd have to learn from scratch.
Batch scripts, one statement at a time
The lexer exposes a statement-splitter used by the mutation-review path, so a multi-statement script can be reasoned about - and safety-checked - statement by statement instead of as one opaque blob.
The performance claims are backed by the test suite
A latency budget test for completions
The IntelliSense crate includes a test asserting cached completions stay under an explicit latency budget - performance regressions on the completion path fail CI, they don't just get noticed by a user months later.
Real-subprocess integration tests
Beyond unit tests, the CLI engine is exercised with real-subprocess integration tests that set actual environment variables and assert on correlated logs - covering ping/response correlation, unanswered-request detection, and graceful-shutdown log trimming.
Golden-file fixtures against a real SQL Server
sqlly-testkit spins up an actual SQL Server container for integration tests, so the query and metadata layers are tested against a real engine, not just mocked responses.
That’s the tour.
If you spot something here that's drifted from the source - the codebase moves fast - tell us. We'd rather fix the page than let it quietly go stale.