// for the nerds

What’s actually running under the hood

No marketing gloss past this point. This is the real architecture - the crates, the data structures, the design decisions - for the SQLly GPUI client, which is the one under active development and the one every screenshot on this site comes from.

// the frontend

A GPU-rendered native UI, not a webview

The window you're looking at isn't Electron and isn't a browser engine wearing a native-looking skin.

SQLly - GPUI editor
SQLly editor rendered natively via GPUI, showing IntelliSense keyword completion

Every pixel here - text layout, the completion popup, the syntax highlighting - is drawn by GPUI's own GPU-accelerated renderer.

πŸ–ΌοΈ
gpui βœ“ functional

The same rendering approach as Zed

SQLly's client is built on GPUI, the Rust UI framework written for the Zed editor. It's a retained/immediate hybrid GPU-accelerated toolkit - not a DOM, not AppKit views underneath, just Rust drawing directly to the GPU.

🧩
sqlly-datatable βœ“ functional

The result grid is its own crate

Rather than hand-rolling grid virtualization inline, the result grid consumes a dedicated published crate - sqlly-datatable - handling virtualization, selection, sorting, and filtering as a standalone component.

πŸ–₯️
one codebase, three OSes βœ“ functional

macOS, Windows, and Linux from one tree

The GPUI client targets all three desktop platforms from the same Rust codebase, with per-OS packaging: an .app bundle on macOS (Apple Silicon only), a .desktop entry on Linux, and a native .exe on Windows.

// open source, running in your browser

Try sqlly-datatable - the actual result grid, as WebAssembly

The result grid is its own open-source crate, and its sample app compiles to WebAssembly on GPUI’s web backend - drawn by the same GPU renderer the desktop app uses, no screenshots, the real component running in a WebGPU canvas. We now host that live build in the docs, where it loads only when you open it and unloads the moment you leave, so the WebAssembly payload never sits idle in the background.

Needs a WebGPU-capable browser (Chrome/Edge 113+, or recent Safari/Firefox). Runs entirely on your machine - nothing is sent anywhere.

// parsing sql, properly

Two lexers, dialect-aware parsers, and a span-faithful tree

SQL isn't one language. T-SQL, Postgres, MySQL, Oracle, SQLite, ClickHouse and DuckDB disagree on quoting, comments, and half their syntax - and Redis isn't SQL at all - so the front of the pipeline treats dialect as a first-class parameter, not an afterthought.

πŸ”€
sqlly-sqllex βœ“ functional

A dialect-table-driven lexer

Tokenization rules - quoting styles, comment syntax, operator sets - are driven by per-dialect tables rather than hard-coded branches, so adding or correcting a dialect doesn't mean forking the lexer.

🌳
sqlly-sqlparse βœ“ functional

Recovering parsers, per dialect

Each supported dialect gets its own recovering parser that stays span-faithful to the source text - meaning the parsed tree keeps exact source positions, so IntelliSense and diagnostics can point at precisely the right character even when the query around it is malformed.

🧭
sqlly-tokenizer βœ“ functional

A second, faster scanner - on purpose

A separate single-pass tokenizer powers the filesystem-overlay object model (see below). It's intentionally simpler and faster than the full lexer/parser stack, because it runs over every file in a schema-as-code repo, not just the query you're actively editing.

πŸ“‘
sqlly-lsp βœ“ functional

A real T-SQL Language Server

On top of the parser stack sits a genuine LSP implementation over JSON-RPC/stdio - hover, diagnostics, semantic tokens, and completion - so the same intelligence that powers the editor is reachable by any LSP-speaking client.

// the schema model engine

Two collectors, one atomic snapshot

This is the part that makes the "local files can outrank the live database" trick in the DBA guide actually work - and stay fast.

// conceptual shape, not literal source
generation += 1;
let db_task   = spawn(collect_from_database(generation));
let fs_task   = spawn(collect_from_filesystem(generation));
let (db, fs)  = join(db_task, fs_task).await;
let merged    = merge_by_provenance(db, fs);  // newer mtime/hash wins
snapshot.swap(merged);  // atomic - readers never see a half-built model
πŸ”€
parallel collection βœ“ functional

Database and filesystem, collected concurrently

sqlly-modelengine runs a database collector and a filesystem collector in parallel, each tagged with a generation counter, then fans them back in through a barrier merge before anything downstream sees the result.

⏳
staleness by timestamp + hash βœ“ functional

Freshness is measured, not assumed

The merge compares modification timestamps and content hashes between what the live database reports and what's on disk, so a local file only overrides the database when it's genuinely newer - not just present.

πŸ“Έ
generation-staged snapshots βœ“ functional

Readers never see a torn model

The merged model swaps in as one atomic snapshot per generation. IntelliSense and validation always read a complete, internally-consistent model - either the previous generation or the new one, never a partial rebuild in progress.

🌫️
incomplete regions βœ“ functional

Confidence is tracked as data, not a guess

Where the model can't be fully validated - a rough sketch of DDL, a migration mid-edit - the engine tags that region with explicit Provenance and Fidelity metadata and runs completion against it through a distinct two-scope validation path, rather than silently upgrading a guess to a fact.

// the on-device ai stack

Local model inference, with cloud as an opt-in fallback

The AI story is genuinely layered - from fully local inference to a cloud provider you explicitly choose. Here's what's real today in the GPUI client, plainly labeled.

⚑
fim completion ◐ in progress

Fill-in-the-middle autocomplete

Code-style FIM completion is designed to run against a small local model (Qwen2.5-Coder class) for inline suggestions as you type. The prompt rendering, model/token catalog, and preferences are fully ported and unit-tested; the streaming generation transport is the piece still being wired.

πŸ—£οΈ
ai chat panel ◐ in progress

An assistant panel with real architecture

The AI panel is a ported, structured chat surface - not a stub - for asking questions about your schema and getting generated SQL back. The panel and its state are real; the provider plumbing behind it is being connected now.

πŸ¦™
ollama ◐ in progress

Bring your own local model via Ollama

If you're already running models through Ollama, SQLly wires straight into it for schema Q&A, embeddings, and search. The health/install state machine is ported and drives the UI; the process-management layer that finds and launches ollama is being finished.

🧭
sqlly-mlx ◐ in progress

An MLX sidecar for Apple Silicon

Local inference on Apple Silicon runs through a small external MLX sidecar process that SQLly talks to over HTTP - keeping the heavy ML runtime out of the main app process. The value types and health reporting are in; launching and managing the sidecar is in flight.

πŸ—ƒοΈ
local vector cache β—‹ planned

On-device vector search (LanceDB)

The plan is a fully local vector store for schema/embedding search via LanceDB. Today that crate is a deliberate stub - every call returns a clear, honest error rather than silently no-opping - while the rest of the pipeline is finished around it.

☁️
cloud fallback ◐ in progress

Cloud providers, opt-in and explicit

When you choose to reach past on-device models, the cloud provider picker (OpenRouter, OpenAI, Claude) and its preferences pane are in place - nothing leaves your machine unless you point it there yourself. The HTTP transport behind the picker is being wired up.

Why call out what's unfinished? Because the alternative is a features page that quietly stops being true the moment someone reads the source. The model picker UI, for instance, is further along in the legacy Swift app than in GPUI today - we'd rather tell you that than have you discover it.
// everything else in the engine

The rest of the Rust workspace

The engine is split into focused crates rather than one monolith - most reachable headlessly through sqlly-cli for scripting or testing without the UI at all.

CrateWhat it owns
sqlly-coreConnection profiles, command/event models, sessions, app preferences, crash-recovery snapshots.
sqlly-enginesThe native drivers behind one transport trait: SQL Server over TDS (Tiberius-based), PostgreSQL, MySQL/MariaDB, SQLite, libSQL/Turso over Hrana, DuckDB in-process, ClickHouse over HTTP, Oracle through Instant Client, and Redis over RESP.
sqlly-queryQuery execution/streaming, IntelliSense completion plumbing, query history, tokenizer glue.
sqlly-metadataObject Explorer and catalog metadata models.
sqlly-scriptingIdentifier quoting, script generation, and the browsable script template catalog.
sqlly-plansExecution-plan (ShowPlan XML) parsing and plan warnings.
sqlly-adminBackup/restore, activity monitor, table editing, query store, extended events, import/export, schema compare, and designer SQL generation.
sqlly-securitySQL and Entra authentication providers, secret-store abstractions, security-catalog SQL.
sqlly-catalogObject "cards," hybrid full-text + vector search over schema, and AI cache invalidation.
sqlly-testkitSQL Server container and fixture/golden-file test helpers used across the workspace's own test suite.
πŸ“
templates βœ“ functional

A real script template catalog

The templating system uses the same <name,type,default> placeholder syntax SSMS templates use - deliberately familiar rather than a novel DSL you'd have to learn from scratch.

πŸ”¬
split_statements βœ“ functional

Batch scripts, one statement at a time

The lexer exposes a statement-splitter used by the mutation-review path, so a multi-statement script can be reasoned about - and safety-checked - statement by statement instead of as one opaque blob.

// tested, not just fast

The performance claims are backed by the test suite

βœ…
βœ“ functional

A latency budget test for completions

The IntelliSense crate includes a test asserting cached completions stay under an explicit latency budget - performance regressions on the completion path fail CI, they don't just get noticed by a user months later.

βœ…
βœ“ functional

Real-subprocess integration tests

Beyond unit tests, the CLI engine is exercised with real-subprocess integration tests that set actual environment variables and assert on correlated logs - covering ping/response correlation, unanswered-request detection, and graceful-shutdown log trimming.

βœ…
βœ“ functional

Golden-file fixtures against a real SQL Server

sqlly-testkit spins up an actual SQL Server container for integration tests, so the query and metadata layers are tested against a real engine, not just mocked responses.

That’s the tour.

If you spot something here that's drifted from the source - the codebase moves fast - tell us. We'd rather fix the page than let it quietly go stale.