The AI data engineer that runs on your machine.
DataRig Lite is free for your desktop: describe a pipeline in plain English, review exactly what it builds, then approve — governed Iceberg tables, orchestrated dbt models, and live dashboards, without your data ever leaving your laptop.
Free forever · Windows & macOS · no account needed · Cloud & Enterprise coming soon
- Zero data retention — nothing leaves your machine
- Review-gated agents — nothing runs without approval
- Open Iceberg tables — no lock-in
- All tools built into the rig
Part of the ExaScale DataRig platform: start free on your desktop, scale to managed Cloud, or deploy in your own VPC — same pipelines, same governance, zero rewrites.
- 01 Spark ingestion →
local.user.orders(Iceberg · Polaris) - 02 dbt model
revenue_daily+ freshness & row-count tests - 03 Airflow DAG
orders_daily, scheduled 06:00 - 04 Superset dashboard “Daily Revenue” over Trino
Everything you need, packaged in one install. Nothing hand-wired.
- Apache Iceberg
- Apache Spark
- Trino
- Apache Polaris
- Kafka + Debezium
- Apache Airflow
- dbt Core
- Apache Superset
From plain English to production, in four steps
You stay in control: DataRig proposes the exact plan, you approve the exact hash, your machine runs it.
- STEP 01
Connect
Drop in CSV, JSON, or Parquet files — or connect PostgreSQL / SQL Server for continuous CDC. All local.
- STEP 02
Describe
Say what you want: “load this into Iceberg,” “add dbt tests,” “build a revenue dashboard.”
- STEP 03
Review & approve
Inspect source analysis, plan steps, and generated files. Approve binds to an exact content hash.
- STEP 04
Run & monitor
Watch live Spark logs, staged progress, and status chips — with error-specific recovery if anything fails.
Explore governed data, without leaving the rig
Browse your Iceberg catalog tree, inspect schemas, and run read-only SQL — right beside the chat. AI-suggested queries load here for review; they never auto-run.
SELECT order_date, SUM(amount) AS revenue FROM local.user.orders GROUP BY order_date ORDER BY order_date DESC;
| order_date | revenue |
|---|---|
| 2026-08-25 | 18,204.50 |
| 2026-08-24 | 16,930.00 |
| 2026-08-23 | 21,417.75 |
| 2026-08-22 | 14,882.10 |
All the tools, pre-connected
Specialist consoles are packaged, wired, and one click away — with the DataRig console as your single home base.
The whole pipeline, one conversation
Sources → Iceberg → Airflow → dbt → Superset. Each stage is a governed skill with its own validated contract.
Ingestion
Land CSV, JSON, JSONL, and Parquet files — or HTTPS APIs — into governed Iceberg tables. Full load today; XML profiling is review-ready.
CDC & streaming
Capture changes from PostgreSQL and SQL Server with Kafka + Debezium. Snapshot, updates, deletes, restart recovery — verified end to end.
Transformation
dbt models generated with tests and docs, executed by Airflow against Trino — every artifact reviewable before it runs.
Orchestration
Airflow DAGs created, scheduled, paused, retried, and refreshed in place. Run-scoped docs and live status inside the console.
Data quality
Freshness checks, row-count parity, and null/duplicate tests proposed as part of the plan — not bolted on afterwards.
Analytics & BI
Superset datasets, charts, and dashboards provisioned from your governed business layer, queried through Trino.
Works while you sleep
Approve a pipeline once and DataRig keeps it running. Airflow schedules every load, dbt tests guard every refresh, and failed tasks retry on their own — all on your machine, with the full history one click away.
- Scheduled, not babysat. DAGs are created, scheduled, paused, and refreshed in place.
- Tested every run. Freshness, row-count parity, and null/duplicate checks run with each refresh.
- Self-healing by default. Transient failures retry automatically; the run log shows exactly what happened.
- ✓ Aug 26 06:00 3/3 tests passed 2m 41s
- ✓ Aug 25 06:00 3/3 tests passed 2m 38s
- ↻ Aug 24 06:00 source timeout · retried 1× · passed 4m 12s
- ✓ Aug 23 06:00 3/3 tests passed 2m 35s
- ✓ Aug 22 06:00 3/3 tests passed 2m 44s
Start on your laptop. Never start over.
The pipelines you build in Lite are the same artifacts every edition runs: Iceberg tables, Airflow DAGs, tested dbt models, Superset dashboards. When you scale, your work scales with you — no rewrites, no migration project.
Your machine
Individual engineers, pilots, local-first teams
- Complete local lakehouse + review-gated AI
- All tools built in — zero setup beyond install
- Zero data retention, telemetry off
Our cloud
Teams that outgrew the laptop but not their budget
- Everything in Lite — fully managed and always on
- Automated upgrades, backups, and monitoring
- Transparent usage metering — no opaque credits, no idle cluster sprawl
- Email + Slack support
Your VPC
Regulated orgs that need sovereign control
- Everything in Cloud, deployed inside your network perimeter
- SSO, organization RBAC, audit export
- Customer-managed encryption + private networking
- Contractual SLA with dedicated support
Pipeline definitions carry forward across every edition. Your data stays in open Apache Iceberg formats everywhere.
Every platform locks its AI inside its own wall
ExaScale is the bridge when your data
doesn't live in one wall.
Query across warehouses
Planned: read-only federation to Snowflake, Redshift, Databricks, and BigQuery catalogs — join lakehouse and warehouse data without ETL sprawl, behind the same review-gate stack.
Bring your semantics
Planned: import existing semantic models and metadata from the platforms you already pay for — your definitions travel with you instead of being rebuilt per vendor.
Gated write-back, last
Federation arrives read-only first. Any future write-back adapter lands behind explicit, hash-bound approvals — never silent synchronization.
Designed, not yet scheduled — join the waitlist to shape the priority.
Ship your first governed pipeline today.
Free forever on Windows & macOS. Your data never leaves your machine.