Cloud-neutral data and AI platform

Your lakehouse, in your jurisdiction, on the engine that fits.

Query one governed Iceberg catalog with DataFusion, Ballista, Apache Spark, Sail or SparkHouse. Run it in the cloud and region you choose, or in your own Kubernetes cluster.

SQL endpoint · engine AUTO Illustrative

        
  • 5 enginesone SQL endpoint
  • GCP · AWS · Kubernetesincluding your own cluster
  • Apache Iceberg + Polarisone catalog per organization
  • Hash-chained audittamper-evident by design

The problem

Regulated and regional teams are offered two bad options.

Take a US hyperscaler's SaaS, and the data leaves your jurisdiction. Build a lakehouse from parts, and you spend a year on platform work before the first analyst runs a query.

Either way, one engine rarely fits every workload, and every vendor catalog turns into its own silo.

SparkHouse is the control plane in the middle. It decides where compute runs, which engine runs each statement, and who can see what, while your tables stay in open Iceberg.

How it works

Choose where. Connect your data. Query.

  1. Choose where

    Pick a region and cloud, or point SparkHouse at your own Kubernetes cluster. Attach a residency policy by country, cloud and region.

  2. Connect your data

    Use the Iceberg catalog every organization gets, or federate AWS Glue, Databricks Unity Catalog, Snowflake or any Iceberg REST catalog.

  3. Query

    Send SQL to one endpoint. The router picks the engine for each statement, and every run shows which one it chose.

Control plane above, data planes below One SparkHouse control plane manages data planes in a Canadian region, an EU region and a customer's own cluster. Customer data stays inside each data plane. SparkHouse control plane API · router · policies · metering · audit · operator console metadata only DATA PLANE Region, country A Engines Iceberg tables Catalog, credentials and data stay in this region DATA PLANE Region, country B Engines Iceberg tables A different cloud is fine. Policy decides placement. YOUR CLUSTER (BYOC) Your Kubernetes Agent Engines The agent calls out to the control plane. Nothing calls in.
The control plane holds metadata. Data, engines and credentials live in the data plane.

Sovereignty

Residency is enforced where compute is placed, not written in a policy PDF.

  • Policies by country, cloud and region. Each organization has a home region and can add extra ones.
  • Placement enforces the policy. A statement that would run outside the allowed set is refused, as the example above shows.
  • Bring your own cluster. The data-plane agent runs engines in your Kubernetes. Compute, data and credentials stay in your environment.
  • Shipped on GCP, AWS and any Kubernetes. GKE Autopilot, EKS Auto Mode, minikube and kind are supported targets.

Workspace

The tools analysts and engineers expect, in one console.

SQL editor, notebooks, jobs and a catalog browser share one workspace. Every statement goes through the same router, policies, audit log and metering.

SQL editor

Worksheets with a schema browser, parameters, explain, formatting, run history and result charts. A command palette gets you anywhere from the keyboard.

Notebooks

SQL, Python and Markdown cells in one document, with results, tables and charts under each cell. Import and export Jupyter .ipynb.

Jobs

Ordered SQL tasks with cron schedules and time zones, retries, timeouts, logs and cancel. Notebooks run as job tasks, with parameters you can override per run.

Catalog Explorer

Schemas, tables, views, columns, sample data and history, with comments, tags and permissions beside them.

Runs and history

Every statement is a run: who ran it, on which engine, where it was placed and how long it took.

Usage

Compute and notebook uptime metered per user, endpoint, engine and placement.

Notebooks that run where your policy says.

  • SQL cells run on the notebook's endpoint through the platform router, so they are routed, audited and visible in Runs.
  • Python cells run in your own Jupyter kernel. A notebook with only SQL and Markdown never starts one.
  • Kernels act as you, with a short-lived token limited to running statements, reading runs and using the catalog in that workspace.
  • With an Apache Spark endpoint, spark in a kernel is a Spark Connect session.
daily_revenue.ipynbIllustrative
SQL
%%sql
SELECT region, sum(amount) AS revenue
FROM sales.orders
WHERE day = '{{day}}'
GROUP BY region
Python
df = sh.sql("SELECT * FROM sales.daily")
display(df)
Markdown
# Notes for the finance review

Engines

Pick an engine, or let AUTO decide.

AUTO routes each statement by lane (write, heavy, interactive) and by lowest effective cost. The run view shows the engine and why it was chosen.

DataFusion

Flight SQL

Fast interactive queries, with warm pools on dedicated data planes.

Ballista

Distributed

The same Arrow-native stack, spread across workers for heavy statements.

Apache Spark 4.1

Opt-in

Familiar engine for heavy ETL, reached over Spark Connect, standalone or on native Kubernetes.

Sail (LakeSail)

Spark Connect

Spark-compatible engine written in Rust on DataFusion, with no JVM. Runs as a single server or as a driver that launches its own workers on Kubernetes.

SparkHouse engine

Preview

Our own Go engine. It is in the product and still maturing, so we label it as a preview.

Open catalog

One governed Iceberg catalog per organization. No table lock-in.

Every organization gets its own Apache Polaris catalog on highly available PostgreSQL. Browse schemas, tables, views, columns, sample data and history in the Catalog Explorer, and manage comments, tags and permissions beside them.

Federate what you already run

  • AWS Glue
  • Databricks Unity Catalog
  • Snowflake Horizon
  • Snowflake Open Catalog
  • Any Iceberg REST catalog

Federated tables use per-table vended credentials, and you choose read-only or writable.

Serverless and cost

Compute that starts when you query and stops when you do not.

  • SQL endpoints in sizes XS to XL, with auto-resume and auto-stop.
  • Clusters with autoscaling.
  • Usage by endpoint, user, engine and placement.
  • A per-statement cost estimate for each placement.

Governance and operations

Auditable for you, operable for whoever runs it.

  • A hash-chained audit log per organization, plus a platform chain.
  • Workspaces, roles, and compute and residency policies.
  • An operator console for tenants and health.
  • Engine version channels (preview, current, deprecated, retired) with pins and upgrades at next start.

For developers

A REST API, Flight SQL, and a CLI.

shctl covers endpoints, clusters, SQL, jobs, runs, catalog, audit and benchmarks.

shctl
# sign in to your SparkHouse host
shctl login --host https://sparkhouse.example --token shp_...

# create a Ballista cluster with four workers
shctl clusters create --name etl --engine ballista --workers 4

# run a statement on it
shctl sql --endpoint etl -e "SELECT count(*) FROM lineitem"

FAQ

Straight answers.

Where does my data live?

In the data plane you choose: a region on GCP or AWS, or your own Kubernetes cluster. The control plane holds metadata, not table data.

Can I bring my own cluster?

Yes. The data-plane agent runs engines in your cluster and connects outbound to the control plane.

Which clouds are supported?

GCP, AWS, and any Kubernetes cluster.

Is it open source?

The engines and formats underneath are open source (Iceberg, Polaris, DataFusion, Ballista, Apache Spark, Sail). The SparkHouse platform itself is proprietary.

Can I run my PySpark unchanged?

Inside a platform notebook, yes: on an Apache Spark endpoint, spark is a Spark Connect session. Pointing your own PySpark at SparkHouse from outside is not available yet, and a Spark Connect gateway is planned.

How is it priced?

Through pilot terms. There is no self-serve signup or public price list yet.

Pilot

Request a pilot.

We are working with a small number of design partners who have a residency or sovereignty constraint and several teams on one lakehouse. Tell us about your setup and we will reply with next steps.

We use your details only to reply to this request.