Fluid Forge vs alternatives
If you already use dbt, Dagster, Terraform, OPA, or Snowpark, you might be wondering which problem Fluid Forge is solving that those tools don't. Honest answer: none of them, individually. Forge's value is unifying the four contracts every data product has — schema, infrastructure, orchestration, policy — into one declarative file so they can't drift from each other.
This page is the comparison page we'd want to read if we were evaluating Forge cold. It includes honest losses (dbt has a bigger ecosystem, Dagster has better asset lineage, Snowpark has tighter Snowflake integration) so you can decide whether the unification trade is worth it for your team.
The 30-second answer
| You have… | Forge fits when… | Forge is overkill when… |
|---|---|---|
| One warehouse, one team, dbt + a CI runner | You'll need governance, agent boundaries, or to add a second cloud later | You'll never leave that warehouse, and policy/agent boundaries aren't on the roadmap |
| Multi-cloud already (BQ + Snowflake + Athena) | You're rewriting the same contract three times in three formats | You have dedicated platform engineers per cloud and they don't mind the duplication |
| Compliance pressure (SOX / GDPR / HIPAA) | You want governance, sovereignty, and AI access boundaries in the same file as the schema | You've already centralised governance on a separate plane (Immuta, Privacera) |
| Building agentic data products | You want declarative agentPolicy gating LLM access at read-time | Your agents query through a separate runtime layer that already enforces this |
| Prototyping, don't know what you need yet | Start local DuckDB, graduate when you do | You're at a code-only POC stage with no contract requirements |
The unification table
| Tool | What it owns | What it doesn't own | What you wire by hand today |
|---|---|---|---|
| dbt Core | SQL transformations, lineage, refs/sources, dbt tests | Provisioning, IAM, multi-cloud, agent governance, sovereignty | Terraform for IAM + Airflow for orchestration + OPA for policy + custom JSON for AI access |
| Dagster | Asset orchestration, asset checks, sensors, schedules | Schema-as-contract, native cloud IAM emission, multi-cloud abstraction | Terraform/Pulumi for infra + dbt for SQL + your own RBAC layer + your own AI gating |
| Airflow / Composer / MWAA | DAG scheduling, task retries, sensors | Schema, IAM, multi-cloud abstraction, contract validation | Resources + provider-specific operators + dbt + IAM + governance code |
| Terraform | Cloud infrastructure (any provider) | Schema, quality rules, transformations, orchestration | Tables + SQL + DAGs + dbt project + SLA checks + lineage emitters |
| OPA / Rego | Policy evaluation engine | Schema, transformations, cloud-native IAM emission, AI/agent boundaries | Compiling Rego → BigQuery row-level security / Snowflake masking / AWS IAM bindings |
| Snowpark / dbt-Snowflake | Snowflake-native data plane (UDFs, stored procs, dbt-cloud features) | Multi-cloud portability, agent governance | Rewriting everything if you ever leave Snowflake; a separate AI gate layer |
| Fluid Forge | All four — schema + infra + orchestration + policy + AI gating — in one contract.fluid.yaml | Bespoke per-vendor extreme tuning (e.g. Snowflake search optimization, BigQuery BI Engine reservations) | Anything genuinely vendor-specific that doesn't have a cross-cloud abstraction |
Forge vs dbt
The most common comparison. dbt is the dominant SQL transformation framework; Forge is sometimes mistaken for a dbt competitor. It isn't.
Where dbt wins
- Ecosystem maturity — 1000s of dbt packages, dbt-utils, dbt-expectations, dbt-snowflake, dbt-bigquery. Forge has none of these; it uses dbt for the SQL layer when
engine: dbtis selected. - SQL-only teams — if your data product is a SQL transformation and nothing more, dbt is simpler. Forge's contract surface (schema, dq.rules, accessPolicy, agentPolicy, sovereignty) is overhead you don't need.
- Community + hiring — dbt has been around since 2016. There are dbt analysts on the market. Forge engineers are still rare.
- dbt Cloud / dbt Mesh — if you're committed to the dbt ecosystem and willing to pay for the cloud product, you get IDE, CI, lineage, and discovery without leaving dbt-land.
Where Forge wins
- Multi-cloud portability — one base contract, with a per-cloud overlay that changes only the binding (platform, format, location, and on fluid-schema 0.7.6 the principals).
fluid apply --env gcpemits BigQuery datasets and IAM;--env awsemits S3, Glue and Lake Formation. dbt's adapter layer handles SQL dialect differences but not the surrounding infrastructure (datasets, IAM, regions). See One contract, two clouds and Switch clouds. - Governance as part of the contract —
fluid applyprovisions native controls from the contract: BigQuery dataset IAM members and Data Catalog policy tags on GCP, Lake Formation grants with excluded columns on AWS. Coverage differs by cloud, and as of 0.18.1 Snowflake reads none of these fields; see the per-cloud table. dbt'sgrantsconfig sets privileges on the models it builds. - Agent governance —
agentPolicydeclares which LLMs can read which fields, with audit logging. dbt has no concept of this. - Sovereignty —
sovereignty.jurisdictionandallowedRegionsare enforced before deploy:fluid validaterefuses an out-of-policy region, and on AWS and GCP so doesfluid apply. dbt does not check where data lives. - Local-first development —
pipx install "data-product-forge[local]"andfluid applywork entirely on DuckDB with no cloud account. dbt-core works locally too, but the typical dbt onboarding assumes a warehouse. - The contract is the source of truth — dbt models describe transformations; Forge contracts describe the entire data product (schema, transformation, exposure, governance). Different scope.
How they fit together
Forge does not replace dbt. The recommended pattern is engine: dbt inside a Forge contract — Forge handles provisioning, IAM, policy, and AI gating; dbt handles the SQL transformation. Both worlds, no overlap.
builds:
- id: customer_metrics
pattern: hybrid-reference # required whenever properties is set
engine: dbt # ← dbt does the SQL
repository: ./dbt
properties:
model: customer_metrics # required: the dbt model to build
target: prod # forwarded as dbt --target
Forge vs Dagster
A sharper philosophical comparison. Dagster's asset-oriented orchestration is the closest thing in the OSS world to Forge's contract-first model.
Where Dagster wins
- Asset orchestration depth — software-defined assets, partitioned assets, asset checks, asset sensors. Forge has builds + exposes; Dagster has a richer asset graph model with native lineage.
- Python-first — write your business logic in Python and let Dagster orchestrate. Forge's primary interface is YAML; Python is for builds via
engine: python. - Dagster Cloud / Plus — hosted control plane, branch deployments, concurrency controls. The Forge CLI has no hosted offering; for a web catalog and run history on top of it, see the Command Center.
- Op-level retries, backfills, and sensors — operationally rich. Forge defers orchestration to the chosen scheduler (
fluid generate schedule --scheduler dagster | airflow | prefect).
Where Forge wins
- Contract-first vs pipeline-first — Forge starts with "what is this data product" (schema, SLAs, governance). Dagster starts with "how is it computed" (assets, ops). Different first principle.
- Native cloud IAM emission — same as the dbt comparison:
fluid applyprovisions the grants and column restrictions the contract declares, where the provider supports them. Dagster has IO managers and resources, but no equivalent IAM compilation. - Multi-cloud abstraction at the contract layer — Dagster's resources are typed to a specific platform per pipeline. In Forge the product is declared once and each cloud is an overlay on the binding.
- Agent governance as a first-class contract field — Dagster doesn't model this.
- Smaller surface for read-only data product producers — if your team's job is to publish a data product (not to operate a complex pipeline), Forge's mental model is lighter than Dagster's.
How they fit together
fluid generate schedule --scheduler dagster emits a Dagster job from the Forge contract. You get Forge's contract + governance + multi-cloud, plus Dagster's runtime. Forge owns the what; Dagster owns the how.
Forge vs Terraform
The infrastructure-as-code comparison. Terraform is universal; Forge is data-specific.
Where Terraform wins
- Universality — Terraform manages everything: VPCs, Kubernetes clusters, IAM roles, S3 buckets, Stripe products, Cloudflare DNS. Forge only manages data products.
- Provider ecosystem — 3000+ Terraform providers. Forge has 4 primary (local, gcp, aws, snowflake) plus a Custom Provider SDK.
- State management — Terraform's state is a battle-tested model with locking, partial apply, drift detection. Forge has lighter state semantics tuned for data products.
- Mature blast-radius controls —
terraform plan,terraform import, workspaces, modules. Forge hasfluid planbut the surrounding tooling is younger.
Where Forge wins
- Data-product-specific abstractions —
exposes,dq.rules,agentPolicy,sovereignty,lineage— try expressing these in Terraform. You can't, except as ad-hoc resource configurations that drift. - Schema-as-contract — Forge validates the schema against the actual deployed table. Terraform doesn't know what a "schema" is.
- One contract, three clouds — Terraform requires three different sets of resource definitions to deploy "the same" table on BigQuery, Snowflake and Athena. Forge keeps one base contract and a binding overlay per cloud.
- Compiles to OpenTofu/Terraform —
fluid generate iacemits a deterministicmain.tf.jsonmodule when you want to inherit your Terraform pipeline downstream.
How they fit together
Forge sits on top of OpenTofu/Terraform: on AWS, GCP and Snowflake, fluid apply runs the module fluid generate iac writes. Teams use Forge for the data-product layer and keep their Terraform pipeline for the surrounding infrastructure (VPCs, networking, organisation policy). To review before anything runs, commit the output of fluid generate iac and apply it with tofu.
Forge vs Snowpark / dbt-Snowflake / dbt Cloud
The single-vendor stack. If you've gone all-in on Snowflake, this is who you're really comparing against.
Where Snowflake-stack wins
- Vendor-specific feature depth — Snowpark UDFs, stored procs, search optimization, query acceleration, time-travel, zero-copy clones. Forge can drive Snowflake but doesn't surface every Snowflake-specific tuning knob.
- Single bill, single support contract — one vendor relationship, one billing system, one support team.
- Snowflake Cortex / native LLM — if your strategy is "Snowflake will be the AI plane too", Cortex is integrated. Forge has no Cortex integration; its agent surface is the MCP output port.
- dbt Cloud — IDE, CI, lineage, jobs all hosted. Forge has none of the hosted UX yet.
Where Forge wins
- Optionality — the day Snowflake pricing changes or a faster engine emerges (DuckDB, RisingWave, Materialize), you can move. With Snowpark you cannot.
- Local development — Forge runs end-to-end on DuckDB with no Snowflake account. Snowpark needs a Snowflake account from day one.
- Cross-warehouse — if part of your portfolio is on BigQuery and part on Snowflake, Forge unifies the contract. Snowpark + BigQuery is two parallel stacks.
- Open-source, Apache-2.0 — no vendor lock at the orchestration layer.
How they fit together
Use binding.platform: snowflake for your Snowflake-resident data products. Snowflake table settings the emitter reads, such as binding.properties.cluster_by, live in the binding, so the rest of the contract stays portable. As of 0.18.1, governance does not carry over: on Snowflake, access grants, column restrictions and masking from the contract are not emitted.
When NOT to use Fluid Forge
The honest list. Adoption decisions are easier when you know where the tool actively isn't right.
You're a one-warehouse, one-team analytics shop
If you have one Snowflake account, one dbt project, and no governance/compliance pressure, the Forge contract is overhead. Stay on dbt. You can still adopt fluid validate standalone if you ever want schema-as-contract testing.
You need real-time streaming with sub-second SLA
Forge's batch-and-mini-batch model fits 5-minute to 24-hour latency. For sub-second streaming (CDC → live materialized view), look at Materialize, RisingWave, or a Kafka + Flink stack. Forge's engine: kafka-connect and engine: debezium cover ingestion; the streaming compute layer is out of scope.
You're doing pure ML feature engineering
Forge's contract surface is general-purpose. For ML feature stores specifically, Feast, Tecton, or Hopsworks have richer concepts (feature views, point-in-time joins, online/offline parity). Forge can express the output feature table as a contract but doesn't have feature-store-specific abstractions.
Your team already has a working four-tool stack
If Terraform + dbt + Airflow + OPA is humming and the team is happy, the migration cost to a unified contract may not pay off. Consider adopting Forge incrementally — start with fluid validate for contract testing, expand only if/when the cross-tool drift starts to bite.
You need a hosted control plane today
Forge is currently CLI + GitHub Actions / GitLab CI / Jenkins / Tekton (any CI). The CLI itself has no hosted offering; the Command Center is a separate web control plane that the CLI can report to. If you need a SaaS UI for your data team to onboard non-engineers, Dagster Cloud / dbt Cloud are mature options today; the equivalent Forge offering is on the roadmap but not shipped.
Bottom line
Pick Fluid Forge when:
- You're shipping data products across 2+ clouds (or expect to within 12 months)
- Governance / compliance is part of the contract, not an after-the-fact audit
- You're building agentic data products and need declarative LLM access boundaries
- You want a single source of truth for schema, infra, orchestration, policy, AI gating
Pick something else when:
- You're committed to one warehouse forever and have no governance pressure → use dbt
- You need rich orchestration with pipeline-first semantics → use Dagster
- You need arbitrary cloud infrastructure (not just data products) → use Terraform
- You need sub-second streaming → use Materialize / RisingWave
- You need a feature store → use Feast / Tecton
Use them together when:
- You want Forge's contract + governance with dbt's SQL:
engine: dbt - You want Forge's contract with Dagster's runtime:
fluid generate schedule --scheduler dagster - You want Forge's contract layered on top of Terraform:
fluid generate iac
The unification value is highest when at least two of (multi-cloud, governance, AI gating) apply to your data products. If only one applies, the dedicated tool for that one thing is usually a better fit.
See also
- Concepts overview — the rest of the conceptual reference
- Get Started — install and ship a data product in 30 seconds
- Demos library — the proof, in 30-second SVG casts
- Universal pipeline — the canonical 11-stage CI flow