BUSFACTOR.TECH
Delivery Metrics

DORA vs SPACE vs DX Core 4: Picking a Framework (2026)

Cycle anatomy

PICKING A FRAMEWORK (2026)

DORA vs SPACE vs DX Core 4, compared fairly: what each framework measures, where each breaks, and how to choose for 2026 without dashboard theater.

6 receipts in this article ↓

TL;DR: Three frameworks, three different jobs. DORA measures delivery outcomes from telemetry: four keys, recomputable by anyone with the data. SPACE is measurement theory: five dimensions and two rules that stop any single metric from becoming a weapon. DX Core 4 is the packaged, vendor-era standard: board-legible, fast to roll out, and, by explicit design, part survey. The 2026 twist is that Core 4 has the distribution (its maker was acquired by Atlassian for roughly $1B), so the question most teams now face is "do we want our flagship numbers to be telemetry, perceptions, or a blend, and can anyone re-derive them later?" Here's the fair comparison, then a decision list.

If you're weighing the DevEx framework (feedback loops, cognitive load, flow state) rather than Core 4, that triangle is covered in DORA vs SPACE vs DevEx. This page is the buyer's-decision version for the framework that took over procurement conversations in 2025-2026.

DORA vs SPACE: same lineage, different jobs

The two are routinely pitched as rivals; they're a scoreboard and a rulebook.

DORA is the scoreboard: deployment frequency, lead time for changes, change failure rate, and failed-deployment recovery time, validated by a decade of research showing speed and stability rise together. Its strengths are hard edges - the keys are computable from pipelines and git history, no surveys required, which is why they anchor most delivery-metrics setups and survive skeptical audiences. Its limits are equally hard: the keys say nothing about why a number moved, and nothing about the humans producing it.

Note also that the ground is moving. The 2025 report retired the famous elite/high/medium/low tiers in favor of seven team archetypes, so "we're elite" is now a 2024-vintage sentence. The full walkthrough, fine print included, is in DORA metrics explained, with the bands unpacked in deployment frequency benchmarks.

SPACE (Forsgren et al., ACM Queue 2021) is the rulebook: productivity cannot be represented by a single metric, so measure across at least three of its five dimensions - Satisfaction, Performance, Activity, Communication, Efficiency - and include at least one perceptual measure. SPACE's sharpest idea is tension by design: pick metrics that check each other so nothing can be gamed in one direction. It ships no metric values of its own; it's the design review you apply to whatever you build, DORA keys included. (The five dimensions in depth: the SPACE walkthrough.)

So "DORA vs SPACE" resolves cleanly: DORA for the outcomes, SPACE for the guardrails around them. The genuinely new decision is the third contender.

DX Core 4: the vendor-era framework

The DX Core 4, announced in December 2024 by researchers from the same DORA/SPACE lineage, packages the field into four dimensions (Speed, Effectiveness, Quality, Impact), each with one key metric and counterbalancing secondaries. Its rise since has been about distribution as much as content: Atlassian acquired DX for approximately $1B in September 2025, and the framework is now procurement vocabulary. Even rival vendors ship Core 4 rollout support, which is what owning a standard looks like.

The honest read cuts both ways. In Core 4's favor: it's coherent, board-legible, SPACE-compliant by construction, and its authors publish real caveats. The research paper explicitly warns against setting targets on the speed metric, and DX's own research team is refreshingly blunt about AI hype.

The eyes-open part: Core 4 collection deliberately layers system metrics with self-reported surveys and experience sampling, and several of its key numbers are perceptions. The Effectiveness dimension is keyed on the DXI, a composite of 14 Likert-scale survey items benchmarked against DX's research cohort, alongside perceived rate of delivery and perceived quality. Perception data is valuable (SPACE mandates some), but perception as a flagship number inherits a known failure mode: developer self-assessment can be wrong in sign, as the trial evidence in the AI productivity paradox shows. And a survey composite cannot be re-derived from system data later, which matters the day the number is challenged.

The deploy view showing the four DORA measures - deployment frequency, lead time, change failure rate, and time to restore - each with its band.The deploy view showing the four DORA measures - deployment frequency, lead time, change failure rate, and time to restore - each with its band.
The DORA four - measured from your own historyLive product · fictional demo org

Side by side

DORASPACEDX Core 4
What it is4 delivery-outcome metricsMeasurement theory (5 dimensions)Packaged standard (4 dimensions)
Primary dataSystem telemetryMixed; mandates ≥1 perceptualTelemetry + surveys + sampling, by design
Best atHard outcomes, trendsPreventing single-metric theaterOrg-wide rollout, board legibility
Weakest atCauses; team healthIt's guardrails, not numbersFlagship numbers are part perception
Re-derivable later?Yes, from historyThe telemetry parts onlyThe telemetry parts only
Watch out forKeys turned into individual KPIsCited in the deck, ignored in practiceTreating survey composites as facts

Which framework in 2026: the decision list

  1. No delivery baseline yet? DORA keys from telemetry, this quarter. Everything else is decoration until the scoreboard exists.
  2. Designing or expanding a metric set? Apply SPACE as the design review - three-plus dimensions, one perceptual, mutual tension. It will veto your worst ideas for free.
  3. Need a standard the whole org and board can share? Core 4 is the defensible pick: adopt the vocabulary, keep the caveats its own authors publish, and label the survey-derived numbers as what they are.
  4. Any number that must survive an audit, a CFO, or due diligence? Only the telemetry half qualifies, whatever framework it wears. Diligence teams apply exactly this bar - see the engineering checklist acquirers run - and so should you, before someone else does.
  5. Anyone proposing any of these for ranking individuals? All three frameworks' authors are unanimously against it. Rare consensus; respect it.
The teams board: a team-by-team review-flow matrix showing which team reads whose code, with the intra-team diagonal dimmed and nobody ranked.The teams board: a team-by-team review-flow matrix showing which team reads whose code, with the intra-team diagonal dimmed and nobody ranked.
The team board - who reads whose code, no stack rankLive product · fictional demo org

…or an honest read that re-derives every number

The quiet fork under all three logos: frameworks tell you what to measure, not whether the measurement can be checked. Most tools in the engineering-intelligence category will happily render DORA charts and framework dashboards with undisclosed models somewhere inside them. We audited the field on exactly this, and recomputable numbers turned out to be rare. Whichever framework you adopt, hold its implementation to the DORA half's standard: same input rows, same numbers, every rerun, receipts attached. A framework is a vocabulary. Evidence is what settles arguments.

Frequently asked

Is DX Core 4 replacing DORA and SPACE?

No, it's built on top of them. DX Core 4 comes from researchers in the same DORA/SPACE lineage and packages the ideas into four dimensions (Speed, Effectiveness, Quality, Impact), each with a key metric and counterbalancing secondaries. It's become procurement vocabulary fast (rival vendors ship support for rolling it out), but it deliberately mixes system telemetry with self-reported surveys and experience sampling, so adopting it doesn't remove the DORA-vs-SPACE question; it inherits both halves of it.

Which metrics framework should we adopt in 2026?

Decide by data source, not by brand. If you can't yet say how often you deploy or how long a change takes, start with the DORA keys from telemetry, nothing perceptual required. If you're designing a broader metric set, apply SPACE as the guardrails: at least three dimensions, at least one perceptual measure, metrics that check each other. Reach for DX Core 4 when you need a board-legible standard and are comfortable that several of its key numbers are surveys.

Are DORA, SPACE, and DX Core 4 numbers auditable?

Only partly, and the split matters. The DORA keys are re-derivable from pipelines and git history: anyone with the data can recompute them. SPACE mandates at least one perceptual measure, which is honest by design but not recomputable from system data. DX Core 4's key metrics include the DXI (a 14-item Likert composite) and perceived rate of delivery: survey instruments, benchmarked against other orgs' surveys. If a number needs to survive an auditor, a CFO, or a diligence team, the telemetry half is the half that qualifies.

Receipts

Keep reading