All Posts
SaaS Platform

The Regulation Is Not Your Schema

EFRAG published the revised ESRS datapoint list in two versions — one of them a mapping back to the 2023 standard. That mapping is a migration file, and most compliance products cannot run it.

15 min read

On 28 August, EFRAG published the datapoint list for the revised European Sustainability Reporting Standards in two versions: a clean one, and one that maps every datapoint back to the 2023 standard it replaces. The second file is the interesting one. It is a crosswalk — a machine-readable account of what became what — and it exists because the standards body understands something most teams building compliance software do not. A regulation is a versioned artifact, and two versions of it are about to be in force at the same time. If your product encoded the 2023 rules as database columns, that mapping describes a migration you now have to run against live audit data. If it encoded them as rows, the mapping is a seed file.

Why this matters now

The Omnibus I package cleared its final Council vote on 24 February, narrowing CSRD scope to companies with more than 1,000 employees and above €450 million in net annual turnover, and granting a transition exemption to the wave-one companies that fall out of scope for 2025 and 2026. On 3 July the Commission adopted the delegated act revising the ESRS themselves. The revised standards apply to financial years beginning on or after 1 January 2027, with early adoption permitted for financial year 2026 once the act enters into force.

Read those two sentences as a product requirement rather than a policy update. Early adoption is optional, which means that in 2027 your customers will not all be on one version of the standard. Some will file FY2026 under the 2023 ESRS. Some will early-adopt the revised standards for the same year. Some dropped out of scope entirely and still want their numbers. Your product has to render, validate, store and score all of them, in the same quarter, out of the same database.

Then on 28 August, EFRAG published the 2026 draft list of datapoints — in a clean version and a mapping version that references both the 2023 ESRS and the 2024 implementation guidance. Fatal-flaw review runs until 23 October, the final list is expected by the end of the year, and a draft XBRL taxonomy follows. The headline number from EFRAG's technical advice is a 61% reduction in mandatory datapoints, with all voluntary disclosures deleted outright.

Almost everyone has read that 61% as relief, and for a company filing a report it mostly is. For a company selling software that files reports, removal is the expensive direction. Adding a datapoint is an insert. Removing one leaves you holding years of answers to a question the standard no longer asks — answers you cannot delete, because they were filed, and cannot show in a 2027 report, because the disclosure is gone. Every score whose denominator included those datapoints is now incomparable to every score you will calculate next year. The simplification is real. The migration it implies is not simple.

A regulation is not a schema

A regulation is a versioned dependency with no semantic versioning, no changelog you can diff, and customers on two versions at once.

The failure mode is easy to describe and nearly universal. A team builds the first version of a compliance product against the standard in force. Disclosure requirements become columns, the scoring model becomes code, the validation rules become a file of conditionals. It ships, and it works, and it keeps working for two or three years — which is exactly long enough for the shape to feel permanent. Then the standard is revised and the team discovers it modelled a moving thing as a fixed one. ALTER TABLE is now the migration path for an audit trail.

The alternative is not clever. It is the same move you already make for currencies, tax rates and feature flags: put the volatile part in rows and let the code read it. Three pieces do most of the work. Datapoints become registry rows keyed by version. Every stored answer records the version it was authored under. And the crosswalk between versions becomes a first-class table, because the interesting relations are not one-to-one.

That last point is where the design earns its keep. Most teams that do version their datapoints still model the mapping as a rename: old id, new id. Real standard revisions split one disclosure into three, merge three into one, narrow a definition until last year's answer is no longer valid, or withdraw a datapoint with no successor at all. A mapping table that can only express renames will quietly produce numbers nobody can defend in an audit.

Here is the shape, in about a hundred lines.

// esrs-registry.ts — modelling a regulation as data instead of as a schema.
//
// The mistake this avoids: `ALTER TABLE reports DROP COLUMN e1_6_gross_scope_1`.
// The revised ESRS removes 61% of mandatory datapoints. A dropped column is a
// dropped audit trail, and you still owe the regulator the years you filed.

/** A published version of the standard. Not semver — the standard picks these. */
type StandardVersion = 'esrs-2023' | 'esrs-2026'

/**
 * Datapoints are rows, not columns. Adding, removing or retyping one becomes an
 * INSERT against a registry table — a seed file, reviewable in a pull request —
 * rather than a migration against live customer data.
 */
interface Datapoint {
  version: StandardVersion
  id: string // e.g. 'E1-6_01'
  dataType: 'decimal' | 'integer' | 'boolean' | 'narrative' | 'enum'
  unit?: string
  required: 'always' | 'if-material' | 'withdrawn'
}

/**
 * The crosswalk is a first-class entity, because the interesting cases are not
 * 1:1. EFRAG shipped exactly this on 28 August: a mapping version of the 2026
 * datapoint list referencing the 2023 ESRS and the 2024 implementation guidance.
 */
type Relation =
  | 'identical' // same meaning, same shape — safe to carry forward
  | 'renamed' // same meaning, new id
  | 'narrowed' // new datapoint asks for less; old answer needs truncation
  | 'widened' // new datapoint asks for more; old answer is incomplete
  | 'split' // one old id -> several new ids; needs human allocation
  | 'merged' // several old ids -> one new id; needs an aggregation rule
  | 'withdrawn' // no successor. Keep the data. Stop reporting it.
  | 'new' // no predecessor. There is no historical value to show.

interface Crosswalk {
  from: { version: StandardVersion; id: string } | null
  to: { version: StandardVersion; id: string } | null
  relation: Relation
}

/**
 * Every stored answer records the version it was authored under. This is the
 * one field teams skip, and skipping it is what makes the migration
 * irreversible: without it you cannot tell a 2023 answer from a 2026 answer
 * once they share a table.
 */
interface Answer {
  entityId: string
  fiscalYear: number
  authoredUnder: StandardVersion
  datapointId: string
  value: string
}

type Resolved =
  | { status: 'ok'; value: string; note?: string }
  | { status: 'needs-review'; reason: string; candidates: string[] }
  | { status: 'unavailable'; reason: string }

/**
 * Rendering a report is a pure function of (version, registry, answers).
 * Note what it does NOT do: silently coerce. A compliance product that guesses
 * across a version boundary is worse than one that reports a gap, because the
 * gap is visible in review and the guess is not.
 */
export function resolve(
  target: { version: StandardVersion; id: string },
  answers: Answer[],
  crosswalks: Crosswalk[],
): Resolved {
  const direct = answers.find(
    (a) => a.authoredUnder === target.version && a.datapointId === target.id,
  )
  if (direct) return { status: 'ok', value: direct.value }

  const links = crosswalks.filter(
    (c) => c.to?.version === target.version && c.to.id === target.id,
  )
  if (links.length === 0) {
    return { status: 'unavailable', reason: `${target.id} has no predecessor` }
  }

  // A clean carry-forward is the only case we resolve automatically.
  const clean = links.find(
    (c) => c.relation === 'identical' || c.relation === 'renamed',
  )
  if (clean?.from) {
    const prior = answers.find(
      (a) =>
        a.authoredUnder === clean.from!.version &&
        a.datapointId === clean.from!.id,
    )
    if (prior) {
      return {
        status: 'ok',
        value: prior.value,
        note: `carried forward from ${clean.from.id} (${clean.relation})`,
      }
    }
  }

  // Everything else is a question for a human, and the report should say so
  // rather than emit a number that cannot be defended.
  return {
    status: 'needs-review',
    reason: links.map((l) => l.relation).join(', '),
    candidates: links.flatMap((l) => (l.from ? [l.from.id] : [])),
  }
}

/**
 * The comparability check nobody builds until an auditor asks for it.
 * A year-on-year trend across a version boundary is only honest when the
 * datapoint survived the boundary intact.
 */
export function comparable(
  a: { version: StandardVersion; id: string },
  b: { version: StandardVersion; id: string },
  crosswalks: Crosswalk[],
): boolean {
  if (a.version === b.version) return a.id === b.id
  return crosswalks.some(
    (c) =>
      c.from?.id === a.id &&
      c.to?.id === b.id &&
      (c.relation === 'identical' || c.relation === 'renamed'),
  )
}

The function that matters is resolve, and what matters about it is that it can return needs-review. A compliance product that guesses across a version boundary is worse than one that reports a gap, because a gap is visible in review and a guess is not. The comparable helper is the other half: a year-on-year trend that crosses a version boundary is only honest when the datapoint survived the boundary intact, and the crosswalk is the only thing that knows whether it did.

This is more machinery than a column, and it is worth saying what it costs. The joins are real, the queries are slower, and a new engineer takes a day longer to find where a number comes from. If you are already fighting table bloat in a shared-schema multi-tenant database, a registry join on every datapoint read is a cost to measure before you accept it. If you serve one jurisdiction, one customer segment, and a standard that has never been revised, the registry is over-engineering and you should skip it. I would only note that ESG reporting looked like that category in 2023.

One boundary worth stating plainly: this is about how to build the product, not about what any company must report. The reporting obligations belong to accountants and counsel, and nothing here substitutes for either.

What a founder or CTO does with this

The practical question is not whether to refactor. It is whether you can serve two versions of a standard in the same quarter, and you can answer it this week without touching code. Pick a datapoint that changed. Ask your team what it would take to show a customer their FY2025 answer next to their FY2027 answer on one screen. If the answer involves a migration, a backfill script, or the phrase "we'd have to check", you have a schema where you need data.

The cost asymmetry deserves to be blunt. Building the registry before you need it is roughly a sprint. Retrofitting it after a revision — against live audit data you are legally required to retain, with customers filing on a deadline — is a quarter, and it is the same quarter those customers spend asking for the new features the revision created. Waiting for the requirement is not the cheap option here, because the requirement arrives with a date attached.

It also generalises well past ESG. Anything that encodes an external standard sits in the same position: tax rules, FHIR profiles, PCI DSS, WCAG success criteria, SOC 2 control sets, the EU AI Act's transparency obligations. Every one has a version history and a body that revises it on a cadence you do not control. If the standard lives in your schema, that body's release calendar is now your migration calendar.

The four-part pattern

  1. Datapoints are rows, keyed by version. The registry holds id, data type, unit and whether the disclosure is mandatory. A new version of the standard is a new set of rows, not a new set of columns, so shipping it is a seed file in a pull request rather than a migration against production.
  2. Every answer records the version it was authored under. One column. It is the cheapest item on this list and the only one that is genuinely irreversible if you skip it, because once two versions share a table without it, nothing can tell them apart.
  3. The crosswalk is a table, and its relations are richer than rename. Identical, renamed, narrowed, widened, split, merged, withdrawn, new. The split and merged cases are the ones that need a human, and a model that cannot express them will fabricate continuity instead of flagging it.
  4. Validation, scoring and rendering are functions of (version, registry). Not branches on a version constant. If supporting a new version means editing conditionals in three services, the standard is still in your code — you have only moved where it lives.

My perspective

I co-founded Klimado, an ESG and CSRD compliance platform, and we shipped three production applications in nine months. We got a lot right under that timeline. This is the thing I would do differently.

We built the scoring engine against the standard as it stood, because the standard as it stood was the requirement in front of us and nine months is not long. The disclosure requirements went into the schema. It was the fastest path to a working product and I would probably take it again under the same constraints — but I would take it knowing the cost, which is not what I knew at the time. The Klimado case study still describes the problem as enterprises struggling to keep up with evolving compliance requirements. That framing was right about the customer and incomplete about us. The platform was subject to the same evolution, and a platform that models a moving requirement as a fixed one inherits the customer's problem instead of solving it.

What changed my mind was not this revision. It was watching the identical pattern in a different domain — multi-tenant Postgres, where the teams that survive their first serious schema change are the ones who put tenant-varying rules in data early. The lesson is domain-independent, and the only real question is which part of your problem is actually the volatile one. Compliance products get this backwards in a specific way. They treat the regulation as the stable part and the customer's data as the variable part. It is the other way around. Your customer's emissions are a number. The question you are permitted to ask them about it changes by delegated act.

Recommended action this quarter

Three things, in order. Download EFRAG's mapping version of the 2026 datapoint list and try to load it into your product — not to migrate, just to see whether the shape fits. If it does not, you have found your gap before your customers did, and you have until the FY2027 filings to close it. Second, add the authoredUnder field to your answers table now, even if you version nothing else this year; it is a single nullable column today and an archaeological dig later. Third, if you are inside the fatal-flaw window, use it. Feedback closes 23 October, and software vendors are the constituency most likely to notice a mapping that cannot actually be implemented, which is precisely what a fatal-flaw review is asking for.

If you are building against a standard that is about to move and want a second opinion on whether your data model can take it, that review is part of the fractional CTO and architecture work I take on.

Check your model before the standard moves

If your product encodes a regulation and you are not certain whether the next revision is a config change or a migration, establishing which one it is usually takes an afternoon. Book a time and we will work through it together.

#Compliance Software#Data Modeling#RegTech#ESRS#Schema Versioning