# The anatomy of a module

How a DODIL business module is built, so that the fifth one looks like the first and a customer
can change any of it later. The [lifecycle](/modules/lifecycle.md) says *what to do in what
order*; this says *what the thing you build looks like*.

Every rule here comes from a module that runs, and each one says where it was learned. Some rules
**post-date** the CRM/GL/ITSM packages, which were built before the pattern existed — those say so
too, because a reference that quietly contradicts its own rules is worse than one that admits it.
Where a rule is currently blocked by an open platform fault, that is stated rather than glossed.

---

## 1 · Three artifacts, no more

A module is three things. Anything else is drift.

| Artifact | Audience | Lives | Answers |
|---|---|---|---|
| **The spec** — `/modules/<id>/spec.yaml` | agents | this site | who uses it, what it must do, what it keeps, what runs on its own |
| **The page** — `/modules/<id>/` | people | this site | the same content, readable, in the customer's words |
| **The package** — `/code/<package-id>/` + its build guide | both | this site | the running reference: models, routes, jobs, UI. A module may ship several (`crm-core`, `crm-suite-app`); its spec names them |

**The platform layer is not part of any of them.** It lives once, in
[`/code/platform-starter`](/code/platform-starter): `db.py`, `auth.py`, `sa_token.py`,
`PLATFORM.md`, the app skeleton with its self-check, an idempotent `migrate.py`, the Dockerfile
and the deploy files. Copy that into a customer's repo and generate the business layer beside it.
It was extracted because the same four files, duplicated across two dozen packages, had already
drifted into three versions of `auth.py` and three of `db.py`, and a repo that copied the wrong
one inherited the difference silently.

The spec and the page are generated from **one source** (`modules/<id>/module.yaml`), because two
hand-written copies of the same truth diverge — ours did, repeatedly, before they shared a source.

A customer's own system adds a fourth, in *their* repo: **`ERP_<MODULE>.md`**, the design file
([template](/modules/template.md)). The reference spec is the starting vocabulary; their design
file is the decision record.

## 2 · One app, a router per module

The shape to generate (the shipped packages are older and do not all look like this yet):

```
customer-repo/
  ERP_CRM.md  ERP_INVOICING.md      the design files — one per module, read by every agent
  main.py                            one FastAPI: mounts each module's router under /api
  db.py  auth.py  sa_token.py        the platform layer — COPIED VERBATIM, never edited
  schema.sql  migrate.py             tables; migrate is idempotent and safe to re-run
  models.py                          SQLAlchemy models, natural keys
  modules/
    customers.py  deals.py  quotes.py  tasks.py  catalog.py  home.py   one router each
    common.py                        permissions, settings, money, ids, retry
  tests/test_stories.py              one test per story id, per role, per invariant
  web/                               Vite + React; screens named after the stories
  .dodil/deploy.yaml  .github/workflows/ci.yml
```

**Split into more than one app only for a stated reason:** a public surface beside a private one,
independent scaling, or a distinct trust boundary. Two apps on two hosts do **not** share a
gateway login session — the cookie is host-only — so splitting for tidiness costs the customer a
second sign-in.

**One bucket per customer, one sign-in pool per customer.** That is what makes a sales opportunity
JOIN a finance invoice with no ETL, which is the whole reason the system is on DataK3.

## 3 · The data layer

### Naming
- **Module-owned tables carry the module prefix:** `crm_quotes`, `gl_journal_lines`. In a shared
  bucket an unprefixed `quotes` table is a collision waiting for the second module.
- **Shared master data is unprefixed and has exactly one writer:** `business_partner`, `product`.

### Keys
- **Natural primary keys, always.** DataK3 has no sequences, and `RETURNING` does not return a
  staged write. If you cannot name the key in the design phase, the entity is not understood yet.
- **Derive an id when one is needed:** `bp_id = int.from_bytes(blake2b(bp_key, digest_size=7))` (a
  56-bit BIGINT), `line_id = f"{quote_id}|{line_no}"`,
  `task_id = f"followup:{quote_id}"`. A derived id makes a re-import or a double-clicked button
  idempotent for free.
- **A counter is a row, not a sequence** — a high-water mark, with the primary key as the
  uniqueness guard and a retry on conflict. **Blocked today:** concurrent inserts of the same key
  are all acknowledged and only one row survives, so four quotes created at once can take the same
  number. Until that fault is fixed, treat strictly sequential numbering as unavailable under
  concurrency and say so to the customer.

### Writes
- **Every write is idempotent**, by one of two means: a **derived id plus a status guard** (the
  common case — re-marking a sent quote as sent returns early, and the follow-up task keyed
  `followup:<quote_id>` cannot double), or `INSERT … ON CONFLICT (pk) DO UPDATE SET col =
  EXCLUDED.col` / `data_table_upsert` where a row may legitimately be rewritten. `DO UPDATE SET`
  accepts only `EXCLUDED.<col>` — put constants in `VALUES`.
- **`DELETE` requires a `WHERE`** (`WHERE TRUE` for a deliberate clear).
- **No read-your-writes inside an open transaction.** Commit, then re-derive in a second
  transaction — and re-derive with `SUM(...)`, never `+=`, so concurrent re-runs land once.
  (A second, open fault goes further today: rows committed seconds earlier are *intermittently*
  invisible to a later request. Where a decision depends on a row you just wrote, re-read it and
  fail loudly rather than acting on an empty result.)
- **Money is `Numeric(18,2)`,** never float. Never add across currencies: totals are per currency.
- **`at` is reserved** — name timestamp columns `event_at`, `created_at`, `sent_at`.

### Derive, don't store
State that a date already decides is computed on read: a quote is *expired* because
`valid_until < today`, not because a nightly job set a flag. Every derived state is one less job,
one less way to be wrong, and one less thing to migrate. Store what a person decided; derive what
follows from it.

## 4 · Identity and visibility

- **No auth code in the app.** The gateway does PKCE and injects `X-Dodil-User`; `auth.py` is a
  header-trust role gate. `db.py` and `sa_token.py` are copied byte-for-byte; `auth.py` is copied
  and then extended with this customer's visibility helpers, which is the one place the platform
  layer legitimately grows. Forged headers are stripped at the edge — verify that once, per deploy.
- **Permissions are named `<namespace>:<object>:<verb>`** (`crm:quote:approve`, `bp:partner:write`)
  — the namespace is the module's declared short form, which may be shorter than its id. The CRM
  package predates this and still ships `quotes:approve`; generated systems use the namespaced
  form. A module joining an existing system **merges** its roles into the pool catalog:
  `roles get` → merge → `roles set`.
  `roles set` replaces the whole catalog, so writing only your own roles silently strips every
  other module's access.
- **Roles map to permissions in one place** (`common.py`), never scattered through routes.
- **Visibility is a decision, and it is recorded.** "Everyone sees everything" is a legitimate
  answer for a five-person team and a wrong one for a sales floor — ask, and write the answer in
  the design file. Where rows *are* owned, scope every owned read by the caller's identity and
  return *not found* rather than *forbidden*: 403 confirms the record exists. The packages ship the
  helper (`owner_scope`) but not the wiring, so this is a build step, not a freebie.
- **The UI hides what the role cannot do,** and the route refuses it anyway. The screen is
  convenience; the route is the control.

## 5 · Routes

- **CRUD plus workflow operations**, named for what the business does: `POST /quotes/{id}/send`,
  `POST /partners/promote`, not `PATCH /quotes/{id}` with a magic status field.
- **The module that owns a table is its only writer, across module boundaries.** Another module
  calls the owner's route; it never writes into the owner's tables. (Inside one module's own app,
  a sibling router writing a task row is fine — the rule is about the seams.) On DataK3 a foreign
  key is accepted at DDL and never enforced, so the sole-writer path *is* the integrity mechanism —
  and note it only holds for callers coming through the app: any principal with `k3.editor` can
  write the table directly.
- **A refusal explains itself in the customer's words** — "Q-1042 is sent and can no longer be
  edited — use Revise to make a new revision". The refusal text is part of the story's acceptance
  criteria, not an afterthought.
- **Guards live on the write path,** because column constraints are limited. A closed period or a
  frozen quote is enforced by the route, which is exactly why every write must go through it.
- **Every generated app ships `POST /` as a self-check** that opens the bucket with its own
  credential and reports row counts. `ignite invoke` calls it; a public-URL 401 proves only that the
  gateway is up. The older packages expose only `GET /healthz` — add the self-check as you generate.
- **`GET /healthz` does no database round trip** — the pod must report healthy before credentials
  resolve.

## 6 · Workflows are data, not code

States and transitions belong in the design file; the *tuning* belongs in rows:

| In code | In rows (the customer changes it without us) |
|---|---|
| the transitions that exist, and who may make them | stage names, order, probability, forecast category |
| that a discount cap exists and is enforced on every edit **and** on send | the cap: 10%, 20% for VIP, per customer |
| that a quote freezes when sent | validity days, follow-up days, expiry warning days |
| that a lost deal needs a reason | the list of reasons |

The test: *can the customer's office manager change it on a Settings screen?* If yes it is a row.
If changing it would change what the tests assert, it is code. This is what makes one reference
module fit a machine dealer, a consultancy and a distributor without a fork.

## 7 · Jobs — there is no scheduler

Ignite is request-invoked. A system story ("every 6 hours…", "every night at 02:00…") must say
**where it runs**:

1. **Derive it instead** (§3) — the cheapest job is the one that does not exist.
2. **In the same request** that caused it — a follow-up task created when the quote is sent, keyed
   `followup:<quote_id>` so a double click makes one task.
3. **An always-on worker** pinned warm (`--reserved 1 --max-replicas 1`) running its own loop.
4. **An external clock** calling a route.

Every job is idempotent: "running it twice changes nothing" belongs in the story's acceptance
criteria, not in a comment. A job that can fail half-way records what it processed, so the retry
knows where to resume.

## 8 · Screens come from stories

Each person story names an actor and a screen; the screens are the list of those, not an invention
of the build phase. Each screen names the stories it serves — usually several — and the design file
carries that table, so a change request can name the screens it touches before any code is written.

What the screens must do is fixed; how they are rendered is a judgement:

- **The default is Vite + React + TypeScript**, talking only to the routes and never to DataK3,
  with a client generated from the OpenAPI schema rather than hand-written fetch calls. Reach for
  it when there are many screens, live updating, or a team using it all day.
- **Server-rendered templates are the right answer for a small internal tool** — two users, a
  handful of screens, no offline behaviour. One less build, one less thing to deploy. State the
  choice in the design file; both are the pattern, and a needless React app is its own cost.
- **Either way the route is the control.** The screen hides what the role cannot do; the route
  refuses it regardless.
- **The rule is visible where it bites:** the discount bar reads "8.00% of 10%", the banner says a
  machine quote needs a site visit first. A rule the user cannot see is a rule they will fight.

## 9 · Tests are the stories

Three suites, all against a **dev bucket**:

- **Acceptance** — one test per story id, named for it (`test_s06_send_freezes_sets_validity_…`),
  asserting its Given/When/Then. A story with no test is a story nobody verified.
- **Access** — per role: sees the right rows, others' rows are *not found*, excluded actions
  refused, no gateway identity is rejected.
- **Invariants** — the trial balance nets to zero, a re-import changes nothing, a frozen record
  refuses edits, totals are computed not typed.

Coverage is checkable because both sides carry the story id. A customer CRM built this way carried
31 tests at first release and 36 after three change requests; in every run its failures were the
same three platform faults rather than app faults — which is only knowable *because* each test
names the requirement it verifies.

## 10 · Scale and failure

- **The tables reader serves 16 concurrent reads.** Above that: `ConfigurationLimitExceeded`. Hold
  a client-side semaphore below it (`DB_CONCURRENCY=10`) and retry transient failures. That pair
  took a 60-way write storm from lossy to clean. No shipped package implements the semaphore yet —
  add it when the workload is wide.
- **Retry the transient, fail the deliberate.** Match on SQLSTATE (`40001` write conflict, `08006`,
  `58030`, `XX000`) and, because object-store blips are sometimes misclassified as syntax errors,
  on the message marker too.
- **One pinned-warm replica** for anything holding a loop or a counter.
- **`data table compact` is a performance step, never a correctness step.** There is no freshness
  flag.
- **When the platform misbehaves, reproduce it, write a brief with a script that exits non-zero
  while the fault is live, and stop.** Do not turn a workaround into a design rule, and do not
  document the bug as a convention.

## 11 · Change, migration, versioning

- **A change request is a story change first** — the design file moves before the code, and the CR
  records who asked, who approved, the impact with real counts, and how to undo it.
- **`ALTER TABLE … ADD COLUMN` is the only cheap schema change** — there is no migration tool and
  no views — **but it currently destabilises reads on that table for minutes afterwards** (an open
  fault: the new column is intermittently invisible, and files written before it fail a schema
  match). Add columns last in the DDL, migrate the test copy first, and expect a noisy few minutes.
- **A rename or type change is: add → backfill → switch reads → drop later.** Nothing is dropped in
  the release that stops using it.
- **`ignite version rollback` returns the app to an earlier deployed version; the data is not
  versioned with it** — which is why the go-live step exports the affected tables first.
- **Dev bucket, then live, and live only on an explicit "go live".**

## 12 · Anti-patterns

| Smell | Why it hurts | Instead |
|---|---|---|
| A CRM-only `customers` table when invoicing is on the roadmap | two customer lists, reconciled by hand forever | reference the shared partner from module one |
| `SERIAL` / autoincrement | no sequences here; the ORM's `RETURNING` path returns nothing | natural or derived keys |
| A nightly job to set a status a date already implies | a job, a race, and a wrong answer between runs | derive it on read |
| Writing another module's table "just this once" | the sole-writer rule is the only integrity there is | call the owner's route |
| `roles set` with only your module's roles | silently strips every other module's access | get → merge → set |
| Hard-coded stage names, thresholds, terms | every customer forks the code | rows the customer can edit |
| 403 for a record the caller may not see | confirms the record exists | 404 |
| Tests named after functions | nobody can tell which requirement is unverified | name them after story ids |
| Verifying a deploy by loading the public URL | proves the gateway is alive, not the app | `POST /` self-check + pod logs |

## 13 · Composing modules into one system

One module is a project; four modules is an ERP. What makes the fourth one cheap:

**Install in master-data order.** A module that *owns* a shared record comes before the modules
that *reference* it. For a typical business:

```
business-partner   the customer/supplier record          (owns it)
   └── crm         sells to them                         (references)
        └── ar     invoices them                         (references; CRM hands it the accepted quote)
             └── gl posts the invoice to the books        (owns the ledger; AR calls it)
```

Getting this order wrong is the one mistake that cannot be fixed cheaply later: two customer lists
are a migration, not a refactor.

**One app, one router each.** Adding a module to a running system is a new router, a new set of
prefixed tables, and its roles merged into the existing pool catalog. Not a new app, not a new
bucket, not a second sign-in — unless one of the split reasons in §2 applies.

**Cross-module reads are joins; cross-module writes are calls.** The point of one bucket is that
the collections list can join invoices to the partner and to the sales owner in a single query, with
no ETL. The moment a module *writes* another's tables, that guarantee is gone, because the owner's
guards no longer run.

**Verified, 2026-09-18 — four modules in one bucket.** `business-partner`, `crm`, `ar` and `gl`
installed in the order above into one DataK3 bucket: **71 tables, no name collisions**, and a
fifth package (`itsm`) adds 23 more with none either. A deal accepted in the CRM promotes the
party, raises the invoice and posts the journal — one query joins all four
(`code/crm-suite-app/v1/tests/test_suite.py`, 10 tests). Two things the exercise found, both
worth knowing before you repeat it:

- **The join needed a `CAST`.** `business_partner.bp_id` is a BIGINT (a graph walk needs an
  integer start node) and `ar_invoices.bp_id` is a VARCHAR. The un-cast join does not return
  wrong rows — it **fails outright**. Agree the type of a shared key across modules *before*
  the second one stores it, because afterwards it is a data migration.
- **The prefix rule above is followed by 21 of 68 module-owned tables.** `quotes`, `products`,
  `tasks`, `ledger`, `periods` and `journals` all sit unprefixed, which by this document's own
  convention claims they are shared master data. Nothing collides today; the convention that
  would keep it that way is not actually being applied.

**Shared settings stay where they are owned.** Tax codes belong to the ledger once the ledger
exists; payment terms belong to invoicing; a customer's currency belongs to the partner record.
Two modules writing the same setting is the same bug as two customer lists, one level down.

**Test the seams, not just the modules.** A suite gets one test per integration story that runs both
sides — an accepted quote produces exactly one draft invoice, and posting it twice changes the
ledger once. Those are the tests that fail when someone "just adds a column" months later.

**Each module keeps its own design file.** `ERP_CRM.md`, `ERP_INVOICING.md`, `ERP_FINANCE.md`. A new
module reads them all before designing anything, which is how it learns what a customer is, who
writes it, and which stories already cross into its territory.

## 14 · Maturity — what a module claims about itself

A module's `maturity` is a promise to whoever picks it up, so it has criteria:

| Level | Means | Earned by |
|---|---|---|
| **planned** | a spec exists; no code | the spec passes `check:modules` and names its gaps |
| **reference** | code exists and runs against a bucket | the package installs, its schema applies, and its routes answer |
| **tested** | every story has a test, and they pass | the acceptance, access and invariant suites run green on a dev bucket |
| **deployed** | it has run behind the gateway, for a real person | the self-check passes through `ignite invoke`, forged identity headers are stripped, and someone signed in per role |

Two rules keep the ladder honest: **a module never claims a level it has not been through**, and
**known gaps are published with it** (`reference.known_gaps`), because a reference that hides what
it does not do costs the next person a day.

A new module starts from the skeleton at
[`/modules/module.template.yaml`](/modules/module.template.yaml), not from a blank file — the
sections are the pattern in question form.

## 15 · The shape of a good module, in one paragraph

One bucket, one sign-in, one app with a router per module. Tables prefixed unless they are shared
master data with a single writer. Natural keys, idempotent writes, money as exact decimals, derived
state over stored state. Every rule enforced on a write path the UI merely mirrors. Every scheduled
thing either derived away or given a named home. Every story carrying its own test, and every
decision carrying its date and its author in `ERP_<MODULE>.md`. Everything a customer is likely to
argue about — stages, thresholds, terms, reasons — living in rows they can edit, not in code we
have to fork.
