How JMAD is built
Four stages turn public source data into the record you read at an asset URL: gather, save, consolidate, publish. Each stage only appends. Nothing earlier in the chain gets rewritten, which is what lets every fact trace back to where it came from.
Gather
Work is split into tiles by 町字 (machi-aza, Japan's smallest administrative unit below a ward), falling back to a JIS X 0410 mesh grid where a tile has no address boundary yet. A gather job runs one source loader against one tile: OpenStreetMap, PLATEAU (LOD1 CityGML, range-read per mesh tile out of the per-ward CityGML ZIP rather than downloading the whole archive), the Address Base Registry, e-Stat (census small-area boundaries, used to resolve real prefecture/ward/machi-aza), and Wikidata are running. J-REIT and EDINET disclosures, the Tokyo green-building registers and the 管理計画認定 register (mankan) add facts to buildings they match exactly. Foursquare and 登記 (touki, the real estate registry) are not implemented. The loader's raw response is written to storage keyed by the hash of its own bytes, so re-fetching an unchanged source writes nothing new. Every fact the loader extracts becomes one row in the observation log.
Save
Observations are append-only. Nothing is ever updated or deleted here; if a source corrects itself, that's a new row with its own retrieval timestamp, sitting next to the old one. Every row carries its source, source URL, retrieval time, extraction method, and a confidence score. There's no path that skips this. A fact without provenance doesn't get written.
Consolidate
A consolidate job takes a tile's observations and clusters them into buildings, using footprint overlap and normalized name matching. One building gets one DNK Asset ID, allocated once and never reused, whether the underlying observations came from three sources or ten. The ID is tied to the cluster's source records rather than to its geometry, so re-gathering a building that has moved or been renamed reuses its ID. When two clusters turn out to be the same building later, they merge: the older ID survives, and the retired one redirects to it permanently. When one cluster turns out to be two buildings, it splits: one keeps the ID and the other is allocated a fresh one, with no redirect between them. Consolidation is queued when no gather jobs remain pending or leased for the same tile and revision. A final gather completion or dead-letter transition commits together with that follow-up job, so an enqueue failure rolls back the transition and leaves it retryable. Concurrent final transitions create one consolidation job.
Dead-lettered sources count as finished for this decision. Consolidation uses the observations available from the run, even when a source exhausted its retries. This permits partial-source results; a completed consolidation does not mean every source succeeded. A new revision can be enqueued for a later run.
For each field, a precedence table picks a winner by source (see Provenance and precedence for the table itself). The values that lost stay attached to the record instead of being thrown away, which is what shows up as conflict on an asset page. A building's location (prefecture, ward, machi-aza) is resolved the same way, from e-Stat and Address Base Registry area observations — see areas and coverage. JMAD marks values it computes rather than reads directly, like GFA from footprint times floor count or a centroid from a footprint polygon, as estimated. Records are bitemporal: they track both when something was true in the world and when JMAD recorded it, so a correction closes the old row and opens a new one rather than overwriting it.
Publish
For each building, PostgreSQL consolidation commits canonical building changes, venue changes and the publish job in one transaction. If enqueueing the job fails, those canonical changes roll back. Concurrent consolidation decisions for the same building are serialized. Every consolidated building gets a publication intent keyed by the building and its latest recorded time across the building, its venues and its photos, so venue-only and photo-only changes publish too, and a retry for the same recorded version reuses that key.
The publish worker then regenerates the read model, markdown twin and change-feed entry asynchronously. A committed canonical update can therefore wait in the queue before readers see it; it cannot commit without its publication intent through this consolidation path. This is a per-building boundary, not a transaction for the entire tile.
flowchart LR Gather[Gather: tile + source loader] --> Save[(Observations, append-only)] Save --> Consolidate[Consolidate: cluster, ID, precedence] Consolidate --> Publish[Publish: page, markdown, change feed]
Database access by workload
Web, worker and migration tasks use separate database credentials (jmad_web, jmad_worker, jmad_migrator). The web role reads published tables such as pages, redirects and the change feed, and can write only rate-limit counters. It cannot read or write canonical building records, observations or jobs, and cannot create tables. The worker role reads and writes application tables but not the migration ledger. A dedicated migration task uses jmad_migrator, which owns the application tables and can alter their schema.
On a fresh database, an operator runs the bootstrap script once, before the first migration, to create the three login roles. Per-table privileges are applied by migrations, so they arrive with the tables they cover.
Persistent storage and startup
With NODE_ENV=production, the web server requires a non-blank JMAD_DATABASE_URL. The worker always requires JMAD_DATABASE_URL, and in production it also requires JMAD_PAYLOAD_BUCKET (or JMAD_PAYLOAD_DIRECTORY) for raw source payloads and JMAD_PHOTO_BUCKET (or JMAD_PHOTO_DIRECTORY) for building photos. Missing configuration stops startup before the web server listens or the worker processes jobs. Errors identify the missing variable names without printing connection strings.
An empty configured database is valid and can return no assets. It is different from a startup failure: if a service cannot start, check its configuration error and supply the required deployment variables. Do not treat an empty asset count alone as a readiness failure. Outside production, the web server uses empty in-memory stores when JMAD_DATABASE_URL is unset; the worker has no in-memory database mode.
Optional reverse enrichment
Nominatim-compatible reverse enrichment is disabled by default. Operators must explicitly enable it and configure a provider endpoint, identity, version and supported rate budget. The public nominatim.openstreetmap.org endpoint is rejected for this systematic workload. An endpoint setting is not evidence of permission; the operator must establish that the provider permits the intended use.
The only supported budget is 60 requests per minute. JMAD paces reverse requests at one per second through a shared Postgres token bucket, so every worker draws from the same budget. Reverse responses are not cached yet (JMA-61). Reverse enrichment reads whichever footprints are already stored when it runs and does not wait for footprint sources in the same run to finish (JMA-60), so an early run can produce few or no reverse results.
Disabled enrichment skips reverse requests and leaves other gather sources usable. Enabled responses retain the actual provider URL in provenance, with credentials and endpoint query secrets removed. Reverse results describe the nearest suitable source object and are fallback address evidence, not proof of an exact building address.
Why this matters if you're reading the data
You can always ask why a value is what it is. Every field points at a specific source record, not "aggregated from public sources." If two sources disagree, you see both, instead of a silent pick you have to trust. And history is never lost: a building's floor count from an earlier revision is still there if you need it, next to whatever's current now.
Next: DNK asset IDs for how the identifier itself works, or Provenance and precedence for the field-by-field precedence table.
Updated 1 day ago
