Agent discovery
JMAD lives at dnk.co/japan; dnk.co itself is DNK's marketing site, a separate deployment that shares nothing with JMAD but the hostname. That decides where discovery files can live: everything JMAD publishes sits under /japan, and the two documents whose location is fixed at the host root by their specifications, robots.txt and anything under /.well-known/, are the marketing site's to publish. Everything here follows conventions already in use elsewhere; nothing is a bespoke JMAD format.
llms.txt
/japan/llms.txt describes the dataset in under roughly 4,000 tokens: what JMAD is, a short "how agents should use JMAD" block (the API is keyless, cite the source URL and observed date on every fact, join on DNK Asset IDs rather than addresses or names, search the API instead of crawling), links to the OpenAPI document and the MCP server, the ID scheme, the provenance model, and a list of coverage hubs by prefecture. Every JMAD page advertises it in its head and its Link header with rel="llms-txt", so an agent that lands on any page finds it in one hop; there is no /llms.txt or /.well-known/llms.txt at the host root.
It deliberately holds no links to individual buildings. With millions of buildings once coverage grows, a flat list of asset links would defeat the file's purpose, which is to orient an agent before it starts making requests, not to serve as an index. Use search and pagination or the MCP search_assets tool for that.
Sitemaps
/japan/sitemap.xml is an index listing one shard per ward, /japan/sitemaps/<prefecture>-<ward>.xml, each capped at 50,000 URLs (a ward that exceeds it gets -2, -3, and so on), plus /japan/sitemaps/recent.xml holding the 1,000 most recently published pages. Every URL's lastmod is that page's actual record time, not a build timestamp, so a crawler that respects lastmod can skip pages that have not changed. The sitemap protocol lets an index list any URL below its own path, so an index under /japan covers every JMAD page; dnk.co/robots.txt points crawlers at it.
robots.txt
robots.txt is only read from the host root, so JMAD does not serve one: https://dnk.co/robots.txt is the marketing site's file, and it carries the Sitemap: line for /japan/sitemap.xml. JMAD's own position on crawling is expressed where it can be: every page and markdown twin is served to every crawler without a key, and the crawler user agents listed in markdown twins and negotiation get markdown by default.
What is not published, and why
JMAD does not publish /.well-known/api-catalog, /.well-known/mcp/server-card.json, /agents.txt, /agents.json, ai.txt, an A2A agent card, ai-plugin.json, /.well-known/mcp.json, or any OAuth discovery metadata. The /.well-known/ paths are root-only by definition and the root belongs to the marketing site; the OpenAPI document at /japan/api/v1/openapi.json, the server card at /japan/mcp/server-card and /japan/llms.txt carry the same links. The rest are competing drafts for the same job llms.txt and the MCP server card already do, or describe an auth flow JMAD does not have: the API and MCP server are keyless, so there is nothing for an OAuth discovery document to point at. Publishing an empty or unused discovery file is worse than omitting it, since a client that finds it has reason to expect it does something.
Next
- MCP server for the server card these files link to.
- Markdown twins and negotiation for how a crawler reaches a page's markdown once it has found it.
Updated 21 days ago
