Skip to main content

Provenance

Sources, licences & attributions

Every endpoint badged real has to name where its facts come from — the registry refuses to register one that doesn't. This page is generated from those declarations, so it lists exactly what the catalog actually serves.

Endpoints badged mock generate synthetic data and appear nowhere on this page: there is no upstream to credit, which is the whole point of the distinction.

Data we redistribute

These are the sources whose data is copied into the Worker bundle and therefore served onward to you. Standards we merely implement — an RFC algorithm, an ISO check-digit procedure — carry no redistribution obligation and are cited on the individual endpoint pages instead.

Source What we redistribute Licence Obligation
IANA protocol registries HTTP status codes, HTTP methods, HTTP field names, DNS RR types, media types, the root zone TLD list, IPv4/IPv6 special-purpose address registries and the complete service name & port registry for ports 0–49151 (every individual registration plus every port-range row). Each bundled registry is pinned to a dated snapshot by SHA-256 and re-verified by a committed parity script. CC0 1.0 (public domain dedication) None. Attribution given as a courtesy and because it is the honest provenance.
mime-db 1.54.0 (jshttp) File-extension and compressibility mappings for media types, plus the de-facto types real servers send that IANA never registered. Merged into the bundled MIME dataset alongside the IANA registry. MIT — Copyright (c) 2014 Jonathan Ong; (c) 2015-2022 Douglas Christopher Wilson The copyright and permission notice must travel with the data. It is reproduced in the bundled dataset, credited here and on the /apis/mime-types page. The pinned version and db.json SHA-256 are recorded in the committed refresh script.
GeoNames A pinned snapshot of the cities500 gazetteer dump: place names and ASCII aliases, ISO country codes, first-order division names, WGS84 coordinates, population figures, IANA timezone identifiers and the PPLC capital classification. From the countryInfo.txt dump (snapshot 2026-07-30, SHA-256 93bafc52…) two more columns: the neighbours column plus the fips column that bridges GEC file codes to ISO 3166-1, which is the land-adjacency graph behind /apis/country-borders, and the Postal Code Format column, which anchors the hand-adjudicated postal-code mask table behind /apis/postal-code-formats. The upstream Postal Code Regex column is deliberately NOT redistributed — 31 of its rows contradict their own mask and several carry another country's data — so every published regex is derived from the mask here instead. CC BY 4.0 Attribution to GeoNames is required and is given here and on the endpoint page. GeoNames states the data is provided "as is" without warranty or any representation of accuracy, timeliness or completeness; population figures are approximate administrative counts, not census values.
Unicode CLDR / ICU, via the runtime's Intl implementation Nothing is copied from the runtime's own copy: timezone, plural-rule, collation and locale-formatting answers are computed by the serving runtime's Intl data at request time. (A pinned CLDR release IS copied for the ISO 3166-2 subdivision tables — see the Unicode CLDR release-48-2 row below.) Unicode License v3 (as shipped inside the runtime) No redistribution by us. Because the data is the runtime's rather than a table we pin, affected responses echo the resolved options so an answer stays auditable.
IANA Time Zone Database (tzdb) Nothing is copied — zone identifiers, offsets and transitions come from the runtime's own tzdb copy via Intl. Release 2026c's zone.tab and backward files are additionally read at snapshot-refresh time to validate each airport's derived time zone against its ISO country and spell it with that country's canonical zone id. Public domain None.
OurAirports The complete dated set of scheduled-service airports that carry an IATA code (4,159 rows, 235 countries and territories): IATA/ICAO/ident, name, served municipality, country, coordinates and elevation. The upstream airports.csv and countries.csv are pinned by SHA-256 and regenerated by a committed refresh script. Public domain None. The snapshot date is stated on the endpoint page.
tz-lookup 6.1.25 Nothing is copied. Its coordinate→timezone lookup runs at snapshot-refresh time only, to derive each airport's IANA zone id from its pinned coordinates; the Worker ships the resulting zone ids, not the lookup's data. CC0 1.0 None.
ISO 4217 currency register Currency codes, names, symbols and minor-unit digits. Factual reference data; the register itself is published by SIX on behalf of ISO. We redistribute the code list and minor-unit digits as facts, not the ISO standard document. Amount arithmetic follows the published minor-unit digits exactly.
SI Brochure (BIPM) and NIST Special Publication 811 Nothing is copied. The unit converter's factors are the defining numeric values of published standards — the 1959 international yard & pound agreement, gₙ = 9.80665 m/s², 1 atm = 101325 Pa, the IT and thermochemical calorie and Btu, the 2019 SI electronvolt — restated in our own words and re-derived from first principles by a committed refresh script. NIST SP 811 is a U.S. Government work (public domain in the US); the SI Brochure is published by the BIPM. Defining values of standards are facts, not expression. None beyond honest attribution, which is given on the endpoint page and here.
National public-holiday rules (61 countries) Nothing is copied. The rules are hand-authored from each country's own statute, gazette or ministry page and the dates are computed at request time; no upstream holiday dataset is redistributed. Statutory facts. Each country's official source is cited in the API response of GET /api/holidays/countries. Cite the per-country official source, which the API does on every listing. The rule set carries a verification date and a committed parity script that re-derives six years for every country.
WHATWG Encoding Standard index tables All 27 legacy single-byte index tables published by the Encoding Standard (every windows-125x, every ISO-8859 part, KOI8-R/U, IBM866, windows-874, macintosh and x-mac-cyrillic), used to diagnose and repair mojibake. Snapshotted from the index files dated 2024-09-18 and pinned per file by SHA-256. CC BY 4.0 (Encoding Standard); the index tables are factual byte-to-code-point mappings. Attribution to the WHATWG Encoding Standard, given here and on the endpoint page.
Unicode Character Database and UTS #39 security data (Unicode 17.0.0) The UTS #39 confusables mapping (6,565 code points onto their prototypes) and the 2,245 skeleton groups derived from it, the 77 intentional-confusable pairs, Identifier_Status and Identifier_Type, plus the UCD data those checks need: Script and Script_Extensions ranges, Default_Ignorable_Code_Point ranges, decimal-digit zeros and 7,559 character names. The Bidi_Class ranges behind /apis/punycode and the character names behind /apis/stress-strings come from the same database and are covered by this row. Eleven upstream files, each pinned by SHA-256 to a dated 17.0.0 snapshot and regenerated by a committed parity script. Unicode License v3 — Copyright © 1991-2026 Unicode, Inc. Terms of use: https://www.unicode.org/copyright.html The copyright and permission notice must appear with all copies of the data files or in the accompanying documentation. It is reproduced verbatim in every bundled dataset header, credited here and on the endpoint pages. The licence also forbids using the copyright holder's name to promote the product, so Unicode is cited as a source only, never as an endorser. Unicode and the Unicode Logo are registered trademarks of Unicode, Inc.
Unicode CLDR (release-48-2) The ISO 3166-2 subdivision code list: 5,046 current subdivision codes across 200 countries with their English names and CLDR's own "NO SUBDIVISIONS" statements, the current containment, the 599 deprecated codes and 27 overlong aliases from the subdivisionAlias tables, and 27 English territory names. All five pinned files are recorded with byte length and SHA-256 and re-verified by a committed refresh script. Unicode License v3 (SPDX: Unicode-3.0) — Copyright © 1991-2026 Unicode, Inc. The copyright and permission notice must travel with the data. It is reproduced verbatim in the header of api/src/data/country-subdivisions.ts, credited here and on the /apis/country-subdivisions page; Unicode is cited as a source only, never as an endorser. CLDR tracks ISO's newsletters with a lag, so a freshly renamed subdivision can be stale — the pinned release and its date are stated on the endpoint page.
Google libaddressinput address metadata The per-country postal address layout rules transformed from RegionDataConstants.java (252 region rows plus the ZZ fallback): the line template and its Latin-script variant, the required-field and upper-case letters, the admin-area/locality/sublocality/postal-code type tokens, the postal-code display prefix, field-width overrides and language hints — plus the English field labels resolved from the project's own address_strings.xml and the rendering algorithm transcribed from cpp/src/address_formatter.cc. The same file supplies the four postal-code facts behind /apis/postal-code-formats: whether a country's address format contains a postal code field, whether it is required, the term that country uses for it and the international display prefix. Eight upstream files are pinned by commit and SHA-256 and re-verified by a committed parity script. Google's Address Data Service (postal-code patterns, post-office URLs, sub-region lists) offers no pinnable snapshot and is deliberately excluded. Apache License 2.0 (source) and CC BY 4.0 (data), per the project README — Copyright (C) 2010, 2013, 2014 Google Inc. Both are attribution licences and both require a statement of changes. The copyright notices, links to both licences and an explicit statement that this is a transformed work — the Java source parsed, its per-region data decoded, keys renamed, fallbacks applied and label ids resolved against a second file — are reproduced in the headers of api/src/data/address-formats.ts and api/src/data/postal-code-formats.ts, credited here and on both endpoint pages. No upstream file is redistributed verbatim. The repository ships no NOTICE file at the pinned commit, so Apache-2.0 §4(d) adds nothing; the parity script re-checks that on every run.
Google libphonenumber (release v9.0.35, 2026-07-17) The numbering-plan metadata from resources/PhoneNumberMetadata.xml: 254 territories on 215 country calling codes — national numbering patterns and possible lengths per number type, national and international dialling prefixes, national-prefix parsing and transform rules, formatting rules and the upstream example numbers. Transformed into packed TypeScript rows; the alternate-format, carrier, geocoding, timezone and short-number datasets are not included. The upstream file is pinned by release tag and SHA-256 and regenerated by a committed parity script. Apache License 2.0 — Copyright (C) 2009 The Libphonenumber Authors The licence notice travels with the data: it is reproduced verbatim in each generated dataset header together with a statement of the changes we made, and credited here and on the /apis/phone-numbers page. The upstream repository ships no NOTICE file at v9.0.35, so §4(d) adds nothing further. The libphonenumber and Google names are used only to identify the origin of the data.
crawler-user-agents (monperrus) A pinned snapshot of 1,498 crawler user-agent regular expressions with the project's own 12-value tag classification, the date each pattern was added and the reference link it carries, plus — for 118 selected rows only — one example User-Agent string. Pinned by commit SHA and JSON SHA-256 and regenerated byte-for-byte by a committed refresh script. MIT — Copyright (c) 2017 Martin Monperrus The copyright and permission notice must travel with the data. It is reproduced verbatim in api/src/data/crawler-agents.ts, credited here, and cited as the primary source on the /apis/crawlers page.
Crawler operator documentation (OpenAI, Anthropic, Perplexity, Google, Apple, Common Crawl, Meta, Amazon) Nothing is copied. Each crawler's robots.txt token, documented purpose, stated robots.txt behaviour, published verification method and published User-Agent string are read from the operator's own page and restated in our own words. The table carries the date it was last checked. Published facts about each operator's own crawlers; the token names are factual identifiers. Cite the operator's own page per token, which the API does on every row.
CIA World Factbook, via the factbook.json mirror Per-pair land-boundary lengths in kilometres and coastline lengths, read from Geography → Land boundaries and Geography → Coastline across the 260 region files of the mirror, pinned to one commit by a content digest over those files. Nothing else is copied. The publication was discontinued in February 2026, so the figures are frozen at its final edition and will never be updated. CC0 1.0 (public domain dedication) for the factbook.json datasets; the underlying World Factbook is a US-government work. None. Credit is given because it is the honest provenance. The lengths are the publication's own figures, not measurements we made; cia.gov is not cited as a source URL because it no longer serves the data.
World Magnetic Model 2025 (NOAA NCEI / British Geological Survey) The model's 90 Gauss coefficient rows — 168 main-field and 168 secular-variation coefficients to degree and order 12 — transcribed by script from the pinned WMM.COF, NOAA's published per-element error-model values, and NOAA's own 112 published test values, committed as the endpoint's regression fixture. US Government work — public domain. NOAA states: "The WMM source code is in the public domain and not licensed or under copyright. The information and software may be used freely by the public." None. Attribution to NOAA NCEI, NGA, the UK Defence Geographic Centre and the British Geological Survey is given because it is the honest provenance. The model is valid only from 2025-01-01 to 2029-12-31; requests outside that window are refused rather than extrapolated.
Published astronomical series — ELP-2000/82B, VSOP87D and the IAU 1980 nutation series Numeric coefficients only, and only a derived truncation of them: 644 lunar terms, 153 solar terms and 36 nutation terms, re-derived from the pinned upstream files by a committed refresh script. No prose, table caption or commentary is reproduced, no upstream file is redistributed, and the API serves only computed results — positions, illuminated fractions and event instants — never the tables themselves. Numeric constants of published astronomical theories. ELP-2000/82B (Chapront-Touzé & Chapront) and VSOP87D (Bretagnon & Francou) are the work of the Bureau des Longitudes / Observatoire de Paris, distributed openly through the CDS as catalogues VI/79 and VI/81; the IAU 1980 nutation series is published by the IERS Earth Orientation Centre. Attribution to each theory and its publication, given here, in the api/src/data/moon-terms.ts header and on the /apis/moon page. Every upstream file is pinned by SHA-256, and the shipped coefficients are proved against JPL Horizons (DE441) and the U.S. Naval Observatory rather than against any secondary source.
Jean Meeus, Astronomical Algorithms (2nd ed., 1998), chapter 27 Numeric coefficients only: the mean equinox/solstice expressions and the 24 periodic corrections behind /apis/sun's /seasons route. No prose, worked example or commentary is reproduced, and the API serves only the computed instants. Astronomical Algorithms is © Jean Meeus / Willmann-Bell. What is reproduced is the numeric constants of a published astronomical expression. Attribution to the author and edition, given here and on the /apis/sun page. Values are verified against the U.S. Naval Observatory's published equinox and solstice table. /apis/moon deliberately takes no coefficients from this book: its worked examples are used only as test oracles, which is comparison against a published number, not reproduction.
W3C Web Application Security specifications (CSP Level 3 and 2, Trusted Types, Mixed Content, Upgrade Insecure Requests) The CSP directive reference table: 30 directive names with their fallback chains, value grammars, Fetch destinations, meta/report-only applicability, nonce/integrity/strict-dynamic flags, IANA registration flag and defining-section anchors. Facts only — every descriptive sentence in the dataset is our own wording and no specification prose is reproduced. Pinned to a dated CSP Level 3 Working Draft by SHA-256, with the IANA Content Security Policy Directives registry as a secondary cross-check. W3C Software and Document License (permissive) for the specifications; the IANA registry is CC0 1.0. No redistribution obligation for the facts. The W3C is credited here and on the /apis/csp page, and each row links to the exact defining section. Re-verified by a committed refresh script.
Published format specifications (IETF, ISO/IEC, ITU-T, ECMA, W3C, The Open Group, Unicode, UEFI Forum, ICC, Microsoft, Adobe, Google, RARLAB, Oracle, Apple, SQLite, Apache, MP4RA) Nothing is copied. The file-signature table holds numeric constants — magic bytes, their offsets and the end-of-file markers some formats define — read from each format's own specification, together with a citation naming that document and its clause: 116 patterns across 99 formats and 84 distinct citations. No specification text, table or third-party signature list is reproduced. Numeric constants are facts; the cited documents remain the property of their publishers. Several are paywalled (ITU-T T.81, ITU-T T.800, ISO/IEC 14496-12) — those rows carry the document and clause with no link rather than a decorative one. None. Every row cites its source document and clause on /apis/file-signatures, and a committed verification script fails if any row loses its citation. Wikipedia's "List of file signatures" was not used, not even as a checklist: its text is CC BY-SA 4.0.

Cited source per endpoint

Trademarks

  • QR Code is a registered trademark of DENSO WAVE INCORPORATED. We implement the ISO/IEC 18004 symbology and are not affiliated with, endorsed by or certified by DENSO WAVE.
  • GS1, EAN and UPC are trademarks of GS1 AISBL. We implement the published encoding and check-digit procedures and are not a GS1 member organisation or an issuer of identifiers.
  • Google, Android and Chromium are trademarks of Google LLC. This API redistributes transformations of the Apache-2.0 / CC BY 4.0 licensed libaddressinput address metadata and the Apache-2.0 licensed libphonenumber numbering-plan metadata, and is not affiliated with, endorsed by or certified by Google.
  • Unicode and the Unicode Logo are registered trademarks of Unicode, Inc. We redistribute Unicode data files under the Unicode License v3 and are not affiliated with, endorsed by or certified by Unicode, Inc.
  • FOCUS is a trademark of the FinOps Foundation. FOCUS — the FinOps Open Cost and Usage Specification — is published under CC BY 4.0 by a Series of the Joint Development Foundation Projects, LLC. We implement the published column vocabulary and are not affiliated with, endorsed by or certified as conformant by the FinOps Foundation.
  • All other product, company and standards-body names are the property of their respective owners and are used only to identify the specification an endpoint implements.

Corrections

Reference data drifts: registries publish new rows, standards get obsoleted, and a snapshot that was right last quarter can be wrong today. Bundled registries are pinned to a stated upstream version or date and re-checked by a committed parity script rather than edited by hand. If you find a value that disagrees with its cited source, the source wins — and it is a bug worth reporting.

The full sourcing policy, including which upstreams are deliberately not used, lives alongside the catalog at /docs.