Built on public dataBuilt on public data. Not affiliated with the USPTO or any government.
Open Nshipyard
Patent Index
frSign in (demo)

Methodology

How the snapshot was built, what was normalized, what was left out, and how each figure is computed.

Data sources

Google Patents (patents.google.com)

The corpus source. Candidate generation used the public XHR search endpoint with G06N queries over 3-day publication windows (the endpoint returns the top 10 relevance-ranked results per query and ignores pagination, so breadth came from 1,000+ distinct queries, not pages); every candidate was then verified by fetching its detail page and keeping only documents whose listed CPC classifications include G06N. Bibliographic fields (title, abstract, assignee, inventors, dates, CPC codes) come from the detail page microdata.

https://patents.google.com/
USPTO (uspto.gov)

The authoritative issuer of US grants and applications in the corpus. The old keyless search API (api.uspto.gov/patents/v1/search) was retired; its replacement requires an ID.me-verified API key, and bulk downloads were unreachable from the build environment, so the build used Google Patents as the retrieval layer and documents this substitution plainly.

https://www.uspto.gov/
Google Patents Public Datasets on BigQuery

Noted as the global-coverage source for future versions. Not used at runtime because it needs a GCP project with billing.

https://console.cloud.google.com/marketplace/product/google/patents-public-data

Scope and normalization

  • CPC class G06N only (computing arrangements based on specific computational models: the machine-learning class).
  • Publication dates from 2025-10-01 to 2026-05-22 (snapshot built 2026-10-09), sampled in 3-day windows for uniform temporal coverage.
  • Documents kept only when the detail page listed at least one CPC code starting with G06N and a non-empty abstract. Figures describe this verified sample, not the full publication universe: the retrieval layer returns the top relevance-ranked candidates per query and ignores pagination, so monthly counts show sample composition, not true USPTO publication volumes.
  • Assignee is the assigneeOriginal string from the record; inventor strings are not merged across spelling variants.
  • Deduplicated by publication number.

How each figure is computed

  • Filings per month: count of snapshot documents by publication-date month (sample composition, not a census of publications).
  • Top assignees: count by normalized assignee string (case-insensitive trim), top 15 shown.
  • CPC subgroups: for each document, its distinct G06N subgroups (e.g. G06N3/08); share = documents carrying the subgroup / total documents.
  • Top inventors: count by raw inventor string, top 15 shown, variants unmerged.
  • Semantic search: cosine similarity between the query embedding and each document embedding (title + abstract, text-embedding-3-large), top 20 returned.

What was left out

  • Full claims text: similarity and summaries use titles and abstracts only.
  • Legal status beyond publication: grant vs application is recorded from the document kind code, not live status.
  • Non-English records are included with their English-language Google Patents metadata where available.