Methodology

How a public body enters the index

Listing URL, scrape, OCR, hash, cite, refresh. The same path for every public body. We do not allege. We do not invent an identifier from a similar name.

We only index public bodies. A supplier, developer, or care-home Ltd appears when a public file or a published identifier attaches it. We do not run a private-company directory.

From a public list to a cited row. Not a private-company directory. Click to enlarge.

From a public list to a cited row. Not a private-company directory.

A paper-archive plate showing six steps from a public listing URL to a cited row: scrape, OCR, hash, attach on a published identifier, cite the original file.
Public bodies in. A private Ltd only with a published number. Click to enlarge.

Public bodies in. A private Ltd only with a published number.

Two columns: public bodies in the index, and private companies out unless a published number attaches them.
How a stranger walks a council. Invented names only. Click to enlarge.

How a stranger walks a council. Invented names only.

A walk diagram using invented names: search a city council, open the body, ask a topic, open a person, follow a company number, open money.

How does a public body enter the index?

A body is listed when we have a public papers URL we can collect from. The Sources catalogue is that list: live sources and planned names. A name without a listing URL stays a name. We do not activate it.

What is a listing URL?

A committee library, a board-papers page, or another public file list. An organisation homepage is not enough. NHS directory rows stay off Sources until a papers URL is confirmed.

How do you scrape?

We collect the files that page already publishes, keep the original URL, and store the file. We do not invent papers. We do not collect from a site that is not already publishing the files.

What is OCR for?

Many packs are image PDFs. We turn the image into searchable text so a name or a topic can be found. Text coverage is never 100%. A title can sit in the index before the text is ready.

Why hash a file?

We keep a hash of the file at the time of indexing. If the publisher amends or removes it, the record of what was published remains. Copyright stays with the body that published the file.

How do you cite?

Every row opens the original file or the official register. Institrace is how you found the record, not the record itself. A published identifier attaches two rows. A name in the text is labelled and is not an attachment.

How often do you refresh?

Tracked websites run on a repeating schedule. Bulk registers refresh when a new official extract is loaded. Homepage counts are a snapshot, not a live ticker.

What do you not ingest?

  • WhatDoTheyKnow is not a public shelf we scrape as if it were the body. FOI in the index is what a public body already published.
  • The national Companies House dump is not poured into People. Officer seats attach when a tracked body already has a company number, or when you paste a number on Companies.
  • Find Case Law judgments are not held. Bulk access needs a grant from The National Archives. Nothing is loaded until that grant is in force.
  • Patient records, social care case files, and email inboxes are not held.

How records attachFAQHow it worksBirmingham CouncilAbout