Sefer APIs

How we establish trust

Every edge carries an honest, earned statement of how far to trust it.

The graph makes a specific promise: it may contain weak claims, but it never conceals their weakness. This page states how far that promise is kept today, and where it is still a design rather than a mechanism. Holding those apart is itself the promise — a trust page that overstated its own machinery would be the first thing to distrust.

Measured against build sefer-serving-v3 on 2026-09-11. The counts and the vocabulary behind this page are on The shape of the graph.

Two kinds of edge, told apart by construction

Every edge carries a provenance, and there are exactly two values.

Inherited citations — 6,536,721 edges, provenance: curated. Sefaria's curated link set, taken as given. These are the citations of record: widely relied upon, and not derived here.

Machine-extracted citations — 2,178,134 edges, provenance: gemini_l1. Extracted here from the text, and 99.87% of them carry evidence_hebrew — the exact Hebrew words in the source that ground the claim. That is what makes an extracted edge checkable rather than asserted.

The two are separable by construction, and more cheaply than most graphs manage: their citation_type vocabularies are disjoint. No type appears under both provenances, so a caller can tell where an edge came from without consulting the provenance column at all.

What confidence means, per provenance

Confidence means two different things depending on the edge, and conflating them would be the easiest way to mislead.

On inherited edges it is an asserted constant. All 6,536,721 curated edges carry exactly 1.0, and the column holds exactly one distinct value. Read it as "this came from the curated set" — never as "this was scored and came out perfect."

On machine-extracted edges it varies and is meaningful in the ordinary sense: a 0.01.0 range with a mean of 0.872. 5,910 extracted edges sit below 0.5, and they are kept, not filtered. An edge that announces its own weakness is more useful than one quietly dropped — and dropping them would leave a corpus that looked more certain than it is.

Evidence, where it exists

For a machine-extracted edge, the proof-text is the point. evidence_hebrew carries the words themselves, so the claim can be checked against the source rather than trusted. Inherited edges carry no evidence field — there is nothing to check against, only a link to rely on. The API never presents the two as though they were the same kind of assurance.

Designed, not yet built

The following are the intended direction and are not implemented today. They are listed here because a roadmap stated plainly is worth more than a capability implied.

  • Both-end verification. The standard an extracted citation should meet — cited words found in the source at their stated location, mapping specifically enough to the target that they could not point elsewhere by coincidence — is not yet a pass that runs over the corpus.
  • An unscored state. Today every edge carries a number. There is no way to say "no legitimate basis for a confidence exists yet", which would be more honest than a low score in some cases.
  • Reviewer sovereignty. Affirm, correct, reject; a correction creating a new edge while preserving the original as an auditable record; confidence allowed to fall when evidence says so. There is no review table in the schema, so no ruling can be recorded against an edge yet.

What this means for you today

Every edge tells you its provenance, and its type tells you the same thing independently. Extracted edges show their proof-text and carry a confidence that genuinely varies. Inherited edges are what Sefaria curated, presented as exactly that.

What the graph cannot yet tell you is that a human has looked at a specific edge and ruled on it. When that exists, it will be said here in the same terms as everything else.

On this page