Open source · Apache-2.0
A reading machine that refuses to overstate what it read.
Give it topics and a frozen list of questions. It finds the literature, obtains what it legally can, and reads every source. You get, per question, the evidence bound to verbatim quotes, with the coverage that evidence rests on.
- every claim has a quote
- no pooling
- a person signs
- quote is an exact substring of the chunk
- every number and inequality appears in the quote
- “1.2” does not appear in the quote
Illustrative example, not a result from a real source.
What it is for
Ask the literature. Check the answer.
When you read the literature there are two mistakes to avoid: claiming more than the sources say, and concluding that an effect doesn’t exist just because no evidence turned up. Claimstone is for when you can’t afford either.
There are hundreds of papers on your topic and no time to read them.
It finds candidates through two independent routes, keyword search and citations, gets the legal copy of every paper it can, and reads them one by one.
AI summaries sound sure of themselves, but you can’t tell what the paper actually says.
Every claim comes with a quote copied word for word from the paper. The code checks that the quote is really there and that every number in the claim appears in it. What fails is thrown out, and the list of rejections stays visible.
“I found nothing” could mean there is nothing, or that the papers couldn’t be obtained.
It measures how much of what it found it actually obtained, against a threshold declared in advance. Below it, it stops and draws no conclusions. And it keeps three easily confused cases apart: the sources don’t answer, the sources say the opposite, the sources disagree with each other.
You need an answer you can defend, not a black box.
For each question it prepares an evidence profile: every result, how many sources point each way, what was rejected and how much was read. A person reads it and signs. Every step writes plain text files you can open and search.
And it isn’t tied to one field. Topics, questions and kinds of source are input files; nothing in the engine is specific to finance or biology.
How it works
From your questions to an answer you can check, in six steps.
You provide the topics and the questions. Claimstone does the rest, one step at a time, and leaves a file at every step that you can open and check.
You provide the topics, a fixed list of numbered questions, and the kinds of source you accept. They are three plain files.
- 1discover
Finds
Searches for papers on your topics in two independent ways: by keywords and by following citations. You get the list of candidates, each marked with what kind of source it is, such as a peer-reviewed paper or a blog post.
writes · candidates.jsonl - 2acquire
Obtains
Gets the best legal copy of each paper, preferring open access, and never goes through pirate libraries. It records every attempt and why it failed, so you know how much it could really read.
Papers behind a paywall can’t be fetched. You can add a copy you got yourself, from a library or by purchase. It goes through the same identity and full-text checks, and is reported on its own line, so it never quietly inflates how much was read.
writes · acquisitions.jsonl · raw/ - 3normalize
Prepares
Turns PDFs and web pages into clean text and splits it into passages, so every quote can be traced back to an exact place.
writes · documents.jsonl · chunks.jsonl - 4extract
Extracts
A model proposes the claims it finds in each paper. The code then checks every claim against its quote. The ones that fail are rejected and listed.
writes · claims.jsonl · rejections.jsonl - 5review
Rechecks
A second, different model rereads each claim in its whole passage, to check it still holds in context.
writes · reviews.jsonl - 6synthesize
Summarizes
For each question it assembles an evidence profile from what was accepted: the results, how many sources point each way, what was rejected and how much was read. No model, no network, no statistics at this step: it only organizes and counts.
writes · profiles.jsonl
You get an evidence profile for each question. A person reads it, decides, and signs.
The answers
Not just yes or no: five possible answers.
A question put to the studies doesn’t always have a yes or a no. Each verdict says what the evidence lets you claim, and no more. A person records it after reading the evidence profile.
The evidence points one way
SUPPORTEDThe evidence says yes.
The profile is convincing, and whoever signs writes down why.
CONTRADICTEDThe evidence says the opposite.
The profile is convincing in the other direction.
The evidence doesn’t decide, and that can happen in three ways
CONTESTED_IN_LITERATUREThe studies disagree.
The studies speak and contradict each other in a way that can’t be reconciled.
UNANSWERED_IN_LITERATUREThe studies don’t settle it.
They were read, and they aren’t enough to decide.
NEVER_ASKEDNobody has studied it.
A person checked that it isn’t just a gap in the search.
These three look the same from outside, and they aren’t. Treating “the studies disagree” as “nothing found” is the mistake Claimstone exists to avoid.
The signature is tied to the evidence the person saw: if the evidence changes later, the verdict is marked out of date. Questions are numbered and frozen, and changing the list is a dated version change.
No verdict has been signed yet. These are the states the project recognises.
Where it stands
Early, and open about it.
Claimstone works from start to finish, but it is young. Here is what has been done and what hasn’t.
Works today
All six steps run, from the search to the evidence profile. One full round was run on open-access articles from PubMed Central about screen time: 37 of the 40 papers found were obtained, above the 80% threshold set in advance. It produced 1,721 accepted annotations and 271 rejected ones, all listed.
Not yet
No verdict has been signed. The first profile is ready for a person to read; the other seven are provisional. Two other collections of papers stayed below their threshold, so Claimstone correctly produced nothing for them. The portal for following the work is read-only for now.
Where help is wanted
Reports, measurements and new fields all help. The ways to take part are just below.
Join in
Three ways to take part, from ten minutes to a whole project.
You don’t need to know the engine to help. Pick the size that fits your time.
Point out what doesn’t add up
Read how the project describes itself and tell us where it says more than it can show, or where a check lets something through. An issue with the page and the sentence is enough.
Open an issue →Put a decision to the test
Every design decision is recorded with the measurement that settled it. Pick one, repeat the measurement or bring a better one, and tell us what you find.
Read the decisions →Use it on your own field
Write the three input files for a topic you know, from public literature, and run it. Whatever the checks reject on your field is the most useful report we can get.
Follow the guide →One rule is not negotiable: no claim enters without a verified quote. The rest is open to discussion. Read the rule