What to expect
What gets recomputed, how findings are worded, and when to run it before a reviewer does.
What you get
When to use it
How it works
Three steps. You do one of them.
Upload and verify
Upload the document. No dataset, no variable list, nothing to anonymise — the check reads the statistics out of the text itself.
The engine reads your document
- Recomputing every number
- Extracting every reported statistic
- Recomputing p-values from your test statistics
- Running GRIM on reported means
- Compiling the findings
- Checking group totals and percentage breakdowns
- Flagging effect sizes that should be reported
- Listing each inconsistency with its passage
An illustration of the stages this engine works through. Most reports are delivered within minutes; a full thesis takes longer.
A report you can act on
Sample report
A real report from this engine — not a mock-up. Scroll it here, or expand it to full length.
See exactly what you get
A real report from this engine — not a mock-up, not a marketing illustration. Scroll it here, or expand it to full length.
Published with permission. The research title is obscured and verbatim extracts of the author’s document have been removed. Short quoted terms are kept, so you can see the engine read this specific document rather than producing generic advice.
Why this check exists
Reviewers increasingly run consistency checks on submitted manuscripts as a matter of routine. If your reported statistics do not agree with each other, it is far better that you find out than that they do.
Numbers drift during write-up. An analysis is re-run and the table is updated, but the sentence describing it is not. A p-value is transcribed from output with a digit transposed. A percentage is calculated against the wrong denominator. Subgroup counts stop summing to the total after an exclusion is applied late.
None of this is misconduct. All of it is invisible to you after the tenth read, because you are checking meaning rather than arithmetic. And all of it is now trivially detectable by anyone who cares to look — which, in peer review, increasingly means the reviewer, the editor, or an automated screen the journal runs before your paper reaches a human.
Finding an inconsistency yourself is a five-minute correction. Having it found for you, after publication, is something else entirely.
What gets checked — every number, every time
Not a sample. Not the ones you flag. Every statistic in the document, recomputed independently of what you typed.
p-values recomputed
Each p-value is recalculated from the test statistic and degrees of freedom you reported. Where the recomputed value disagrees with the printed one, it is flagged.
GRIM — means against sample sizes
A reported mean has to be arithmetically possible given the number of participants and the scale used. Some are not. GRIM finds those.
Group sizes and percentages
Subgroup counts are checked to total, and percentage breakdowns are checked to sum correctly against the denominators you state.
Effect sizes
Where an effect size should conventionally be reported and is absent, the report says so — increasingly the first thing a methods reviewer looks for.
Findings are inconsistencies, not accusations
Everything this check reports is framed as something to re-check, never as an error and never as misconduct. That framing is accurate, not diplomatic: the overwhelming majority of the differences it finds are rounding, transcription slips, or a table updated after a re-run while the sentence describing it was not.
Those are precisely the mistakes that are invisible to you after the tenth read and obvious to a reviewer on their first.
The statistical reporting errors that appear most often
Research into published papers has repeatedly found that a substantial share contain at least one internally inconsistent statistical result. Almost none of that is misconduct. It is arithmetic drift across a long document, and it is invisible to the person who wrote it.
- p-values inconsistent with the test statistic. The reported p does not follow from the reported test statistic and degrees of freedom. Usually a transcription slip or a value left over from an earlier analysis run.
- Impossible means. A mean that cannot arise from the stated sample size and scale. The GRIM test identifies these, and they are more common than most authors expect.
- Group sizes that do not sum. Subgroup counts that fail to total, typically after exclusions were applied late and one table was updated while another was not.
- Percentages against the wrong denominator. Percentages calculated on the full sample when the analysis used a subset, or the reverse.
- Degrees of freedom inconsistent with reported n. Often the clearest signal that the analysis was re-run and the text was not fully updated.
- Missing effect sizes. Increasingly the first thing a methods reviewer checks, and increasingly a condition of acceptance.
- Rounding that changes significance. A p-value rounded across the threshold it sits beside.
Every one of these is detectable from the document alone, without your dataset. That is precisely what this check does — on every number, not a sample.
Who it is for
Before submission
The check takes minutes and removes an entire category of reviewer objection.
Before resubmitting after revision
Revision is exactly when numbers get updated in one place and not another.
For multi-author papers
When several people have touched the results, no single author has checked the whole set.
When your analysis was re-run
The single most common source of inconsistency is a re-run analysis with partially updated text.
You do not need to send your dataset
This check works entirely from the statistics reported in your document. There is no raw data to upload, no variable list to prepare, and nothing to anonymise. Upload the document and the check reads the numbers out of it.
The minimum is around 200 words — enough text to contain reportable statistics.
What it does not do
This is a consistency check, not a statistical review. It tells you whether the numbers you report agree with one another. It does not tell you whether you chose the right test, whether your assumptions held, or whether your interpretation is sound.
Those are design questions, and they belong to the Methodology Check, which reviews your analysis and whether your conclusions follow from it. The two checks are complementary and many people run both.
What happens to your document
Your file is used for one purpose: producing your report. Once the report has been generated, the source document is deleted from our servers. It is not kept for training, it is not shared, and it is not readable by anyone who has not verified the email address the report belongs to.
This matters more for academic work than for most things people upload. An unpublished manuscript or an unexamined thesis is the one document in your career you cannot afford to have circulating, and a service that quietly retained it would be a liability rather than a help.
See a real report before you buy
Sample reports showing how inconsistencies are flagged, explained and prioritised.
View Statistical Consistency samples →Our guarantees
Common questions
Do I need to upload my raw data?
No. The check reads the statistics reported in the document itself.
What if it flags something that is actually correct?
Then you have spent a minute confirming it. Findings are reported as inconsistencies to re-check, not as errors, precisely because some will have an explanation you know and the document does not state — which is itself worth knowing, since a reviewer will not know it either.
Does it check my statistics are appropriate?
No. It checks internal consistency. Whether the test was the right one is a design question covered by the Methodology Check.
What is GRIM?
A check on whether a reported mean is arithmetically possible given the sample size and the scale used. Some reported means cannot occur with the stated number of participants, and GRIM identifies those.
Can I run it on the results section only?
Yes. Around 200 words is the minimum.
How long does it take?
This is the fastest of our checks — it is mostly computation rather than model calls. Minutes.
Minutes, not months
The alternative to this is waiting. Waiting for a supervisor with six other students, waiting for a reviewer who has your manuscript for four months, waiting for a viva to discover what you should have known before you submitted.
Upload your document and the report exists before you have finished your coffee. You do not book anything, you do not wait for a slot, and you do not explain your project to anyone.
Ready to run it?
No dataset, no setup, no waiting. Upload the document and the check reads the numbers straight out of it.
Start your check ↑