Reviewing AI-generated or AI-assisted text is not a matter of spotting a few unusual phrases. People can write formulaically, and generative systems can produce many styles. Editing, translation and collaboration blur the surface clues further. A responsible review therefore asks two separate questions: is the content trustworthy and fit for purpose, and was it produced in a way that follows the relevant rules?
This guide gives educators, editors, content teams and writers a repeatable human-review workflow. It uses automated detection only as an optional signal, not as a verdict. The aim is to protect quality and integrity while reducing the risk of an unsupported accusation.
Start with a policy, not a suspicion
You cannot decide whether AI use is acceptable until the rule is clear. “AI-generated” can cover very different activities: brainstorming, outlining, translation, grammar correction, source discovery, drafting, coding, image creation or rewriting. Some settings permit certain tasks if they are disclosed. Others require wholly independent work. A publisher may care primarily about factual reliability and rights, while an assessment may be designed to measure a student's unaided ability.
Before reviewing a document, write down:
- which forms of assistance were allowed, restricted or prohibited;
- what disclosure or citation was required;
- which evidence will be considered if a concern arises;
- who makes the final decision and whether there is an appeal;
- how submitted text and review records will be protected.
Apply the same process consistently. A policy invented after seeing a detector score is neither clear to the writer nor a reliable basis for a consequential decision.
The eight-step human-review workflow
1. Preserve the original submission and context
Keep an unaltered copy of the text with its submission time, prompt or brief, permitted resources and relevant policy version. If the document includes quotations, references, tables or boilerplate, note that before using any automated tool. Do not silently clean, shorten or combine material and then treat the new result as though it describes the original.
For a workplace or publishing review, record the audience, intended claims, author, editor and approval route. For education, record the learning outcome and whether drafting support was available to all students. Context helps a reviewer distinguish a quality problem from a process or policy problem.
2. Read once for meaning before looking for AI
Begin as an ordinary reader. Summarise the central claim in your own words. Check whether each section advances that claim and whether the conclusion follows from the evidence. This first pass prevents a detector result from anchoring the entire review.
Possible warning signs include abrupt changes in terminology, examples that do not fit the setting, confident claims without support, repeated abstractions or citations that appear unrelated. None proves AI use. They are editorial questions worth investigating regardless of how the text was produced.
3. Audit facts, quotations and sources
Generative systems can produce fluent statements that are incorrect, outdated or unsupported. Verify names, dates, statistics, quotations and technical claims against primary or authoritative sources. Open every cited source. Confirm that it exists, says what the text claims and is cited at the correct location. Check whether the evidence is current enough for the subject.
This fact-check is often more useful than trying to infer authorship from style. A human-written error still needs correction; a well-supported AI-assisted draft may be permitted under the applicable policy. Keep the quality judgment separate from the process judgment.
4. Check originality separately from AI classification
Plagiarism, copyright, unattributed reuse and AI assistance overlap in some cases but are not synonyms. A person can copy text without using AI, and an AI system can generate wording that does not match a source while still introducing factual or disclosure problems. Use a dedicated plagiarism check when source similarity is relevant, then inspect the underlying matches and citation context.
Do not turn a similarity percentage into an automatic plagiarism finding. Common phrases, reference lists and correctly quoted material can create matches. As with AI classification, human interpretation remains necessary.
5. Review process evidence
Authorship is a process question, so examine process evidence when the stakes justify it. Depending on the setting and privacy rules, this may include:
- an outline, notes or research questions created before the draft;
- document version history or tracked changes;
- source annotations and records of how evidence was selected;
- earlier work showing the writer's knowledge and style;
- the writer's explanation of a paragraph, calculation or design choice;
- a disclosure of tools used and the purpose of each tool.
Each item has limitations. Version history can be incomplete, style changes with subject and a nervous writer may struggle to explain good work. Evaluate the whole pattern and avoid demanding intrusive access that is disproportionate to the decision.
6. Use automated detection, if appropriate, as one signal
If the policy permits automated analysis and the text meets the tool's input requirements, an AI detector can add a classifier signal to the review. Read the provider's definition of its score, supported languages, minimum length and limitations first. EraseGPT describes its intended use in its detection methodology.
Record the exact submitted text and date. Do not interpret a percentage as the share of words produced by AI unless the tool specifically defines and validates it that way. Do not repeatedly alter the input or select the result that best fits an initial suspicion. The guide to understanding an AI percentage explains why apparently precise scores can differ.
Independent studies have shown that detector performance varies and that false classifications occur. Research has also raised concerns about bias affecting non-native English writers. These findings make a classifier unsuitable as sole evidence for a disciplinary, employment or publication decision.
7. Hold a fair, open conversation
When the review concerns another person, share the concern without presenting an automated result as established fact. Ask open questions tied to the work:
- How did you choose and organise the sources?
- Can you explain the reasoning behind this conclusion?
- Which tools did you use while planning, drafting or editing?
- How did you verify the factual claims?
- Do you have notes or an earlier version that would clarify the process?
Give the writer the relevant policy, the material being questioned and a meaningful opportunity to respond. Consider language background, accessibility needs and whether expectations were communicated in advance. Escalate only through the organisation's established process.
8. Decide and document on the full evidence
Separate findings into three columns: confirmed facts, unresolved indicators and reasonable alternative explanations. A confirmed fabricated citation is a quality finding. A classifier score is an indicator. A standard template may explain repetitive structure. State which policy clause applies and why the combined evidence meets—or does not meet—the required standard.
The record should identify who reviewed the material, the sources checked, any automated system used, the writer's response and the reason for the outcome. Avoid recording a stronger conclusion than the evidence supports. Retain data only as long as the governing privacy and records policy allows.
A practical quality checklist for AI-assisted drafts
When AI use is permitted, human review should improve the work rather than merely relabel it. Before publication or submission, confirm that:
- Purpose: The draft answers the actual brief and provides something useful to its intended reader.
- Accuracy: Every material fact, calculation and quotation has been checked.
- Evidence: Claims link to suitable sources, and citations have not been invented or misrepresented.
- Original contribution: The author has added relevant analysis, examples or experience rather than accepting generic output.
- Voice and clarity: Terminology is consistent, vague filler is removed and important limitations are stated plainly.
- Fairness: The text does not reproduce stereotypes, exclude relevant groups or make unjustified inferences.
- Rights and confidentiality: Inputs and outputs do not expose protected information or reuse material without permission.
- Disclosure: Tool use is acknowledged when the policy, publication or audience requires it.
- Accountability: A named person accepts responsibility for the final version.
Responsible revision is not detector evasion
A good revision makes the text more accurate, specific and useful. It does not aim to disguise prohibited assistance. If a draft is generic, return to the evidence and add analysis that the author can defend. If the tone is wrong, rewrite for the audience. If a passage is unclear, simplify it without changing the claim. A rewriting assistant may suggest alternatives, but a person must verify the output and follow the applicable disclosure rules.
Do not add errors, awkward phrasing or random variation in an attempt to influence a classifier. Do not replace words mechanically when that makes the meaning less precise. Passing a detector is not a quality standard, and a changed score does not prove that a revision became more human.
How the workflow changes by setting
For teachers and academic integrity teams
Design assessment instructions that state permitted assistance before students begin. Prefer evidence of learning—drafts, discussions, source notes, oral explanation and in-class work—over surveillance of prose style. If a concern arises, follow the institution's established process and do not impose a penalty from an automated result alone.
For editors and content teams
Focus on accountable publication: fact-checking, source quality, first-hand expertise, reader benefit, rights and disclosure. Search systems reward helpful content rather than a particular production method; Google's published guidance tells creators to focus on quality and purpose. Maintain an editorial trail showing who verified and approved the final work.
For writers reviewing their own work
Keep research notes and disclose assistance as required. Treat automated feedback as a prompt to inspect clarity, not a target to game. If a checker returns a surprising result, save it, read the limitations and review the draft against your source trail. Your strongest evidence is a transparent process and work you understand well enough to explain.
A compact review record
A useful record can fit on one page. Include the document and policy version, reviewer, review date, factual or citation issues found, process evidence considered, automated tools and their stated limitations, the author's response, changes requested and final decision. This makes later moderation or appeal more consistent and helps a team improve its policy.
The central principle is simple: assess what the content says, verify where it came from and judge the process against a clear rule. Technology may help prioritise attention, but responsibility for the conclusion remains with people.
Sources and further reading
- UNESCO, Guidance for generative AI in education and research (2023)
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024)
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023)
- Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023)
- Google Search Central, Creating helpful, reliable, people-first content