The Detector Cannot See the Drawing Board
On AI writing, authorship, and the false certainty of a percentage
A few days ago, just before I could publish an essay, I was offered a scan for AI-written text. The instruction looked harmless. Press a button, wait a few seconds, receive a percentage. The interface suggested that a difficult question had already been solved. Somewhere inside the finished article, the machine would find the machine.
Then the score appeared.
What it could not show was how the essay had come into being. It could not see the original argument in my notes. It could not see the references I had collected, the instructions I had written, the drafts I had rejected, or the paragraphs I had returned because they sounded polished but said nothing. It could not know which sentence began with me, which one began with a language model, which one had been rewritten six times, or which claim I had removed after checking the source. It could not know whether I had read every word or merely skimmed the surface and pressed publish.
It saw the wall. It did not see the drawing board.
This is the basic problem with the debate around AI detectors. We are asking a tool that examines the statistical surface of language to answer a moral question about authorship. These are not the same task. A detector may estimate whether a passage resembles text associated with a language model. It cannot reconstruct the chain of intention, labour, rejection, research and judgment that produced the passage. It cannot tell us where the idea originated. It cannot tell us who carried responsibility for it. Yet the percentage arrives with the visual authority of a concrete test result, and institutions often treat it as though the machine has discovered a hidden fact.
That is a dangerous category error.
A score is not a history
AI detection is not magic. Most detectors look for patterns in the submitted text: statistical predictability, linguistic features, model signatures, or combinations learned from examples of human and machine writing. The field is developing quickly, and some current systems perform much better than the crude detectors that appeared after ChatGPT. A 2026 study comparing Pangram, GPTZero, Copyleaks and Turnitin found that one tool performed strongly on its test set and that false positives were rare across the four systems. But the same study found substantial differences between tools, especially with hybrid and “humanised” text, and concluded that detection should be an initial flag rather than sole evidence in a high-stakes decision.
Authorship is a layered process, not a binary result. The score reads the top sheet; the judgment lives in the thick zone beneath it.
Even Turnitin tells educators that its report may misidentify human, AI-generated and AI-paraphrased writing, and that the score should not be the sole basis for action against a student. It no longer displays exact percentages below twenty per cent because that range is less reliable. OpenAI withdrew its own text classifier in 2023 because of its low accuracy. Research has also shown that paraphrasing, new models, unfamiliar domains and ordinary human editing can weaken detectors. Earlier systems were especially troubling for non-native English writers: a Stanford-led study found that several widely used detectors repeatedly misclassified their work.
None of this means that every detector is useless. That would be another lazy binary. Detection can be useful when examining industrial spam, fake reviews, coordinated misinformation, mass-produced assignments or a sudden break in a student’s established work. A smoke alarm can be useful without being allowed to convict someone of arson. The problem begins when a probability becomes a verdict.
Wikipedia’s own field guide on signs of AI writing calls itself descriptive rather than prescriptive. It warns that many supposed signs also occur in human writing, advises against relying solely on detectors or instinct, and says the patterns may point towards deeper problems rather than constitute the problem themselves.
A paragraph is not poor because it contains a word that a model often uses. It is poor when the word performs work the thought has not earned. “This marks a significant shift” fails when nothing significant has shifted. “This highlights” becomes empty when it repeats the previous sentence with a raised eyebrow.
We have begun confusing symptoms of bad writing with evidence of machine origin. That can punish careful writers for being orderly and non-native writers for being linguistically predictable.
The detector cannot tell those stories. It receives only the final page.
Architecture has never been made by one hand
Architects should understand this problem better than most people because architecture has never been a pure act of solitary production.
Consider a building. One person originates the concept. Another develops the drawings. Engineers alter the grid and services. A contractor proposes a different junction. A craftsperson resolves a detail no drawing anticipated. Software calculates, coordinates and warns. The client approves one option and kills another. Decisions accumulate until the building stands.
Who designed it?
What the detector sees, and what it cannot: the source cards, the rejected directions, the redlines and the iteration log that produced the building.
We do not answer by locating the hand that drew the largest number of lines. We look for design intent, control over consequential decisions, coordination and accountability. We ask who can explain why the courtyard is there, why the section drops at that point, why the column moved, and what was surrendered to preserve something more important. Authorship in architecture is distributed, but it is not meaningless.
The same distinction applies to writing with AI.
Suppose I begin with an argument: AI has made production abundant, which makes judgment more valuable. I develop the architectural analogy, collect evidence, define the audience, set restrictions on tone, reject generic drafts, correct weak reasoning, verify quotations, change the sequence and approve every sentence. A language model may have arranged much of the prose, but the work is not adequately described by the phrase “AI wrote it.” That description ignores the brief, the criticism, the revisions and the person who accepts responsibility for the published claim.
Now consider another process. Someone enters a topic, accepts the first fluent response, reads the opening and conclusion, fixes two commas and submits the document under their name. Here, “I wrote this” becomes difficult to defend. The person selected an output. Selection is a decision, but one decision does not automatically amount to authorship.
The difference is not the presence of AI. It is the location of judgment.
This is also why time spent is relevant but not decisive. I may spend ten hours instructing and revising a model, and that effort can contain real authorship. I may also spend ten hours polishing a borrowed argument without adding any intellectual ownership. Authorship is not billed by the hour. It rests in meaningful control over the work and in the willingness to answer for it.
We keep searching for a thin line because institutions like thin lines. In practice, there is a thick zone. Grammar correction is not argument generation. Research assistance is not invented evidence. Asking for five openings is not accepting an essay unread. A model drafting from a detailed human brief is not identical to manual composition, but neither is it intellectual absence.
The categories must become more precise than “AI” and “not AI.”
The old problem of influence
Architectural writing makes the detector’s limits even clearer because influence has always moved through language.
Le Corbusier wrote that “a house is a machine for living in.” Decades later, Charles Correa described the house as a machine for dealing with India’s often hostile climate. The word machine remains, but the argument moves. In one formulation, the metaphor points towards standardisation and functional precision. In Correa’s hands, it is pulled into heat, shade, monsoon, ritual and the everyday intelligence of the open-to-sky space. One architect does not merely repeat the other. He takes an inherited frame and makes it answer to another place.
Louis Kahn asked a brick what it wanted to be, and the brick answered: “I like an arch.” A language model can reproduce the device. It can make concrete speak or ask timber what span it desires. The syntax is easy to imitate. The intellectual act is harder. Kahn was training architects to attend to the possibilities and limits contained in material.
This gives us two very different kinds of imitation.
A student may read Kahn and write a dialogue with laterite. An architect may absorb Correa until every project begins with shade and climate. This is not automatically plagiarism. Influence becomes original work when it passes through a new mind and a new problem.
Now add an LLM.
I can tell the model: do not copy Correa’s sentences; use the underlying spatial logic. Begin with climate as a generator. Treat the veranda as inhabited infrastructure rather than decoration. Explain an AI system as a threshold between human intention and machine production. Keep the language clear. Avoid borrowed metaphors unless they are transformed by the argument.
The model can produce a paragraph carrying my selected influences. A detector cannot tell whether that inheritance was chosen by me, retrieved by the model, or absorbed through years of reading. It cannot know whether I rejected twelve versions because they used Correa as decoration rather than thought.
The reverse is also true. I can manually write a paragraph that imitates the clean symmetry, generic significance and smooth transitions associated with AI output. It may be entirely human and entirely bad. A detector might accuse it of machine origin, but the actual failure would be mine.
This is why AI detection must not be confused with plagiarism detection, source verification or criticism. A human can plagiarise without touching AI. A model can generate a sentence that is not copied from any identifiable passage but still carries a borrowed idea without attribution. A detector can give a low score to a fabricated essay. It can give a high score to an honest one. It tells us almost nothing about whether the sources exist, whether the argument is original, or whether the writer understands what has been published.
Those are the questions that matter.
The hypocrisy is not where we think it is
There is an obvious irony in using AI to detect AI. But irony alone is not an argument. We use machines to inspect machine-made objects all the time. A laser can check a component cut by another machine. Software can find errors in software. There is nothing inherently hypocritical about using computation to identify computational patterns.
The hypocrisy begins elsewhere.
It begins when an institution uses a probabilistic model to claim moral certainty. It begins when a university permits grammar correction, translation, citation tools and automated feedback, then acts as though authorship remained untouched until ChatGPT. It begins when an editor accepts invisible machine assistance but treats one percentage as proof of dishonesty.
Most of all, it begins when the institution refuses to examine process.
A detector is attractive because it converts a difficult conversation into a number. A teacher need not compare drafts, question sources or listen to a student defend the argument. A publication need not define acceptable collaboration. The score appears, and uncertainty is pushed onto the accused.
Writing has always been evaluated through evidence larger than the finished page. Supervisors know their students’ development. Editors see drafts. Studio juries ask designers to explain decisions. Researchers preserve notes, citations and versions. We already possess a better starting point.
Architecture schools, in fact, have an advantage. We have never believed that the final rendering alone proves design ability. We ask for plans, sections, models, iterations and a verbal defence. We watch how a student responds when a critic removes the seductive image and asks why the building is organised that way. A student who cannot explain the project is exposed, even if the boards are beautiful.
The design jury asks for intent, evidence, reasoning, revision and accountability. A detector asks only whether the surface resembles a machine.
Writing in the age of AI needs the equivalent of the design jury.
Ask the writer to explain the claim without reading it. Ask why one source was trusted and another rejected. Ask for a draft. Change a condition and see whether the argument can be revised. Ask what the AI contributed and what the human changed. These questions examine understanding. A detector examines resemblance.
One method produces evidence of authorship. The other produces a score.
Not all AI assistance is equal
We also need to stop pretending that every use of AI occupies the same ethical position.
When I write a paragraph manually and use a tool to correct grammar, the machine acts after the thought has formed. When I use it to locate papers or organise notes, it enters research, so every source must be checked. When I ask it to challenge my argument, it acts as a critic. When I ask it to draft from a detailed brief, it enters composition itself.
These processes are different. They should be described honestly.
The full-draft method places a greater burden on the human author. Fluency can conceal errors, and polished paragraphs can make weak connections feel complete. The writer must read slowly, verify every external fact, remove claims that exceed the evidence, restore the specific experiences the model smooths away, and recognise sentences that sound like them without containing their thought.
This may take as long as manual writing. Sometimes it takes longer.
Equal effort does not make the methods identical. In manual writing, the author owns the first formation of sentences and the struggle that often clarifies thought. In AI drafting, some of that struggle moves into briefing, selection and revision. We may gain range and alternative structures. We may lose discovery through composition and the accidental phrase that appears before the argument is settled.
The honest position is not that these methods are the same. It is that difference does not automatically establish inferiority, fraud or absence of authorship.
The stronger test is this: who originated the problem? Who determined what counted as evidence? Who made the consequential choices? Who verified the claims? Who can defend the finished work, revise it under pressure and accept responsibility when it is wrong?
Recent research on human-AI co-writing has begun moving in this direction. It treats authorship as a question of agency, ownership and control across different stages of writing, rather than reducing collaboration to the presence or absence of generated sentences. Research on attribution also recognises human-LLM text as a distinct and difficult category because human and machine contributions can begin, alternate and blend in many ways.
That is much closer to reality than a binary percentage.
What we should look for instead
The answer is not to abandon standards. AI makes standards more necessary because it can produce finished-looking work before the thinking is finished.
We should be severe about fabricated citations, borrowed ideas, factual errors, generic argument, undisclosed automation where disclosure is required, and the submission of work the named author cannot explain. We should be equally severe about institutions that make accusations without sufficient evidence.
For academics, disclosure should be tied to the purpose of the work. If an assignment is intended to test unaided writing, using an LLM to compose it may defeat the exercise even when the resulting prose is accurate. If the task is to develop an architectural argument with permitted AI assistance, the process should be documented and assessed. A thesis, journal article, professional report and personal essay do not carry identical obligations. Policies must state what they are protecting.
For public writing, I care less about whether a model touched a sentence than whether the writer stands behind it. Did they verify the factual claims? Are the sources real and correctly represented? Is the argument theirs to make? Did they use the machine to deepen the work or to avoid doing it? Could they continue the conversation after the article ends?
The common signs of AI writing remain useful as editorial warnings. Remove declarations that have not been earned. Replace vague authority with named evidence. Watch for automatic contrasts, convenient sets of three and conclusions that announce future challenges because the model has run out of thought. Keep the simple verb when it is right. Use clear language and one variety of English. Wikipedia’s Manual of Style asks for straightforward, understandable language and consistency, while its policy page says rules should be applied with reason and common sense.
Good AI-assisted writing does not come from learning how to fool a detector. It comes from refusing the model’s average answer. It requires personal material, exact claims, real sources, deliberate structure and the courage to delete a fluent paragraph that has no reason to exist.
The goal is not to make AI writing look human.
The goal is to make the human contribution impossible to remove without collapsing the work.
Who can answer for the sentence?
We are entering a period in which the finished artefact will reveal less about its means of production. Images, buildings, films and essays will move through human and machine systems in different proportions. A clean border around “purely human” work may feel reassuring, but it will not describe how serious work is being made.
That does not mean authorship disappears. It means authorship moves.
It moves from manual production towards direction, selection, verification, transformation and responsibility. This movement is already familiar to architects. The architect does not place every brick, calculate every beam or draw every line. Yet the architect cannot escape responsibility by pointing to the draughtsperson, engineer, software or contractor. The more distributed the production becomes, the more clearly responsibility must be located.
Writing should be treated the same way.
A person who asks a model for an essay, skims the result and publishes it has delegated both labour and judgment. Calling that person the sole author is difficult. A person who develops the argument, establishes the sources and constraints, interrogates the drafts, rewrites where necessary, verifies every claim and accepts public responsibility has done something fundamentally different, even when the model generated many of the sentences.
A detector cannot distinguish those two people because their final texts may look statistically similar.
That is not a temporary bug at the edge of the system. It is a limit in the question we have asked the detector to answer. We handed it a finished wall and asked for the history of the building.
The better future will not be built on purity tests. It will require clearer disclosure where the context demands it, visible process in education and research, stronger source verification, and a more mature account of collaboration. Detectors may remain one instrument among many. They should never become the judge.
When we read a piece of writing now, we should ask more than who or what produced the first draft. We should ask where the argument came from, how it changed, what was rejected, what was verified, and who is prepared to stand behind it when challenged.
The machine may have arranged the words.
The author is the one who can answer for them.
Key Sources
Detection accuracy — Comparative study of AI-text detectors, including hybrid and humanised text (International Journal for Educational Integrity, 2026).
Turnitin — Using the AI Writing Report, the company’s own guidance on misidentification and the twenty per cent threshold.
OpenAI — New AI classifier for indicating AI-written text, carrying the company’s own note that the classifier was withdrawn on 20 July 2023 “due to its low rate of accuracy”.
Bias against non-native writers — Liang et al., GPT detectors are biased against non-native English writers (Patterns, Cell Press, 2023). Seven widely used detectors; over half of the non-native English samples misclassified as AI-generated.
Wikipedia editorial guidance — Signs of AI writing, Manual of Style, and Policies and guidelines.
Human–AI co-writing — Research on authorship, agency, ownership and control across the stages of writing.
Le Corbusier — The text containing “a house is a machine for living in”.
Louis Kahn — Collected quotations, including the exchange with the brick.
I’m Sahil Tanveer. I run RBDS AI Lab, where we study what AI is actually doing to architectural practice — through built work, research, and teaching.
The book, the courses, the channels, the free tools, past talks — everything is in one place: rbdsailab.com.
Enquiries: sahil@rbdsailab.com






