Read the AI-Generated Article 100% AI-generated · not peer-reviewed · click to expand
Abstract
Fluent, on-demand text generation lands differently in qualitative research than elsewhere, because the field note is where observation becomes data: the choice of what to record and the words that fix a fleeting expression are interpretive acts, not clerical ones. The published debate has given us a vocabulary of concern and a few hard boundary-markers — no AI authors, disclose your use, keep humans accountable — but it argues at the level of grand methodological principle or of publication policy, and skips the daily craft where researchers already improvise. I defend a position I call task-sensitive authorship. “Voice,” I argue, is not a stylistic property of prose but the evidentiary trace of a situated interpreter’s attention, and different AI operations touch that trace in radically different ways. The ethical line falls not between AI and no AI, nor between analysis and writing, but between operations a disclosed, accountable human continues to author and operations in which authorship has migrated to the model. The refusal camp has the right diagnosis — interpretive delegation corrodes, and its corrosion hides in plausibility — but overgeneralizes it into prohibition; the responsible-use frameworks have the right remedy of disciplined, audited practice but cannot say which uses are dangerous. Joining that diagnosis to that remedy at the level of the specific operation, I classify the tasks of record-keeping: expansion of jottings, primary coding, and vignette-drafting belong on the refuse side; transcription, anonymization, and reformatting are permissible under privacy safeguards. Disclosure should be itemized by operation and sorted by whether the model supplied interpretive content. What remains unsettled is empirical: whether reviewers can catch a substituted construal in fluent prose.
Introduction
A field note is a strange document. It looks like a record — a plain account of what happened in a clinic waiting room, a union meeting, a refugee shelter — but anyone who has written one knows it is nothing of the kind. The choice of what to write down and what to leave out, the phrasing that fixes a fleeting expression as “wary” rather than “tired,” the aside that connects a gesture to something a participant said three weeks earlier: these are interpretive acts performed in the moment of writing, and they carry much of the analytic weight of qualitative inquiry. The field note is where observation becomes data, and it does so through a particular human being’s situated attention. This is why the arrival of fluent, cheap, on-demand text generation lands differently in qualitative research than it does in, say, drafting a grant summary or polishing a literature review. When a large language model can expand three bullet points of shorthand into two paragraphs of readable prose, or smooth a raw jotting into a publishable vignette, it is not merely accelerating a workflow. It is inserting itself into the exact place where a discipline locates its epistemic authority.
The stakes are being felt across the qualitative community, and the published responses already span the full range from cautious embrace to outright refusal. A collectively authored editorial in the Information Systems Journal, written by an eleven-member team who openly disagree among themselves about whether generative AI should be used for qualitative analysis at all, nonetheless converges on a shared map of the ethical terrain: data ownership and rights, data privacy and transparency, interpretive sufficiency, bias, and researcher responsibility and agency (Davison et al. 2024). At the more permissive end, some scholars have proposed structured frameworks for responsible adoption, arguing that the sensible response to a powerful tool is disciplined use rather than prohibition (Eacersall et al. 2024). At the restrictive end, prominent methodologists have published flat rejections: a group including the originators of reflexive thematic analysis argue that generative AI is constitutively incompatible with reflexive qualitative work because such work depends on the researcher’s own situated subjectivity (Jowsey et al. 2025), and a separate position piece contends that embracing these tools betrays the commitments of engaged and responsible scholarship (Nguyen and Welch 2025). Meanwhile the governance infrastructure of academic publishing has already moved on one narrow point: editorial bodies have ruled that an AI system cannot be an author, because authorship entails an accountability that no non-human tool can bear (International Committee of Medical Journal Editors, n.d.), and commentators in Nature have called for continuously updated “living” guidelines and independent auditing so that oversight stays with the scientific community rather than the vendors (Bockting et al. 2023).
What this literature has established is a vocabulary of concern and a set of hard boundary-markers — no AI authors, disclose your use, keep humans accountable. What it has not established is how any of this applies to the specific, granular act of composing field notes and other primary qualitative records. The existing debate is pitched largely at two levels that skip over the middle: at the level of grand methodological principle (is generative AI compatible with reflexive inquiry?) and at the level of publication policy (can a chatbot be a co-author?). Between the philosophy and the masthead lies the daily craft — the jotting, the expansion, the memo, the coded excerpt, the vignette — where researchers are already making choices with no settled guidance. The refusal arguments treat “using AI” as a single undifferentiated act to be accepted or declined, while the responsible-use frameworks tend to offer principles broad enough to endorse almost any careful practice. Neither tells a researcher standing over a day’s messy jottings which specific operations preserve the interpretive voice that gives qualitative data its worth and which quietly dissolve it.
This article defends a position I will call task-sensitive authorship. I argue that the question “can qualitative research use AI without losing its voice?” has no single answer, because “voice” in qualitative work is not a stylistic property that lives in the prose — it is the evidentiary trace of a situated interpreter’s attention, and different AI operations touch that trace in radically different ways. Some uses of generative AI leave the interpretive core untouched and merely relieve mechanical burden; others silently substitute the model’s statistically typical construal for the researcher’s particular one, and these are the uses that hollow out the note while leaving its surface intact. The ethical line, I contend, does not fall between “AI” and “no AI,” nor even between analysis and writing, but between operations that a disclosed, accountable human continues to author and operations in which authorship has actually migrated to the model. Getting that line right requires attending to the field note as a specific genre, not to “qualitative research” in the abstract.
This is a conceptual and normative analysis. It builds on the published debate — the ethical map, the refusal arguments, the responsible-use frameworks, and the authorship rulings summarized above — and works through their implications for one under-examined object, the primary qualitative record, to arrive at a defensible practical position and a disclosure protocol that follows from it. I proceed as follows. I first characterize the field note as an interpretive act, so that the thing at risk is clearly in view. I then lay out the five-part ethical map that the literature offers and use it to locate the specific vulnerabilities of synthetic field notes. I take up the authorship question directly, then give the refusal position its strongest form before turning to the responsible-use frameworks. From that engagement I develop the task-sensitive account and a corresponding disclosure protocol. I close with the counterarguments and limits that remain genuinely unresolved.
The Field Note as an Interpretive Act
To see what is at risk, we need to be precise about what a field note is and does. In the received understanding of ethnographic and interview-based research, the written record is not a transcription of reality but a construction of it. The researcher attends selectively, and that selection is theory-laden: what strikes an observer as worth recording depends on the questions they carry, the participants they have come to know, and the analytic hunches already forming. The prose in which an observation is fixed does interpretive work — a note that a caseworker “hesitated before approving the claim” commits to a reading that “there was a two-second pause” does not. Field notes are, in this sense, the first analysis, not the raw material that precedes analysis.
This matters for the AI question because it dissolves a comforting distinction that a lot of casual thinking relies on: the idea that there is a clean division between data, which is collected and fixed, and writing, which merely dresses the data for presentation. If that division held, then using a language model to improve the writing would be ethically trivial, like using a spell-checker or hiring a copy-editor. The trouble is that in qualitative inquiry the division does not hold. The writing is where the data is made. When a model rewrites a jotting into fluent prose, it is not polishing a finished object; it is intervening at the moment of construction, and it does so by supplying the connective tissue — the causal verbs, the affective adjectives, the implied stance — that constitutes the interpretation. This is precisely the concern that Davison et al. (2024) name interpretive sufficiency: the worry that a generative system cannot supply the depth of construal that qualitative analysis demands, because that construal is grounded in a knowing that the model does not have.
The reflexive tradition sharpens the point to a fine edge. In reflexive thematic analysis, the researcher’s subjectivity is not noise to be minimized but the very instrument of analysis; themes are actively generated by an interpreter reflecting on data through the lens of their own position, not passively “found” lying in the transcripts. On this account, as Jowsey and colleagues argue, there is no coherent sense in which a language model could produce the reflexive reading, because the reading just is a particular situated person’s engagement with the material (Jowsey et al. 2025). A synthetic field note, in this frame, is not a lower-quality note; it is a category error — an artifact that has the grammatical form of interpretation while lacking the interpreter whose stance the form is supposed to express.
I think this is the correct place to start, but I want to resist collapsing the whole discussion into it too quickly, because the reflexive argument proves both more and less than it appears to. It proves that the interpretive core of a field note cannot be delegated without loss. It does not, by itself, establish that every operation involved in producing a field note is interpretive in this constitutive sense. Consider the range of things a researcher actually does around a field note: they scribble timestamped fragments during observation; they expand those fragments into fuller prose hours later, from memory; they anonymize names and places; they attach analytic memos; they later pull coded excerpts into a display; they craft a polished vignette for a paper. Some of these operations are saturated with situated judgment. Others are closer to clerical labor performed under interpretive constraints the researcher has already set. The reflexive argument, pushed to its conclusion, treats them all alike. Much of my argument in what follows turns on the claim that they are not alike, and that ethics done at the right grain must tell them apart. But before I can defend that, I need the fuller ethical map.
Five Fault Lines: The Ethical Map and Where Field Notes Sit on It
The most useful synthesis of the ethical terrain comes from the Information Systems Journal editorial, precisely because its authors disagree with each other. When eleven scholars who cannot agree on whether to use a tool nonetheless agree on what is at stake in using it, the resulting map has a claim to being more than one faction’s manifesto. Davison et al. (2024) identify five categories of ethical concern for generative AI in qualitative analysis: data ownership and rights; data privacy and transparency; interpretive sufficiency; biases manifested in generative AI; and researcher responsibility and agency. Their framing insight — that these ethical questions overarch the technical question of how the tools could be used — is one I take as a premise. The interesting work is figuring out how each fault line runs through the specific practice of writing field notes, because the answers differ by category in ways that the abstract list conceals.
| Fault line | General concern | How it bites specifically on field notes |
|---|---|---|
| Data ownership and rights | Who owns the material fed to and produced by the model? | Field notes contain participants’ words, bodies, and settings; pasting them into a commercial model may transfer or expose material participants never licensed for that use. |
| Data privacy and transparency | Where does the data go, and is its handling disclosed? | Raw notes are the least anonymized documents in a study; they are exactly the wrong text to upload to an opaque third-party service. |
| Interpretive sufficiency | Can the model supply the depth of construal analysis requires? | The note is the interpretation; a model-supplied construal is not merely thin but not the researcher’s at all. |
| Bias in generative AI | Models reproduce dominant patterns in training data. | Model expansion pulls idiosyncratic observation toward the statistically typical, erasing the marginal and the surprising — the very things fieldwork exists to capture. |
| Researcher responsibility and agency | Human accountability for AI-assisted content. | If the note’s stance is the model’s, the researcher signs an interpretation they did not form, and can no longer stand behind it as witness. |
Two of these fault lines — ownership and privacy — are, I will argue, the most tractable, and yet in the field-note context they are also the most acute, which is a combination worth pausing on. They are tractable because they concern where the text goes rather than what the text means, and questions of the first kind have concrete answers: run models locally, strip identifiers before any upload, use tools with contractual guarantees against training on your inputs. But they are acute because the field note is the single most exposed document a qualitative study produces. A published paper is anonymized; an interview transcript is at least partly cleaned; the raw field note, written in haste, is thick with real names, recognizable locations, and unguarded detail about people who consented to be studied by a named researcher, not to have their lives streamed through a corporate inference API. The concern Davison et al. (2024) group under privacy and transparency is not abstract here. It is the difference between a participant’s confidence being kept and being casually broken in the name of a faster draft.
The other three fault lines — interpretive sufficiency, bias, and responsibility — concern meaning rather than location, and they are the ones that the reflexive argument of the previous section engages most directly. Bias deserves a specific note because it interacts with interpretation in a way that is easy to miss. A language model expands terse input by predicting the most probable continuation, which means it pulls toward the center of its training distribution — the typical, the expected, the already-said. But the entire point of fieldwork is often to record what is atypical: the practice that does not fit the official account, the population whose experience is absent from the corpus that trained the model, the surprising gesture that a theory did not predict. A model asked to flesh out a jotting will tend to render it more conventional, more like the average of everything it has read, which is to say less like the specific and possibly marginal reality the researcher went out to observe. This is not a failure the model can be prompted out of; it is what statistical text generation is. So the bias fault line and the interpretive-sufficiency fault line reinforce each other: the model does not merely fail to supply the researcher’s construal, it actively supplies a construal biased toward the mainstream, and it does so invisibly, in fluent prose that gives no sign of the substitution.
Having the map laid out this way lets me state the analytic strategy for the rest of the article. The tractable fault lines (ownership, privacy) are matters of infrastructure and disclosure; they generate rules that a competent research ethics process can specify and enforce, and I will treat them as constraints rather than dilemmas. The meaning-related fault lines (interpretive sufficiency, bias, responsibility) are where the genuinely hard normative question lives, because they cannot be settled by better infrastructure — a perfectly private, locally hosted, contractually clean model still cannot supply the researcher’s situated interpretation. It is on this second cluster that the refusal arguments and the responsible-use arguments actually join battle, and it is there that I need to do the real work.
Who Can Be the Author of a Field Note?
Before adjudicating between refusal and responsible use, it helps to fix one point that is, for once, relatively settled — and to notice how much and how little it settles. Academic publishing’s governance bodies have converged on a clear rule about AI and authorship. An AI system cannot be listed as an author, because authorship carries responsibilities — accountability for the integrity of the work, the ability to approve the final version, and the capacity to answer for it — that a non-human tool cannot assume; and the humans involved must take responsibility for all content, including any portions produced with AI assistance, and must disclose how the tools were used (International Committee of Medical Journal Editors, n.d.). This is the operative benchmark across biomedical publishing and, increasingly, well beyond it.
What the rule settles is the byline. What it leaves wide open is everything upstream of the byline — and field notes are as upstream as it gets. The authorship rule is framed around the published article and its named contributors. It tells us that a chatbot cannot appear on the masthead and that a human must be answerable for whatever the chatbot helped produce. But it says nothing directly about whether a particular sentence in a particular field note, months before any article exists, was interpretively authored by the researcher or generated by the model and waved through. The accountability principle that the rule enshrines, though, has a bite that reaches all the way back to the note, and this is the connection I want to draw out.
The reason a machine cannot be an author is that it cannot be accountable, and accountability in qualitative research has a specific texture. When a researcher writes “the nurse seemed to resent the new protocol,” they are doing something more than asserting a proposition; they are staking their credibility as a witness on an interpretation they formed by being there. If challenged, they can say why — what they saw, what they knew of the nurse, what made “resent” the right word rather than “question” or “tire of.” This capacity to answer for the interpretation is exactly what the authorship rule means by responsibility. Now suppose the word “resent” was supplied by a model expanding the shorthand “N unhappy re: protocol.” The researcher signs the note, and under the rule remains responsible for its content. But responsibility here has become hollow: they are accountable for an interpretation they did not form and cannot fully reconstruct, because the specific construal came from the model’s prediction, not from their situated judgment. They can defend that they wrote down “unhappy”; they cannot really defend “resent” as their reading, because it was not.
This is why the authorship rulings, though narrowly about bylines, carry a broader implication that the responsible-use literature has not fully absorbed. Bockting et al. (2023) argue that oversight of generative AI in science must remain with the research community rather than with technology companies, and propose living guidelines and independent auditing precisely because responsibility cannot be outsourced to systems whose behavior the vendors control and change at will. Put the authorship principle and the oversight argument together and you get a demanding standard: a researcher may use a model only in ways that leave them genuinely able to answer for the result as their own interpretation. That standard does not prohibit AI use. But it does prohibit a specific and tempting kind of use — the kind where the model supplies interpretive content that the researcher then adopts without being able to reconstruct why it is right. The authorship framework, read for its underlying principle rather than its literal scope, thus already contains most of the task-sensitive distinction I will develop: it separates uses that a human continues to author from uses in which authorship has quietly migrated.
The Case for Refusal, Stated at Its Strongest
The most intellectually serious response to synthetic field notes is not caution but refusal, and it deserves to be met at full strength rather than dismissed as technophobia. Two recent statements make the case. Jowsey et al. (2025) reject the use of generative AI for reflexive qualitative research on the ground that such research is constituted by the researcher’s situated subjectivity and interpretive labor — the “voice” is not decoration on the analysis but its substance, and a machine cannot supply it. Nguyen and Welch (2025) argue, from the standpoint of engaged and responsible scholarship, that qualitative researchers should not embrace these tools at all, because doing so is incompatible with the epistemic and ethical commitments that define the enterprise.
Let me build the refusal position up rather than knock it down. Its first premise I have already granted: field notes are interpretive acts, and interpretation in qualitative work is the trace of a particular knower’s engagement. Its second premise is subtler and, I think, the strongest thing in the refusal camp’s arsenal — a claim about the corrupting subtlety of fluent text. The danger of a language model is not that it produces obviously bad interpretations that a careful researcher would catch and reject. The danger is that it produces plausible ones, phrased in exactly the confident, readable register that field notes are supposed to have, and that this fluency is disarming. A researcher reviewing a model’s expansion of their jottings is in a poor epistemic position to notice what has been lost, because the loss is invisible: the note reads well, it is consistent with what they remember, it says something a competent ethnographer might have said. The model has, in effect, filled the gaps in memory and attention with statistically probable content that the researcher then cannot distinguish from their own recollection. On this view, the “review and edit” safeguard that responsible-use frameworks lean on is close to worthless, because the thing that needs catching is precisely the thing fluent prose hides.
The third premise concerns the corrosion of skill and disposition over time. Engaged scholarship, in the sense Nguyen and Welch (2025) invoke, is not only a set of outputs but a practiced way of attending to the world. The discipline of writing your own field notes — struggling to find the right word, sitting with the ambiguity of what you saw, noticing when you cannot honestly reduce a moment to a phrase — is formative. It is how researchers develop and maintain the interpretive sensibility that makes them good witnesses. Outsource the writing and you do not merely offload a task; you stop practicing the thing that makes you capable of the task, and over a career, or over a discipline’s collective career, the capacity atrophies. This is a serious argument that no amount of careful prompting answers, because it is about what the tool does to the researcher, not about what the tool does to any particular note.
Taken together these three premises make a case that I regard as substantially correct about the interpretive core and substantially overreaching about everything else — and the overreach is instructive. The refusal position derives a blanket conclusion (“do not embrace these tools”) from premises that, examined closely, bear only on a subset of operations. The corrupting-fluency argument is devastating against using a model to generate interpretive content — to supply the affective reading, the causal connection, the thematic construal. It has no purchase at all against using a model to, say, transcribe an audio recording of your own spoken field notes, or to flag every proper name in a document for anonymization, or to reformat coded excerpts into a table. In none of those cases is the model supplying interpretation that fluency could disguise; there is no construal to be hidden. The skill-atrophy argument is powerful against habitual delegation of the interpretive writing that builds sensibility; it is weak against delegating the mechanical anonymization pass, which builds no sensibility worth protecting and which many researchers already outsource to find-and-replace.
The refusal camp will respond that these “safe” uses are a slippery slope: allow the model in for transcription and anonymization and you normalize its presence, and the interpretive uses follow by increments. I take the slope seriously — normalization is real, and the history of technology adoption in research is full of tools that entered as conveniences and ended as dependencies. But a slippery-slope worry is an argument for a bright line, not for prohibition, and the whole burden of my position is that a bright line can be drawn in the right place: at the boundary of interpretive content. The refusal position’s own best argument — the corrupting subtlety of fluent interpretation — is what tells us where to draw it. Where the model can hide a substituted construal in plausible prose, refuse. Where there is no construal to substitute, the argument falls silent. Refusal generalizes a valid local prohibition into an invalid global one, and in doing so it forfeits the chance to give researchers usable guidance about the operations it has no real objection to.
The Case for Responsible Use, and Why Principles Are Not Enough
The opposite camp offers frameworks rather than refusals. Eacersall et al. (2024) propose the ETHICAL framework, a principles-based scheme for responsible generative-AI use across the research process, designed to help researchers make disclosure, consent, and integrity decisions rather than to tell them yes or no. Bockting et al. (2023) supply the governance complement: living guidelines that update as the technology changes, and independent scientific auditing of the tools, so that the community retains oversight instead of ceding it to vendors. The animating idea of this camp is sensible and, I think, ultimately correct: a capable tool that is already in wide use is better governed by disciplined norms than by prohibitions that the honest will follow and the careless will ignore.
The responsible-use camp has three real advantages over the refusal camp. First, it is realistic. Generative tools are already embedded in the software researchers use; a prohibition that depends on collective abstinence will fail unevenly, penalizing the scrupulous. Second, it keeps oversight where Bockting et al. (2023) rightly insist it belongs — with scientists who can audit and update, rather than with a norm of refusal that, once breached, offers no fallback guidance at all. Third, it is honest about the diversity of the qualitative field: not all qualitative work is reflexive thematic analysis, and a framing that treats reflexive commitments as definitional of the whole enterprise, as the strongest refusal arguments implicitly do, over-claims.
And yet the frameworks have a characteristic weakness that mirrors the refusal camp’s characteristic overreach. Where refusal generalizes a valid local prohibition into an invalid global one, responsible-use frameworks tend to state principles so general that they endorse almost any conscientious practice and forbid almost none. “Be transparent,” “obtain appropriate consent,” “maintain integrity,” “take responsibility” are unimpeachable and nearly empty operationally. They do not tell the researcher standing over a day’s jottings whether expanding shorthand into prose is a transparency-plus-review matter or a bright-line violation. The ETHICAL framework, as a principles-based scheme, is a genuine advance over having nothing; but a principle that a researcher can satisfy by adding a disclosure sentence does not discriminate between the operation that preserves interpretive authorship and the operation that destroys it. Both can be disclosed. Both can be done “responsibly” in the thin sense. The frameworks, in other words, tell you to be transparent about what you did without telling you which things you do are the dangerous ones.
This is the crux of my disagreement with both camps, and it is worth stating sharply. The refusal camp has the right diagnosis (interpretive delegation is corrosive and its corrosion is hidden by fluency) but the wrong remedy (prohibit the tool). The responsible-use camp has the right remedy (disciplined, disclosed, audited use) but an under-specified diagnosis (it does not say which uses are corrosive, so its discipline lands everywhere and nowhere). What is missing from both is an account keyed to the actual operations that make up qualitative record-keeping — an account that inherits the refusal camp’s diagnosis and the responsible-use camp’s remedy, and joins them at the level of the specific task. Building that account is the work of the next section.
Task-Sensitive Authorship: A Proposal
My proposal starts from a single distinction that both camps blur: between operations in which a language model supplies interpretive content and operations in which it performs interpretively constrained labor on content the researcher has already authored. Interpretive content is any construal that stakes a claim about meaning — an affective reading (“resented”), a causal connection (“because the protocol changed”), a thematic assignment, a selection of what matters from a mass of detail. Interpretively constrained labor is work whose parameters the researcher has fully set, such that the model has no latitude to construe: transcribing spoken words, locating proper names, reformatting an existing table, checking whether a passage the researcher has already coded contains a term.
The reason this distinction does the ethical work is that it maps exactly onto the fault lines that survived the earlier analysis. Interpretive-content operations are where the corrupting-fluency danger lives (a substituted construal hides in plausible prose), where the bias-toward-the-typical danger lives (the model pulls the idiosyncratic toward the mainstream), where interpretive sufficiency fails (the construal is not the researcher’s), and where accountability goes hollow (the researcher signs a reading they did not form). All four of the meaning-related concerns from Table 1 fire together on exactly these operations. Interpretively constrained operations trip none of them, because there is no construal for the model to substitute, bias, or hollow out — there is only a defined transformation of text the researcher already made their own. What remain, for the constrained operations, are the two tractable fault lines of ownership and privacy, and those are handled not by refusal but by infrastructure: local or contractually clean models, and anonymization before any upload.
Table 2 sets out how this distinction sorts the actual operations of qualitative record-keeping. The classifications are my own reasoning, applying the criteria just stated; they are not empirical findings, and reasonable researchers will contest the placement of the borderline cases, which is exactly where judgment is required.
| Operation | Does the model supply interpretive content? | Dominant fault lines engaged | Proposed stance |
|---|---|---|---|
| Transcribing the researcher’s own spoken field notes | No | Privacy, ownership | Permitted with privacy safeguards and disclosure |
| Flagging proper names/places for anonymization | No | Privacy, ownership | Permitted with privacy safeguards and disclosure |
| Reformatting existing coded excerpts into a display | No | Ownership | Permitted with disclosure |
| Copy-editing prose the researcher has fully authored | Borderline (word choice can shade meaning) | Interpretive sufficiency (mild), responsibility | Permitted only if the researcher confirms every substantive word choice remains theirs |
| Expanding terse jottings into fuller prose “from memory” | Yes — supplies the connective construal | Interpretive sufficiency, bias, responsibility | Refuse for the primary record |
| Generating an affective/causal reading of an observation | Yes | All four meaning-related fault lines | Refuse |
| Proposing themes or codes from data | Yes | Interpretive sufficiency, bias, responsibility | Refuse for the primary analysis; at most a disclosed, secondary prompt to challenge the researcher’s own coding |
| Drafting a polished participant vignette from raw notes | Yes — selection and framing are interpretive | All four, plus consent | Refuse for anything presented as the participant’s reality |
Three features of this scheme need defending. The first is the borderline status of copy-editing. Word choice can shade meaning — the refusal camp is right that “resent” and “question” are not interchangeable — and so copy-editing is not automatically safe. My position is that copy-editing is permissible only when the researcher retains and confirms authorship of every substantive word, which in practice means treating the model’s suggestions as queries to accept or reject rather than as replacements to wave through. The moment a researcher stops being able to say why each retained word is the right one, copy-editing has crossed into interpretive delegation and the corrupting-fluency danger reappears. The line is real even though it runs through the middle of a single task.
The second is the treatment of expansion, which I place on the refuse side and which is probably the most consequential and most contested placement. Expanding terse jottings into fuller prose is exactly the operation researchers most want to delegate, because it is tedious and time-consuming, and it is exactly the operation the refusal camp’s best argument condemns. When a model expands “N unhappy re: protocol” into a paragraph, it supplies the connective construal — how unhappy, in what register, connected to what — and it supplies it by prediction, biased toward the typical, in fluent prose that the researcher cannot easily audit against a memory the prose is now overwriting. This is the pure case of the corrupting-fluency danger, and I do not think disclosure saves it, because disclosure tells the reader that expansion happened without restoring the interpretive content that was lost. The honest workflow is the slow one: the researcher expands their own jottings from their own memory, while the memory is fresh, in their own words.
The third feature is the qualified opening I leave for AI in coding. I do not treat “propose themes or codes” as flatly identical to “expand jottings,” because there is a genuinely different use available: not letting the model generate the primary coding, but prompting it, after the researcher has coded, to challenge or extend that coding — to ask “what would someone who disagreed with these codes say?” Here the interpretive authorship stays with the researcher; the model functions as a disciplined interlocutor whose suggestions the researcher accepts or rejects on their own analytic grounds. This is close to what the responsible-use frameworks envision at their best, and it is compatible with the accountability principle because the researcher can still answer for the final coding as theirs. But it is a narrow and secondary use, sharply distinct from delegating the primary analysis, and it depends on the researcher being genuinely willing to reject the model’s proposals — which returns us to the skill-atrophy worry, since a researcher whose own coding sensibility has weakened cannot play the adjudicating role the use requires.
From classification to disclosure
The task-sensitive scheme dictates a disclosure practice that is more informative than the blanket statements the frameworks currently produce. The prevailing norm, endorsed by the authorship bodies, is that authors disclose that and how AI was used (International Committee of Medical Journal Editors, n.d.). This is necessary but, for qualitative work, far too coarse: “generative AI was used to assist with writing” tells a reader nothing about whether the interpretive core was preserved or delegated. What the task-sensitive account requires is disclosure keyed to the distinction that matters — a statement of which operations the model performed, sorted by whether it supplied interpretive content. A reader assessing a study’s field notes needs to know not that AI was “used to assist with writing” but that, say, “AI transcribed the researcher’s spoken notes and flagged names for anonymization; all expansion of notes, all coding, and all interpretive framing were done by the named researcher without AI.” That disclosure lets a reader locate the study on the authorship map. The current norm does not.
This disclosure practice also answers, in part, the slippery-slope worry the refusal camp raised. A regime in which researchers must itemize which operations the model performed, sorted by interpretive content, makes interpretive delegation visible rather than absorbable into a vague acknowledgment. The slope is slipperiest when all AI use is disclosed at the same low resolution, so that the transcription and the theme-generation look identical in the record. Raise the resolution and the dangerous uses stand out, which is precisely the condition under which a community can police them. Here the responsible-use camp’s governance instinct and the refusal camp’s diagnostic vigilance actually combine: Bockting et al.’s (2023) call for living guidelines and auditing is exactly the mechanism by which a task-sensitive disclosure norm could be maintained and updated as the tools change, and audited so that the itemization is honest.
Consent and the participant’s stake
One fault line from the ethical map has so far stayed in the background and now needs its own treatment, because field notes are not only about the researcher’s voice — they are about other people’s lives. The ownership-and-rights and privacy concerns that Davison et al. (2024) identify have a participant-facing dimension that the authorship debate, focused on the researcher, tends to underweight. When a person consents to be observed or interviewed, they consent to a particular relationship with a named researcher and, usually, to specified downstream uses. They do not ordinarily consent to having the record of their words and conduct processed by a third-party commercial model whose data handling they cannot inspect. The raw field note, as I noted earlier, is the least anonymized document in the study, and it is exactly the document most likely to be fed to a model if expansion or drafting is delegated. So the interpretive-delegation problem and the consent problem are not separate; they compound. Delegating expansion both hollows the researcher’s voice and routes un-anonymized participant material through infrastructure the participant never agreed to.
This gives a second, independent reason for the placements in Table 2, and it is worth making explicit because it does not depend on the reflexive argument at all. Even a researcher who rejected everything the refusal camp says about interpretive voice would still owe participants a duty not to expose their un-anonymized words to opaque systems without consent. That duty, on its own, forbids uploading raw notes for expansion or vignette-drafting unless the material has been anonymized first — and anonymizing raw notes well enough to make them safe to upload is itself interpretive labor that changes what the notes are. The consent consideration thus reinforces the constrained/interpretive boundary from a different direction: the operations I place on the refuse side are also, and not coincidentally, the operations that most endanger participants. The vignette case is the sharpest. A polished vignette “drafted from raw notes” is presented to readers as a window onto a participant’s reality, and if a model supplied the framing and selection, the reader is being shown the model’s construal of the participant dressed as testimony about the participant. That is a wrong to the reader and to the person depicted at once, and it is why I place it on the refuse side for anything offered as the participant’s reality.
Limits, Objections, and What Remains Unsettled
The position I have defended is contestable at several joints, and honesty requires naming the strongest objections rather than the convenient ones.
The first objection is that my central distinction — interpretive content versus constrained labor — is not as crisp as Table 2 makes it look, and that its apparent crispness does real argumentative work I have not earned. Copy-editing already appears as borderline; but so, on reflection, is transcription, since a transcriber makes decisions about where sentences break and how to render ambiguous speech, and those decisions can shade meaning. If the distinction leaks at the edges, perhaps it cannot bear the weight of a bright-line prohibition. I concede the leakage and deny that it sinks the position. Distinctions that guide practice do not need to be sharp everywhere; they need to identify clear cases at the poles and to flag the middle as requiring judgment, which is exactly what a scheme with a “borderline” row does. The prohibition falls on the clear interpretive cases — expansion, primary coding, vignette-drafting — where no reasonable person doubts that construal is being supplied. That the boundary is fuzzy around transcription does not make expansion any less a delegation of interpretation.
The second objection comes from the responsible-use camp and is more troubling. If the corrupting-fluency danger is as insidious as I have said — if fluent prose really can hide a substituted construal from the researcher who wrote it — then how can I trust the “review and confirm” safeguard that my own borderline category (copy-editing) relies on? Either fluency is disarming, in which case copy-editing is unsafe too, or it is not, in which case expansion might be salvageable by careful review. This is a real tension in my position. My response is that the danger scales with how much interpretive content the model supplies. In copy-editing prose the researcher has already authored, the construal is theirs and the model’s suggestions are local and few, so the review task is bounded and tractable: the researcher can hold each retained word against their own intention. In expansion, the model supplies most of the interpretive content at once, and the review task becomes unbounded — there is no prior authored construal to check the output against, only a memory the fluent output is actively reshaping. The safeguard works where the burden it carries is small and fails where the burden is large. That is not a fully satisfying answer, and I think the boundary between “small enough” and “too large” is an empirical question about human reviewers that the current literature does not resolve. The survey and human-versus-AI comparison studies the field needs on this exact point have not, to my knowledge, been done at the grain that would settle it.
The third objection is that my scheme is culturally and methodologically parochial. It is built around the reflexive, interpretivist understanding of field notes that Jowsey et al. (2025) and the tradition behind them represent, and it privileges a conception of “voice” that not all qualitative researchers share. Some work in more structured, post-positivist qualitative traditions treats coding as closer to reliable classification than to situated interpretation, and for such work the delegation of coding to a model looks less like hollowing out a voice and more like automating a defensible procedure. I accept that my argument is strongest for interpretivist and reflexive work and weaker as one moves toward qualitative traditions that make smaller claims about situated subjectivity. But I would resist the inference that the scheme is therefore merely one school’s preference. Even in structured traditions, the participant-facing consent and privacy considerations hold with full force, and the accountability principle from the authorship bodies (International Committee of Medical Journal Editors, n.d.) applies regardless of methodological school. The interpretive-voice argument narrows as one leaves the reflexive tradition, but it does not vanish, and the infrastructural and consent arguments do not narrow at all.
A fourth objection concerns feasibility and equity. Task-sensitive disclosure asks researchers to itemize their AI operations at a fine grain, and it asks them to run models locally or under contractual guarantees to satisfy the privacy constraint. Both are easier for well-resourced researchers at wealthy institutions than for a lone scholar, a student, or a researcher in an under-funded setting who has access only to consumer AI tools. There is a real risk that a demanding norm becomes a marker of privilege — that scrupulous practice tracks resources rather than virtue. This worries me, and it is one place where the governance proposals of Bockting et al. (2023) are not optional but essential: living guidelines maintained by the community must come with shared infrastructure — vetted, privacy-preserving tools available to all researchers, not only to those who can afford institutional deployments — or the norm will entrench inequality while congratulating itself on rigor. I do not have a full solution, and I flag this as a place where the ethics of AI in qualitative research connects to the political economy of research itself, which is beyond what I can settle here.
The final and deepest uncertainty is temporal. Every judgment in this article is indexed to what current generative models do — their tendency to pull toward the typical, their opacity, their inability to supply a situated construal. Bockting et al. (2023) built their entire proposal around the fact that these systems change faster than static guidance, and that is as true of my scheme as of anyone’s. If future systems could be genuinely local, auditable, and — implausibly but not incoherently — capable of representing a specific researcher’s evolving interpretive stance, some placements in Table 2 might warrant revisiting. I do not expect the interpretive-voice objection to dissolve, because it is grounded in what interpretation is for the reflexive tradition rather than in any contingent limitation of the tools; a model that perfectly mimicked a researcher’s construal would still not be that researcher being accountable as a witness. But I hold the infrastructural placements more provisionally than the interpretive ones, and the right way to maintain a scheme like this is exactly the living, audited, community-governed process that the oversight literature describes rather than a fixed rule I could hand down here. The distinction between authoring an interpretation and delegating it is, I have argued, the stable core; where precisely each operation falls against that distinction is the part that must stay open to revision as both the tools and our understanding of the reviewers who use them improve.
Conclusion
The question I began with — can qualitative research use AI without losing its voice? — turns out to be badly posed, and seeing why is the point. “Voice” is not a stylistic finish that a model might either preserve or spoil; it is the evidentiary trace of a situated interpreter’s attention, the thing that lets a researcher answer for “resent” rather than “question” by saying what they saw and why the word is right. Once voice is understood that way, the yes/no framing collapses. A model that transcribes your spoken notes touches no interpretation; a model that expands your shorthand supplies interpretation you cannot reconstruct and dresses the substitution in fluent prose. These are not two settings on a single dial called “AI use.” They are different acts, and only the second one hollows the record while leaving its surface intact.
What the argument establishes as a whole, and what neither camp sees on its own, is that the refusal position and the responsible-use position are each half right and mismatched: refusal has the correct diagnosis of interpretive delegation — that its corrosion hides in plausibility — but generalizes it into a prohibition its own premises do not support; the frameworks have the correct remedy of disciplined, disclosed, audited practice but cannot say which practices are dangerous, so their discipline lands everywhere and bites nowhere. Task-sensitive authorship joins the refusal camp’s diagnosis to the responsible-use camp’s remedy at the level of the specific operation. The ethical line does not fall between AI and no AI, nor between analysis and writing, but between operations a disclosed, accountable human continues to author and operations in which authorship has migrated to the model. That line is stable even where its application to a particular task — copy-editing, transcription at the margins — is not.
For readers who do qualitative work, the practical consequences are concrete. Disclosure should be itemized by operation and sorted by whether the model supplied interpretive content, because the prevailing “AI was used to assist with writing” tells a reader nothing about whether the interpretive core survived (International Committee of Medical Journal Editors, n.d.). Expansion of jottings, primary coding, and vignette-drafting belong on the refuse side of that line for the primary record; transcription, anonymization flagging, and reformatting are permissible under privacy safeguards; and the one genuinely productive interpretive use — a disclosed, secondary prompt that challenges coding the researcher has already done — depends on a sensibility the researcher must keep exercising. Maintaining such a norm is not a matter of a fixed rule but of the living, audited, community-governed process that Bockting et al. (2023) describe, coupled with shared privacy-preserving infrastructure so that scrupulous practice does not become a privilege of the well-resourced.
How far this reaches is bounded in two ways I want to keep in view. The interpretive-voice argument is strongest for reflexive and interpretivist work and narrows as one moves toward traditions that treat coding as reliable classification — though the consent, privacy, and accountability claims do not narrow at all. And the whole scheme is indexed to what current models do; the placements keyed to opacity and statistical typicality are the ones I hold most provisionally, while the distinction between authoring an interpretation and delegating it I expect to outlast any particular generation of tools.
The problem I cannot settle, and that the field should take up, is empirical: whether human reviewers can in fact catch a substituted construal in fluent prose, and where the burden becomes too large for review to bear. Until we know that, the honest workflow for the primary record remains the slow one — the researcher writing down what they saw, in their own words, while the memory is still theirs to keep.
References
Citation Verification Summary
Bockting, Claudi L., Eva A. van Dis, Robert van Rooij, Willem Zuidema, and Johan Bollen. 2023. “Living Guidelines for Generative AI — Why Scientists Must Oversee Its Use.” Nature 622 (7984): 693–696. https://doi.org/10.1038/d41586-023-03266-1.
Davison, Robert M., et al. 2024. “The Ethics of Using Generative AI for Qualitative Data Analysis.” Information Systems Journal. https://doi.org/10.1111/isj.12504.
Eacersall, Douglas, Lynette Pretorius, Ivan Smirnov, Erika Spray, Sam Illingworth, Ritesh Chugh, Sonia Strydom, Dianne Stratton-Maher, Jonathan Simmons, Isabella Jennings, Robby Roux, Ruth Kamrowski, Abigail Downie, Chee Ling Thong, and Kellie A. Howell. 2024. “Navigating Ethical Challenges in Generative AI-Enhanced Research: The ETHICAL Framework for Responsible Generative AI Use.” arXiv preprint arXiv:2501.09021. https://arxiv.org/pdf/2501.09021.
International Committee of Medical Journal Editors. n.d. “Defining the Role of Authors and Contributors” and “AI Use by Authors.” Accessed from ICMJE Recommendations. https://www.icmje.org/recommendations/browse/roles-and-responsibilities/defining-the-role-of-authors-and-contributors.html.
(Non-scholarly URL reference – not checkable in Crossref/OpenAlex/arXiv; excluded from fabrication accounting. Reference cites a web resource; scholarly indexes (Crossref/OpenAlex/arXiv) cannot verify this type. Excluded from fabrication accounting.; URL appears reachable (HTTP 200))Jowsey, Tanisha, Virginia Braun, Victoria Clarke, Deborah Lupton, and Michelle Fine. 2025. “We Reject the Use of Generative Artificial Intelligence for Reflexive Qualitative Research.” Qualitative Inquiry. https://doi.org/10.1177/10778004251401851.
Nguyen, Dung C., and Catherine Welch. 2025. “Engaged and Responsible Scholarship: Why Qualitative Researchers Should Not Embrace GenAI.” https://doi.org/10.1177/00076503251386539.
Reviews
How to Cite This Review
Replace bracketed placeholders with the reviewer’s name (or “Anonymous”) and the review date.
