Skip to content

· By Ricardo Torres Oliva

The researcher's and the student's challenge in the AI era

Where a language model contributes to a piece of research and where it contaminates it without a trace, and the minimum log that makes assisted work defensible.

Abstract

Artificial intelligence entered academic research before the criterion for governing it did. The result is an asymmetry: a method is required to be declared, replicable and auditable; a model gets one sentence at the end of the manuscript. Researchers and undergraduate and postgraduate students who already work with language models have to answer for what they produce before a committee, a board of examiners or an editor, and today most of them cannot. This document separates where the model contributes from where it contaminates, provides a matrix of uses with their epistemic risk and what must be disclosed, and sets out the minimum five-field log that makes assisted work defensible. The rule that orders everything else: the disclosure is written while the work is being done, not once it is finished.

A method is defended: it is explained, replicated and audited. A model, in most of the work that already uses one, is not defended: it is mentioned. This document is about that asymmetry. Where artificial intelligence improves a piece of research, where it contaminates it without leaving a trace, and what one has to be able to declare.

1. The problem

The tool arrived before the criterion. That is nobody's moral failing: it is a historical sequence. But the result is that research is now under way in which the most influential instrument in the project is the only one that passes through no control at all.

A method is required to be declared, to be replicable and to withstand an audit. A sample is required to have a size, a provenance and a known bias. A model is required, at best, to have a sentence at the end of the manuscript.

The dominant failure mode in academic work assisted by artificial intelligence has been identified and has a name: uncritical acceptance of the model's outputs [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era]. It is not that the researcher believes everything the model says. It is that they check what sounds odd and let through what sounds right, and a language model is optimised precisely so that everything sounds right.

There is a second asymmetry, and it is more uncomfortable. A sample can be measured again. A piece of code can be run again. A conversation with a model in a version that has since been withdrawn cannot be reconstructed. Irreproducibility is not an accident of careless use: it is a default property of the instrument. Whoever does not deliberately counteract it inherits it.

The real divide of 2026 is not between those who use artificial intelligence and those who do not. It is between those who consume it and those who understand how it produces its results [source: AI Literacy A Multi-Dimensional Analysis...]. A researcher in the first group is not badly equipped: they are defenceless when the questions come.

2. Why the usual approaches fail

The binary policy. Most institutions have responded with permission rules: what is allowed, what is forbidden, in which courses. A permission rule answers "may I?". The question that decides whether the work stands is a different one: "how is this claim supported?". No permission policy produces traceability.

The detectors. Detecting synthetic text is an arms race with two guaranteed outcomes: false positives that harm honest students and false negatives that acquit the skilful. A detector, even one that works well, proves origin, not rigour. A paragraph written entirely by a person can be indefensible, and an assisted paragraph can be impeccable.

The tools course. Usage literacy — which model is best, how to write a prompt — goes out of date in months. The capability that does not expire is metacognitive: looking at an output and knowing when to trust, when to doubt, when to verify and when to discard [source: The Phoenix Doctrine v1.1]. That is not learnt in a three-hour workshop.

The disclosure written at the end. That is product-oriented reporting. The standard now taking hold is the opposite: process-oriented documentation, which captures how the researcher critically evaluated the alternatives the model generated [source: AI Literacy A Multi-Dimensional Analysis...]. A disclosure drafted once the work is finished is a reconstruction from memory. A reconstruction is not evidence.

The idea that verifying means reading carefully. It does not. Verifying is a procedure with a criterion defined in advance and a result expressed as a rate, not as a verdict. That is the difference between believing something works and knowing how often it works [source: Evals Ingeniería de Confiabilidad y Evaluación de Sistemas Agénticos].

3. Where it contributes and where it contaminates

There is a division of labour that does withstand scrutiny. The machine drives the expansive phase — generating variation, widening the space of alternatives — and the human drives the contractive phase: eliminating what cannot be sustained [source: AI Literacy A Multi-Dimensional Analysis...].

From that comes a one-line operating rule:

*Artificial intelligence is legitimate where its output will be filtered by an independent criterion. It contaminates where its output is the criterion.*

Looking for twenty rival hypotheses is expansion: all twenty will later go through the experimental design. Deciding which of the twenty is the right one is contraction, and there the machine has no authority. Drafting a discussion from results that have already been verified is expansion of form. Generating the discussion and accepting its inferences is delegated contraction.

That rule intersects with a second distinction, the one worth keeping in mind during daily work: the model can operate as an instrument, as a partner or as a proxy [source: AI Literacy A Multi-Dimensional Analysis...].

  • As an instrument, it performs a bounded task with an external criterion of correctness. It is calibrated, like any instrument.
  • As a partner, it supplies alternatives that the researcher tests. It is contrasted, and the contrast is documented.
  • As a proxy, it replaces a cognitive task that the researcher does not perform. The failure mode has its own name: offshoring of the cognitive task [source: AI Literacy A Multi-Dimensional Analysis...].

Use as a proxy is not forbidden by any law of nature. But there are only two honest destinations for it: to declare it precisely or to withdraw it. What is not defensible is that it happens and does not appear.

4. Matrix of artificial intelligence uses in research

This table can be used as it stands, without hiring anyone. The right-hand column is the one worth having written down before starting, not afterwards.

UseEpistemic riskWhat must be disclosed
Literature search and screening (assisted systematic review)Automation bias and gaps in the training data: what the model did not see does not exist for the projectTool and version, query strings, inclusion and exclusion criteria, and how many records were discarded without human review
Data extraction and construction of evidence tablesSilent reading errors; the error is not randomly distributed, it concentrates in the worst-written studiesProportion of records verified manually and the error rate found in that sample
Synthesis and drafting of sectionsFluency that masks the absence of support; plausible references that do not existWhich sections were assisted and to what degree; that the verification of each reference was human
Hypothesis generation and exploration of ideasUnquantifiable uncertainty: not distinguishing a possible future from an invention of the machine [source: AI Literacy A Multi-Dimensional Analysis...]The decision point: which alternatives were discarded and by what criterion
Analysis codeCorrect result by an incorrect path; technical debt that nobody auditsThe complete code and the tests that verify it, not the result
Synthetic data or imputationGenerated material treated as a primary sourceExplicit use, method, and a visible separation between observed data and generated data
Translation and copy-editingLow: it does not touch the inference. Beware of shifts in nuance in technical termsMinimal mention, according to the editor's rules
Assisted peer reviewConfidentiality of someone else's manuscript; self-preference of the model that judges [source: Evals Ingeniería de Confiabilidad...]Usually restricted or prohibited: check the editor's policy before uploading anything

5. Traceability: the unit to keep is not the prompt

Saving the prompts is an archive, not a defence. The unit that makes a piece of work defensible is the decision point: the moment when the researcher had a model output in front of them and decided something about it [source: AI Literacy A Multi-Dimensional Analysis...].

A minimum usage log has five fields and fits in a spreadsheet:

  1. Date and tool with the exact version. Without a version, no reproducibility is possible.
  2. The assignment. What was asked, in what context, with what material uploaded.
  3. The raw output. Before editing. It is the only thing that later allows a comparison.
  4. The decision and its criterion. Accepted, corrected, discarded — and why.
  5. The verification. Who checked what, against which source, with what result.

That log serves three different purposes, which is why it is worth keeping even if nobody asks for it: it allows an error to be reconstructed when it surfaces, it allows one to learn from the pattern of one's own failures, and it serves as evidence of diligence if anyone asks [source: Evals Ingeniería de Confiabilidad...].

Three rules make it work:

Whoever judges is not whoever built. It is the doctrinal rule that governs the verification of non-deterministic systems [source: Evals Ingeniería de Confiabilidad...], and it translates effortlessly to academic work: the assisted section is reviewed by someone who did not write it, with the explicit brief of looking for unsupported claims. Self-reviewing one's own assisted text gathers two biases in the same person.

Verification produces a rate, not a verdict. Checking twenty references out of a total of one hundred and twenty and finding three that do not exist is not an isolated incident: it is a fifteen per cent error rate that projects onto the rest. That figure has to be calculated, written down, and it has to be decided in advance which value forces the work to be redone.

Context degrades. A model's accuracy falls as the conversation fills with redundant history and previous corrections [source: AI Literacy A Multi-Dimensional Analysis...]. Long sessions are the ones that generate the most confidence and deserve it least. New task, new session, material uploaded again.

6. The disclosure: committee, board of examiners, editor

They are three audiences with three different questions. A single generic disclosure answers none of the three well.

The ethics committee asks where the data went. Uploading interviews, clinical records or material from human subjects to a third-party service is a data transfer, whether or not that was the intention. What must be declared is what left the institution, to which provider, under what retention conditions and whether the material was used for training.

The board or thesis panel asks which part of the reasoning is the author's. Here the international norm is already fixed and it is categorical: a model cannot appear as an author, because authorship implies taking responsibility for the integrity of the work and a model cannot take it [source: AI Literacy A Multi-Dimensional Analysis...]. The practical consequence is not one of form: it means that every assisted output has a named human responsible for it, including the ones the author did not review.

The editor asks where, how and with which version. The standards of the main publishing bodies and of the international committees of editors establish that use is declared in the methods section, specifying where it was used, how, and which version of which tool; that generated material is not treated as a primary source; and that the use of artificial intelligence in the handling or analysis of data is declared separately [source: AI Literacy A Multi-Dimensional Analysis...].

To that is added, for anyone publishing or disseminating in the European Union, the transparency obligation on synthetic content that comes into force in August 2026 [source: AI Literacy A Multi-Dimensional Analysis...].

The rule that binds it all together is one of sequence, not of content: the disclosure is written while the work is being done. Written at the end, it is a reconstruction. And if halfway through the project the disclosure paragraph cannot be drafted, the problem is not one of drafting: it is that the work is not traced.

7. What to do on Monday

  1. Open, in the current project, a usage log with the five fields from section 5. Start with Monday's session, not with the backlog: the backlog cannot be reconstructed and the attempt eats the week.
  2. Classify every current use as instrument, partner or proxy. Those that turn out to be proxy have two destinations: to be declared precisely or to be withdrawn.
  3. Take a sample of twenty references produced or screened with the model's help and verify them by hand, one by one, against the source. Note the error rate. That figure is the most important piece of data in the project this week.
  4. Set down in writing, before going on, what error rate would force the section to be redone. A threshold decided after the result is not a threshold.
  5. Write the disclosure paragraph now, with the work half done, and keep it with the manuscript.
  6. Ask a colleague who did not write the assisted section to review it with a concrete brief: locate claims without a verifiable source.
  7. Read the current policy of the editor one intends to submit to and that of the institution. Note the two biggest differences and settle on the stricter one.

What we have not covered here

This document does not cover discipline-specific regimes: clinical trials, personal data of human subjects, material subject to industrial confidentiality agreements or classification. In those fields the applicable rule prevails over any general criterion, this one included.

Nor does it compare specific tools or recommend any. The landscape changes in months, and a document that named products would age before it became useful. The distinction between models run locally and cloud services — relevant for sensitive data — is left out for the same reason.

It does not contain the regulations of each country or of each university in the region. It contains the criterion with which to read them.

And it does not replace the work of applying all of this to a real project, with its log, its measured error rate and its written disclosure. That work, done on the participant's own project, is Phoenix RETx.

Further reading

Back to the index