Compliance

Who Is Accountable When AI Assists an Investigation?

Lewis Smith · · 12 min read

The investigator is accountable. Software can surface an observation; it cannot hold a view, be cross-examined, or answer for a finding about a person. In SentinelOps, an indicator is not a finding: the investigator reviews every indicator and records their own assessment, and that recorded assessment is the conclusion the organisation must defend.

Who Is Accountable When AI Assists an Investigation - the Software or the Investigator?

The question is not rhetorical, and it is not new. It is being asked right now by the institutions that will eventually test the answer.

The Victorian Law Reform Commission tabled Artificial Intelligence in Victoria’s Courts and Tribunals in the Victorian Parliament on 3 February 2026. It is the first inquiry by an Australian law reform body into AI use in courts and tribunals, it makes 30 recommendations, and it draws a hard line at the point of decision: AI tools can support judicial officers, but the report recommends against their use for judicial decision-making itself, because of the risk to judicial independence and to confidence in the administration of justice.

The Australasian Institute of Judicial Administration reached the same place from a different direction. Its guide AI Decision-Making and the Courts: A Guide for Judges, Tribunal Members and Court Administrators, prepared by researchers at the UNSW Faculty of Law and Justice and first published in 2022, examines AI in the courtroom against core judicial values including accountability, procedural fairness and equality before the law.

The Australian Human Rights Commission’s Human Rights and Technology Final Report (2021) addressed the liability question directly, recommending that the law be clarified so that a decision-maker’s legal responsibility for a decision is not diluted by the fact that the decision was AI-informed. Australia’s AI Ethics Principles, published by the Department of Industry, Science and Resources, carry an accountability principle to the same effect: the people responsible for each phase of an AI system’s lifecycle should be identifiable and accountable for its outcomes, and human oversight should be enabled.

And the Report of the Royal Commission into the Robodebt Scheme, presented on 7 July 2023, is the Australian case study nobody in this market needs explained. Among its 57 recommendations it called for a consistent legal framework for automation in government services, with review pathways and transparency mechanisms, and for a body to monitor and audit automated decision-making.

Read together, the direction of Australian institutional thinking is unambiguous. Accountability does not transfer to software. It stays with the person who made the decision, and with the organisation that stands behind them. That is the settled position, and every design choice in an investigation tool should be tested against it.

For a workplace investigation team, this is not an abstract point. When the Fair Work Commission reviews a dismissal under section 387 of the Fair Work Act 2009, it examines the process a person followed and the reasons a person gave. It does not examine a tool. Whatever assisted the investigation, the investigator is the one who has to explain the finding.

Where Does AI Genuinely Help in an Investigation?

Refusing to let AI reach conclusions is not the same as refusing to use it. The useful applications are real, and they are all upstream of the finding.

  • Triage. Deciding what an investigator should look at first, out of a volume nobody can read end to end in the time available.
  • Surfacing what a human might not have looked at. Not deciding what it means. Putting it in front of someone who can decide.
  • Pattern recognition across large volumes. The same subject appearing across unrelated matters, clusters by site or period, sequences that recur. A single complaint looks isolated; twelve over eighteen months in one business unit is a different matter entirely.
  • First-draft drafting. Assembling a structured draft from the case record so the investigator edits rather than starts from nothing.
  • Search and retrieval. Finding the relevant material in a case file without needing to construct a technical query.

Every one of those is a way of making sure a person had the chance to look. None of them is a conclusion. This is the line SentinelOps applies across its AI-assisted investigation capabilities: AI processes information at scale, and every output is presented for investigator review and confirmation rather than applied automatically.

Where Must AI Never Be Trusted in an Investigation?

Three categories, and they are not close calls.

Findings. A finding is a proposition that an organisation asserts and must defend. It has to be reached by someone who can explain how they got there, in a forum that can ask them.

Credibility. Assessing whether a witness is truthful, weighing a hesitation in an interview, deciding what to make of an account that changed between the first and second conversation. This is human judgement built on experience, and treating it otherwise is how AI in investigations becomes a liability.

Any conclusion about a person. Whether the threshold was met. Whether the conduct occurred. Whether the document handed to you during a workplace investigation was altered before it arrived. These are the conclusions that end employment, refuse claims and support regulatory action.

Australian courts have already begun drawing this line in operational terms. The Supreme Court of New South Wales Practice Note SC Gen 23, which commenced on 3 February 2025, prohibits the use of generative AI to generate the content of affidavits, witness statements and character references, and requires prior leave of the court before generative AI is used in preparing an expert report. It also requires practitioners to verify that every citation and authority exists and is accurate - and specifies that the verification cannot itself be done with a generative AI program.

That last provision is the whole principle in one line. A machine output can be checked by a person. It cannot be checked by another machine output.

Why Is a Score the Wrong Output for Evidence?

Start from what evidence has to do, rather than from what any product does.

Evidence has to be defended. Somewhere downstream - a Fair Work Commission hearing, an internal appeal, a regulator’s review, a solicitor’s letter - a named person has to state a conclusion and answer questions about how they reached it. That single requirement imposes four conditions on anything that feeds into the conclusion:

  • It must be assessable. The person has to be able to weigh it, not just receive it.
  • It must be unpackable. If asked why, they need something to open up and explain.
  • It must be attributable. The reasoning has to belong to a person who can own it.
  • It must not pre-empt the decision. It has to leave room for the person to reach a different view.

A number of the shape “87% likely altered” fails all four, and it fails them in ways that are easy to miss.

It presents itself as a conclusion. A percentage is grammatically a finding. It arrives pre-formed, and the investigator’s task silently changes from assessing evidence to agreeing with a figure. Under time pressure, agreement is the path of least resistance.

It is not reproducible reasoning. A score compresses many considerations into one figure, and the figure cannot be decompressed back into the considerations. Asked why, it can only answer 87. Reasoning that cannot be unpacked cannot be reviewed, and reasoning that cannot be reviewed is not evidence of anything.

It cannot be cross-examined. Cross-examination and internal review both operate on propositions a person can defend. Put a score in a report and the questions arrive immediately: what would 86 have meant? What would have had to be different for this to read 60? Is the gap between 87 and 62 the gap between a finding and no finding, and who decided that? The investigator cannot answer, because those answers live inside a system they did not build. The investigator is the person giving evidence - not the vendor, and not the software.

It moves the decision without anyone deciding. This is the most serious failure, and the quietest. When a number does the deciding, accountability has landed on something that cannot carry it. No person formed the view, no person can be asked to justify it, and because the transfer happened without a decision point, there is no moment in the record where anyone took responsibility.

And it can manufacture assurance. A high figure reads as an accusation. A low one reads as a clearance, which is worse, because it invites an investigator to stop looking on the strength of a number that was never capable of clearing anything.

None of that makes scoring a bad engineering choice everywhere. Where an output feeds an automated decision inside a high-volume processing pipeline, a probability and a threshold are exactly the right tools, and no human is being asked to explain the result. That is a different setting with different requirements. An investigation is the other setting, and SentinelOps is built for it: SentinelOps Authenticity Check reports what it observes, and produces no score and no verdict. Where the evidence is not conclusive, the system says less, not more.

A score or confidence percentageAn observation for assessment
What the investigator receivesA conclusion to agree withA proposition to assess
What appears in the case recordA numberA recorded human assessment, in the investigator’s words
Who is accountable for itUnclearThe named investigator
Behaviour under uncertaintyThe number movesThe system says less, not more
What can be cross-examinedNothing that anyone present can explainThe investigator’s stated reasons
What a low value impliesAssurance the evidence may not supportNothing. Absence of indicators is not evidence of authenticity

What Does the Investigator Actually Record Against an Indicator?

This is where the argument stops being a principle and becomes a product fact.

In SentinelOps, an investigator dispositions each indicator as one of exactly three things: “Confirmed concern”, “Possible”, or “Discounted”. There is no “authentic” disposition and no “genuine” disposition. None exists, and none may be added, because no analysis of this kind can establish that a document is genuine.

The vocabulary is doing deliberate work.

  • Confirmed concern records that the investigator, having looked, holds the concern.
  • Possible records genuine uncertainty as uncertainty, rather than rounding it into either an accusation or an all-clear.
  • Discounted records that the investigator has set the concern aside. It is the investigator’s judgement of the analysis - not a statement about the document, and not a finding that anything is authentic.

Each of the three is a human act, recorded in a person’s name, at a point in time. That is the difference between an investigation and a pipeline: at the end of it there is a person who decided, and a record of them deciding.

Corroboration is also required before the strongest characterisation. A single observation cannot escalate on its own.

Why Is “Human in the Loop” an Evidentiary Requirement, Not a Product Feature?

Software vendors list human review as a feature, usually near the bottom of a page. In an investigation it is not a feature. It is the mechanism by which the output becomes evidence at all.

An unadopted machine output is not a finding. It is an artefact sitting in a system. It becomes part of an investigation only when a person examines it, forms a view, and records that view as their own. That act of adoption is what a court, tribunal, commission or internal appeal later examines - not the software’s output, but the human reasoning applied to it.

This is why optional human review is not the same thing as mandatory human confirmation. If review can be skipped, then in the cases where it was skipped there is no reasoning to test, and the organisation is left defending a number.

In SentinelOps, an indicator is not a finding. An investigator reviews every indicator and records their own assessment before anything enters the case record.

One design commitment underpins all of it: in SentinelOps, evaluative assessments are computed by documented rules rather than by a language model, so a model cannot shape an outcome it does not produce. The language model’s role is limited to describing observations in investigator language. This is the single most important admissibility fact about the capability, and it is why the analysis can be re-run and reconciled rather than merely repeated.

Why Can the Absence of an Indicator Never Mean a Document Is Genuine?

Because the two are not opposites.

An indicator is a positive observation about a file. Its absence is not a positive observation about anything. Treating “nothing surfaced” as “this document is genuine” converts silence into a finding, and it is the most dangerous thing a tool of this kind could do - dangerous precisely because it feels like good news, arrives without a decision point, and closes an avenue of inquiry that nobody consciously chose to close.

It would also invert the burden. Under the uniform Evidence Acts, authenticity is proved by the party relying on the document. Section 58 of the Evidence Act 1995 (Cth) allows a court to examine a document and draw any reasonable inference from it, including an inference as to its authenticity or identity. That is a power the court exercises on the whole of the material. It is not a power that transfers to a software output, and no vendor should write copy implying that it does.

SentinelOps Authenticity Check surfaces indicators that a document may have been altered. It does not decide, and it does not accuse. The floor of the scale is “no indicators found”, carrying a standing limitation notice. There is no value above it that says “authentic”, and there never will be, because absence of indicators is not evidence of authenticity.

There is a practical consequence worth acting on. Alteration is frequently visible only as the difference between two files - an original and an altered copy, which in the real world usually arrive separately and at different times. Investigation teams routinely do this comparison by eye, across two screens or two printouts, which is slow and depends on how tired the person is at four o’clock. If a comparison item exists, get it, and put the two side by side. Two items of evidence compared against each other answer a question that neither answers alone.

How Should AI Involvement Be Documented So the Investigation Stays Defensible?

The defensibility of an AI-assisted analysis rests on whether, years later, someone can reconstruct what was run, on what, by whom, and what the human made of it. That is a record-keeping problem, and it is solvable.

What to recordWhy it matters
The integrity fingerprint of the file analysedEstablishes that the item analysed is the item in the report
Which analysis produced the observation, and whenAnchors the observation to a point in time
Tool version, rules version and model identifiers in effect at the timeAnswers “which engine produced this, and can it be reproduced?”
The identity of the person who requested the analysisAttributes the step to a person, not a system
The investigator’s disposition of each indicator, in their wordsThis is the finding. Everything else is input to it
The indicators the investigator discounted, and whyAdverse material addressed rather than ignored is a procedural fairness expectation, not an optional courtesy
Whether the run completed or was partialA degraded run that presents as complete is a defect in the record
That AI assisted, and in what roleDisclosure obligations increasingly attach to this, and forums differ

SentinelOps version-stamps every run - file hash, tool versions, model identifiers, rules version, timestamp and requesting user - so an analysis can always answer which engine produced it and whether it can be reproduced. The record is append-only: completed analyses are never modified, and re-analysis writes a new record, so what an investigator relied on and cited remains exactly as produced. A partial run declares itself in the record rather than presenting as complete. This sits inside the same immutable audit trail that carries the rest of the evidence record.

All analysis is performed in Australia. Evidence never leaves the country.

What Can an Investigator Do Tomorrow?

Practical, and none of it requires new software.

  1. Write your finding in your own words, from the evidence. If a sentence in your report could have been copied from a tool output, rewrite it or justify it.
  2. Record a disposition against every indicator, including the ones you set aside. The discounted ones are what demonstrate that you assessed rather than adopted.
  3. Interrogate any number before you rely on it. Ask the supplier what would have had to be different for the figure to read materially lower. If nobody can tell you, it does not belong in an investigation report.
  4. Never treat “nothing found” as a clearance. Record it as what it is: no indicators surfaced, on this item, on this date, under this version.
  5. Go and get the comparison item. If an original or an earlier version might exist, the difference between the two files is often where the answer is - and comparing them side by side beats comparing them by eye.
  6. Note the tool and version in the case record at the time of the analysis, not at the time of the report. Versions change; memories do not improve.
  7. Check the rules of your forum before you file. Practice notes on generative AI now differ between Australian jurisdictions, and the restrictions bite hardest on witness statements and expert reports.
  8. Keep the roles separate in serious matters. The person who ran the analysis, the person who assessed it and the person who decides the outcome should not collapse into one, for the same reason procedural fairness requires the investigator and the decision-maker to be different people.
  9. Know where your evidence is being processed. If you cannot answer that question about a tool you are using, you cannot answer it for the person whose document you uploaded.

Frequently Asked Questions

Who is accountable if an AI-assisted analysis turns out to be wrong?

The investigator and the organisation. Australian institutional thinking is consistent on this point: the Australian Human Rights Commission recommended that a decision-maker’s legal responsibility should not be diluted by the fact a decision was AI-informed, and Australia’s AI Ethics Principles require that responsible people be identifiable and accountable for AI outcomes. A tool that surfaces observations for human assessment keeps that accountability where it already sits. A tool that issues verdicts obscures it without removing it.

Does SentinelOps decide whether a document is genuine?

No. SentinelOps Authenticity Check surfaces indicators that a document may have been altered. It does not decide, and it does not accuse. The conclusion is the investigator’s. The software’s role is to make sure they had the chance to look.

Does SentinelOps produce a confidence score or a percentage?

No. SentinelOps reports what it observes and produces no score and no verdict, because a score invites adoption as a conclusion and cannot be cross-examined by the person who has to defend it.

What does an investigator record against an indicator in SentinelOps?

One of three dispositions: “Confirmed concern”, “Possible”, or “Discounted”. There is no “authentic” or “genuine” disposition, and none can be added. “Discounted” means the investigator has set the concern aside - their judgement of the analysis, not a statement that the document is genuine.

Does the AI produce the assessment?

No. In SentinelOps, evaluative assessments are computed by documented rules rather than by a language model, so a model cannot shape an outcome it does not produce. The language model’s role is limited to describing observations in investigator language.

Is “no indicators found” the same as a document being authentic?

No, and the distinction is critical. The system will never tell you a document is genuine. Absence of indicators is not evidence of authenticity, and treating it as such would be the most dangerous thing this tool could do.

Do I have to disclose that AI was used in an investigation?

Increasingly, yes, and the requirements differ by forum. The Supreme Court of New South Wales Practice Note SC Gen 23, in force since 3 February 2025, restricts generative AI use in affidavits, witness statements and expert reports and imposes verification duties on practitioners. Assume disclosure will be expected, record AI involvement in the audit trail at the time, and check the practice notes or procedural rules of the forum before you file.

Does SentinelOps provide an expert witness?

No. The investigator is the witness. SentinelOps produces a version-stamped, append-only record of what was analysed, by whom, under which engine version, and how the investigator dispositioned it - so that the person giving evidence can speak to their own reasoning with the record behind them.

Where is the analysis performed?

In Australia. All analysis is performed in Australia and evidence never leaves the country. Customer data is stored in Australia, and the platform is built, developed and tested in Australia.


Related reading: AI-assisted investigation · Workplace investigations · Evidence management · Audit trails · Procedural fairness in Australian workplace investigations

Your Next Investigation Deserves Better

See how SentinelOps transforms investigation management in a 30-minute investigator-led walkthrough. No sales pitch. Just the platform, your questions, and straight answers.

Currently serving Australian enterprise, government, and regulated industry organisations.