
Time to a Verified Finding: The Right Way to Measure Investigative AI
The urgent work of an investigation is rarely visible from the outside. It could be a list of records, messages in an encrypted app, a set of timestamps or a questionable eyewitness account. What sits on the other side of that important work is a victim who may or may not be identified, a perpetrator who may be causing more harm or a threat network that continues to operate. Time in this work is not a performance metric: it is the distance between harm continuing and harm stopping, whether the finding ends up in a courtroom or in a briefing where a commander must make a careful decision. And this is why I’m careful about how our industry should measure it.
When I look at AI for investigations, the question I care about is how much sooner a person can reach a finding they have checked and can explain. That is a harder test than how quickly a system produces an answer. It is also a much more useful one.
An investigator still has to decide what an answer means, whether the supporting material says what the system claims, and what relevant information might be missing. AI creates value when it reduces the effort required to do that work well. If it moves effort from searching to correcting, the apparent gain may disappear before the work is finished.
The Path to a Verified Finding
By verified, I mean that an accountable reviewer has checked the supporting sources and recorded the relevant limitations. That does not make the conclusion infallible or determine its legal admissibility.
Consider a hypothetical investigation in which an analyst asks whether a device was near a particular location during a specified period. An AI system returns a confident narrative with several source references. The analyst opens them and finds that one timestamp uses a different time zone, two records describe the same event, and a message refers to a place with a similar name. A device identifier also appears under two names. Each issue changes what an investigator can reasonably conclude.
The response might have arrived in seconds. The finding is still unfinished. A reference may point to a real record while failing to support the interpretation attached to it. Even a correctly interpreted device location does not, by itself, establish who was carrying the device. The distinction matters to everyone whose life an investigation touches.
I would measure the whole path from a defined question to a finding reviewed by an accountable person. Start with the permitted evidence and a clear task. Include preparation, processing, source inspection, corrections and the handoff to the next reviewer. Record elapsed time as well as active human effort. A system that makes the first analyst faster can still create more work for the examiner or prosecutor who follows.
Speed Is Only Part of the Evaluation
Speed needs an accompanying account of quality. Which relevant findings did the system identify? Which did it miss? Did it preserve information that challenged the initial theory? Could the reviewer see what failed to process? A fast result that quietly leaves out difficult material can give the team an inaccurate picture of how much work remains.
This means testing difficult cases deliberately. Include ambiguous identities, conflicting timestamps and incomplete records. Include a task for which the available material does not support an answer. Give reviewers a way to record an unresolved question as the correct outcome. Otherwise, the evaluation rewards a system for sounding certain when uncertainty is precisely what it should communicate. Show how often the system completes the task to the required standard, including where it cannot.
The comparison also has to be fair. Use the same question and authorized source material. Make the quality requirements explicit before the evaluation. Count setup and review costs for each approach. Disclose the model and tool configuration and repeat the task enough to understand variation. A carefully chosen demonstration can help someone understand a product, but it cannot establish its performance across investigations.
The Standard Investigative AI Should Meet
For buyers, this creates a practical request: show me a completed, reviewed piece of work and let my team inspect the route to it. For technology providers, including Cellebrite, it creates an obligation to make our claims testable. We should be willing to show where assistance helps, where it introduces extra review, and where a task remains outside the system’s reliable scope.
I am optimistic about what AI can do for people who spend their days working through complex digital information. Reducing repetitive work could give them more time to examine an inconsistency, pursue an unanswered question or explain a finding properly. Those benefits become credible when we measure them through the work itself. The next time you evaluate investigative AI, ask when the answer was generated, and then ask when someone had enough support to stand behind the finding.