AI-Assisted Browser Forensics: From Plausible Stories to Verifiable Evidence

.avif)
.avif)
When an endpoint security tool alerts on communication with a suspicious domain, we often need to know what the user was doing in the browser around that time. History and download records can help answer that question, but finding the right profiles, parsing the files, and building a dependable timeline takes time. At Daylight, we started exploring whether AI could speed up that work.
Speed alone is not enough in incident response. Every conclusion has to be traced back to evidence. An invented step in a timeline can change what we investigate or remediate. An unsupported claim about impact can influence remediation decisions and find its way into an incident report.
Browser history has gaps of its own. A recorded visit does not show everything the user saw, and a nearby browser event does not automatically explain a network alert. Our first attempts at AI-assisted analysis crossed those gaps without telling us where observation ended and inference began. That became the problem we set out to solve.
Why Browser Forensics
Browser forensics gave us a bounded task with useful evidence to work from. It is especially relevant to initial-access cases such as ClickFix, where the pages visited before and after suspicious activity may help explain how it began.
The questions are simple to state: What did the user visit? Was there a redirect or download? How does the browser timeline relate to the alert? Getting dependable answers takes work. An analyst must find the relevant browsers and profiles, retrieve history and download artifacts from an endpoint, interpret different SQLite schemas and timestamp formats, and join events without losing the provenance of each result.

The Failure That Changed the Design
Our first approach gave the analyst an LLM copilot. It could explain Chrome structures, generate SQL, and help frame questions. The analyst still had to correct for missing conditions: another profile, an unexpected schema, an incomplete join, or context the model had dropped. In one iteration it took roughly 14 conversational turns to get a usable query.

A skill reduced some repetitive prompting by packaging instructions for the task. It still depended on the analyst noticing when the model had overlooked a profile, chose the wrong join, or carried an assumption into its summary. We could make the conversation more efficient without making its answers verifiable by design.
The copilot made parts of the work faster, yet it did not give us a repeatable forensic process. More importantly, the conclusions it produced were not inherently tied to the rows and artifacts behind them.
We then gave the model a client’s Chrome history file that was difficult to open manually and asked it to answer investigative questions. It parsed enough to produce a report with a clean timeline and plenty of detail, plausible enough to send for peer review. The reviewer found the neatness suspicious and asked what supported each step. Some details came from the artifact; others were connections the model had invented to fill gaps in the data. That was the breaking point: We could no longer take the model’s account at face value.
The problem was more specific than the familiar observation that LLMs hallucinate. The invented details survived an initial human read because they matched a plausible incident pattern. Under the time pressure of IR, that kind of explanation can reinforce what an analyst already suspects. A stricter prompt might reduce these mistakes, but the workflow still gave the model room to mix evidence with unsupported assumptions. Its claims were not consistently linked to the source records, leaving the reviewer to reconstruct those links manually. We needed to change how evidence reached the model and how its conclusions could be checked.
Evidence First, AI Second
We redesigned the workflow so each conclusion could be traced to evidence and gave the model a narrower job. The result has six stages, color-coded in the graphic by who or what does the work:

1. Collect (agentic and deterministic): An agent finds the browsers, users, and profiles on the endpoint and retrieves the artifacts through EDR tooling.
2. Analyze (deterministic): Code parses the SQLite files, converts timestamps, and joins records using predefined SQL queries.
3. Scope (deterministic): The alert’s time, host, domain, and process decide which records matter. This is a lead-based investigation, not a free-form one.
4. Narrate (agentic): The LLM explains the scoped findings using only the evidence it was given.
5. Enrich (agentic and deterministic): Shortlisted domains are checked against sources such as Google Threat Intelligence and URLScan.
6. Validate (human): An analyst checks the evidence links before any conclusion is accepted.
Collection handles the messy endpoint reality. Browsers, users, profiles, operating systems, and directory paths vary, so an agent can inspect what is present and use EDR tools to retrieve the relevant artifacts. Enrichment can also combine tool-driven exploration with fixed checks. Both stages need observable tool actions and retained outputs.
Browser history is usually stored in a SQLite database file. Our code queries its tables, converts timestamps, joins related records, and normalizes the results. Given the same artifact, these steps should produce the same output. The alert supplies anchors such as a time, host, domain, or process; deterministic queries then scope the records to nearby URLs, redirects, downloads, referring pages, search activity, and direct matches. This reduces the evidence set before the LLM sees it.
The handoff between these steps is as important as the steps themselves. The agent returns collected artifacts and records of its tool actions. The parser produces structured results from those artifacts. Scoping retains the source references for each event rather than passing the model a free-form digest. If the narrator says a URL was visited at a particular time, the reviewer should be able to follow that statement back through the query to the browser record.
The model can now narrate structured findings, explain a possible relationship to the alert, and point out gaps. It cannot turn an unqueried browser database into its own set of facts. A human validates consequential claims against the underlying evidence. The color split in the graphic is deliberate: We use flexible agents where the environment varies, repeatable code where evidence is transformed, and human judgment where conclusions are accepted.
Grounding Has to Survive Peer Review
Instructing the model not to hallucinate is not a control; it is a hope. The output needs enough structure for someone else to check it. Several things do the controlling: enforcing a schema the output must satisfy, requiring fields such as timestamp and source, and creating a link from every important claim to the artifact, query, and row behind it, like footnotes in an academic paper. A claim about a download’s source page, for instance, needs a visible relationship in the data rather than a plausible sequence of URLs.
This also limits what the model is being asked to do. It may connect several supported observations into a tentative explanation, but it needs to label the connection as an inference. It may identify a missing artifact or suggest a follow-up query. Neither action authorizes it to assert that the missing event occurred.
We also distinguish observation from inference. “This URL appears in browser history at this time” reports an artifact. “This visit may explain the network alert” proposes a relationship. The second statement can be useful, provided the analyst can see why it was proposed and what remains unproven.
It helps to think of the model as an eager entry-level analyst: confident, keen to hand over a finding, and ready to present an unproven conclusion as fact if nobody checks. That changes the reviewer’s task. Reading the narrative to decide whether it sounds reasonable would reproduce our original failure. The reviewer checks whether each important assertion is supported, whether a relationship is observed or inferred, and whether the evidence leaves another explanation open. We design the output for that inspection from the start.
When a claim lacks support, we treat it as a product bug. AI output gets the same defect-tracking standard as any other software, not a lower bar because a model produced it. Perhaps the scoped evidence was insufficient, a query or tool result was summarized incorrectly, or the output format let an inference pass as an observation. Each failure gives us a concrete part of the system to improve.
A Useful Answer Can Be: We Don’t Know Yet
One case showed why this boundary matters. An endpoint alert involved a suspicious domain, but the domain itself did not appear clearly in the browser history. A system optimized to deliver a tidy explanation might have treated nearby browsing as proof of the connection.
Our workflow instead surfaced activity near the alert time, identified another domain worth checking, and suggested an external investigation. That follow-up exposed ClickFix-related infrastructure. It did not retroactively prove a browser visit to the original alert domain; it gave the analyst a productive next step while preserving the gap in the evidence.
There was still investigative judgment in deciding which nearby activity warranted a follow-up. The difference was that the system exposed the uncertainty instead of disguising it. An analyst could then check the other domain, examine the result, and decide what the new evidence did or did not establish about the original alert.
This is the kind of assistance we want from an agent. It can recognize that a direct correlation is missing, explain the limit of the available artifact, and suggest a check that could reduce uncertainty. It should not fill the gap with a story.

What This Changes for Agentic IR
Starting with browser history gave us a bounded workflow whose inputs and outputs we could inspect. It also showed where flexibility helps and where it becomes a liability. Agents are valuable for navigating inconsistent environments and pursuing leads. Evidence extraction and reduction need predictable code. Narration becomes more useful once the model has a constrained set of findings to explain, and review becomes meaningful when the supporting records remain accessible.
We also learned that human review is only as strong as the material available to the reviewer. A polished narrative with an “approve” button offers little protection. A timeline with source references, explicit inferences, and a clear record of unresolved questions gives the reviewer something to challenge. That is a more useful measure of progress than how naturally the model writes.
We are testing whether these boundaries hold across other forensic artifacts and investigation workflows. The goal is not to make an AI-generated report indistinguishable from an analyst’s report. It is to make the investigation faster and the resulting claims easier to challenge.
Do not ask AI to be the investigator. Use it to make the investigator faster and harder to fool.







