Raghav Maini
All writing

Evidence, authority, and the cost of another question

What I put around the agent to keep evidence, authority, and research spending consistent as an assessment changes.

An expensive theory

A personal project started with a question I couldn’t quite shake: could rebuilding a salvage car ever make financial sense for me? I like exotics, I don’t have the budget for one, and watching someone else rebuild a wreck is an excellent way to develop an expensive theory about what you could do yourself.

I haven’t bought a car at auction. Before getting anywhere near that point, I wanted to see which assumptions made a rebuild plausible, what evidence was missing, and how much I could justify bidding given my circumstances. A shop’s access to equipment and specialist labor doesn’t become mine because an analysis uses its repair costs.

The useful unit of work became a durable assessment: the vehicle, buyer constraints, evidence, investigations, reviews, spending allowance, and decision history. An agent session works on that record. When the session ends, another can continue with the same constraints and the money already committed.

I can inspect what changed between two decisions and which evidence supported each one. The application also has somewhere to enforce those relationships, without relying on a model to reconstruct them from a long conversation.

A valid quote can stop supporting the bid

Consider a repair quote. The source is authentic, the price correctly extracted, and the owner’s review reasonable. Then an inspection uncovers additional damage. Nothing about the original document has changed, but its approval may no longer cover the job.

My implementation records the damage scope when an inspection or quote is reviewed. If that scope changes, the application reopens the relevant requirements and withholds the supported bid ceiling. The old evidence and decision remain in history, so I can explain both yesterday’s recommendation and why I can’t stand behind it today.

One document, two different decision states
When the quote was reviewed
The approved quote covers the recorded repair scope. It can support the assessment.
After additional damage is recorded
The quote remains evidence. Its review must be revisited before it can support the current bid.

The implementation uses a fingerprint of the recorded damage scope. It is deliberately conservative: some changes may prompt another review even when a particular quote still applies. A more selective rule would reduce that burden, but I’d want tests showing which changes it can safely ignore.

The same problem appears in purchasing. A supplier quote can remain valid for an earlier quantity or delivery schedule after the order changes. Maintaining the relationship between the document, its approval, and the current decision takes explicit application behavior.

The budget has to survive an interruption

The coordinator receives a compact packet: the assessment’s revision, buyer constraints, open requirements, recent evidence, remaining allowance, and available investigations. Fuller history and source material stay in durable storage. A research worker gets a bounded question; an identity check can use an ordinary API adapter.

Before an investigation starts, the application checks that the choice is still available against the revision the coordinator read, verifies the allowance, and saves a reservation. A plan made against an older record cannot authorize spending against a changed one.

That accounting becomes consequential when a provider may have performed paid work but the application never received a usable result. The system conservatively consumes the reserved allowance for uncertain interrupted work. A new session doesn’t reset the budget, and a late result cannot restart an assessment the owner stopped. These allowances control execution; they aren’t measured provider invoices.

Give the agent a useful decision to make

A repair-research worker can return a claim with a citation. That alone doesn’t make the claim a captured source observation or put its number into the cost ledger. Photo analysis can identify visible damage without establishing physical repairability. Review belongs to the authenticated owner, and even the owner cannot turn model inference into documentary verification.

There is still substantial room for agent judgment in choosing what deserves investigation. If qualified market evidence shows that the observed auction bid is already too high under the optimistic repair scenario, more paid research into those repairs cannot rescue the purchase under the current assumptions. The application removes that work from the available actions. Elsewhere it uses a simpler heuristic: larger repair-cost ranges get earlier attention.

I don’t yet have evidence that a model selects better research than that policy. A useful comparison would give both the same actions, evidence environment, and allowance, then measure progress toward a justified decision alongside spending and unnecessary work. Fluent explanations would not settle it.

The tests answer different questions

Synthetic decision cases challenge the business rules. Runtime evaluations exercise the actual transport, tools, saved evidence, revision changes, cancellation, and recovery. I inspect the resulting records because an agent can describe successful recovery without leaving the application in the right state.

Forecast accuracy needs a separate outcome loop. An observed repair cost or sale result must be matched to an earlier decision revision. Reporting counts each assessment once per statistic, preserves relevant exclusions, and compares net sale proceeds with a forecast that also accounts for selling costs. Recording one lot three times should never look like three successful predictions.

Below five observations, the interface shows the sample count instead of a statistic. That restrains the presentation; five observations still don’t establish confidence. The reports don’t retrain the system or change its assumptions.

I haven’t demonstrated profitable rebuilds with this project. What I have built is a way to preserve the evidence needed to assess its recommendations honestly. I can change the coordinator, research adapters, or context policy and still examine whether the decision is supported, whether its approvals remain applicable, and whether another investigation was worth paying for.