When the Evaluator Hallucinates: The TRAX Protest and the Accountability Void in AI-Assisted Acquisition
Back to Signal
AIDefenseGovernmentComplianceInnovation

When the Evaluator Hallucinates: The TRAX Protest and the Accountability Void in AI-Assisted Acquisition

August 31, 2026Spartan X Corp

The Department of War has spent the last several years building AI into its kill chain: decision-support tools, autonomous weapons coordination, AI-enabled ISR processing. It has invested comparatively little in governing how that same class of AI is used inside the acquisition machine that funds those capabilities. A lawsuit filed in the U.S. Court of Federal Claims in late July is the first significant legal test of that oversight gap, and the facts on record are uncomfortable for anyone who believes AI adoption in government can outpace the accountability frameworks surrounding it.

TRAX International Corporation alleges that the Army's Source Selection Evaluation Board used an AI tool to score proposals for the White Sands Missile Range mission support contract — a five-year award worth approximately $450 million — and that the tool fabricated weaknesses in TRAX's bid that do not appear anywhere in the company's actual submission. The Army has acknowledged that at least one of the weaknesses it assigned to TRAX is unsupported by the record. That concession is material: in federal acquisition, a weakness is not an opinion, it is a documented finding that directly shapes the competitive ranking and therefore the award decision. An AI-generated weakness that no evaluator verified against the proposal text is, functionally, a made-up finding carrying the weight of a government determination. TRAX, which submitted a $420 million bid and lost to a joint venture of Systems Application & Technologies, Amentum, and EWA Warrior Services at approximately $450 million, now argues the evaluation was structurally compromised.

The Army's defense — that the AI tool was "experimental" and did not influence the final award decision — may prove legally sufficient, but it raises its own set of questions that the acquisition community should sit with. If an AI tool is present in a source selection process but experimental, what is the oversight protocol? Who reviewed its outputs against the proposal text before those outputs became official evaluation findings? Was the tool's involvement disclosed in the source selection documentation provided to GAO, and if not, why not? Records provided during the GAO protest phase reportedly did not explain whether specific strengths assigned to the winning offeror originated from a human evaluator or the AI system. The Government Accountability Office denied TRAX's protest in May; the Court of Federal Claims has broader discovery authority and is a different forum. What emerges from discovery could be more detailed than what GAO saw.

The Federal Acquisition Regulation was designed for human evaluators. It specifies what source selection records must contain, how evaluators must be credentialed, and how findings must be traceable to the proposal text. None of that framework contemplates a language model as a participant in the process, because until recently no such participation existed at any significant scale. The Army is not alone: agencies across the federal government are quietly integrating AI tools into acquisition workflows — drafting evaluation narratives, flagging proposal sections, identifying potential weaknesses — without published guidance on disclosure requirements, human review standards, or how AI-assisted findings are to be documented in the official record. The result is a governance structure where the most consequential AI outputs the government produces may be the ones inside selection decisions, not the ones in battlefield systems, and they are the least regulated.

What sound governance looks like is not mysterious. Source selection records should disclose the use of any AI tool and document the specific human review applied to its outputs before they are treated as official findings. Evaluators who rely on AI-generated analyses without independent verification of those analyses against the source material should not be considered to have made a finding at all — they have ratified an algorithm. Protest forums, from GAO through the Court of Federal Claims, need clear standards for what a contracting agency must produce when an AI tool's involvement is alleged, equivalent to what they already require for records of human evaluator deliberations. None of this is technically difficult. It is organizationally inconvenient, because it would require agencies to document and defend something they have largely treated as administrative support.

The TRAX case will not resolve the policy question regardless of its outcome. A court ruling for or against TRAX settles the facts of this protest; it does not create binding guidance for the hundreds of other source selections now in progress where AI tools may be present. That guidance has to come from the Federal Acquisition Regulatory Council, the CDAO, or Congress — and the TRAX lawsuit is the most visible pressure yet to accelerate that work. The DoW has already shown it can move fast when a legal or operational gap is identified. The same institutional capacity that rebuilt the acquisition pathways for software and autonomous systems can write clear AI governance standards for the contracting process itself. The question is whether a lawsuit at one missile range is enough to make that happen before the next hallucination ends up in a far larger award decision.

Share this article
LinkedIn

BUILD WITH US

Ready to Solve Hard Problems?

Spartan X builds AI systems, autonomous platforms, and cybersecurity solutions for defense and national security.