Prototype · AI-assisted engineering · Developer-reported local results
The failure
HOMININI explores how automated workflows can produce inspectable reports. During development, its verifier accepted appended statements including “all tests passed” and “no security vulnerabilities” without supporting evidence.
A report containing source references could still overstate what the workflow had established. The problem was in the verification path.
The repair
I repaired the verifier and added targeted negative tests for those unsupported additions. The reported regression results show that the tested additions are now rejected.
What this establishes
The repair supports a narrow conclusion: the verification path rejects the specific unsupported additions covered by those tests. It does not establish that every false claim will be detected.
Related prototype work
Repository → private report
One isolated run processed the public psf/requests repository through four workflow steps and generated a private report in 3.469 seconds. It made no model calls and no changes to the external repository. This is one observed run, not a general performance benchmark.
Traceable source inspection
Six Python files from a pinned repository commit were inspected, their content hashes checked and source references included. The downloaded code was not executed.
Complete findings in the report
A formatting limit that removed later findings was reproduced and corrected. A six-file regression test checks that every fixture’s source-linked finding survives report generation.
Artifact visibility
In an isolated application-level test, the submitting user could see the generated artifact and an unrelated user could not. Deployed multi-user security has not been established.
A separate diagnostic repair
An earlier repair explained a recorded execution limit without restarting work. Its developer-reported 56-test suite covered the gateway and diagnosis helper. Those tests are separate from the report-verifier repair described above.
Research direction
Runtime Accountability for Persistent LLM Multi Agent Systems is a working paper by Kiran Yenaganti proposing requirements for human authority, evidence, recovery and independent verification. It is not peer reviewed and does not claim empirical validation. A public reading link is not available on this site yet.
That working paper is distinct from the runtime-governance manuscript shared on ResearchGate.
These isolated demonstrations do not establish production readiness, deployed multi-user security, independent-process verification or fully autonomous engineering. The timing above is a single observation, not a comparative benchmark.