The Agentic Trust Index
AI agents are being given real authority faster than anyone - outside the companies building them - can verify whether they deserve it, and what harms may result if they don’t.
The Agentic Trust Index is our answer: An independent methodology, developed jointly between GIE Foundation and KAMO Tune AI, for assessing what an agent can actually be shown to do.
-
The Agentic Trust Index is an independent assessment index of whether an AI agent or agentic tool can be trusted with the authority it's been given - built on evidence, not on what it is marketed as doing.
-
Every assessment is built on one “instrument”: An identified subject, holding bounded authority, took a checked action, on a record that survives a skeptic. We evaluate ten domains against that standard - identity, authority, delegation, secrets, policy, isolation, provenance, supply chain, observability, and resilience - and trace one impactful or consequential action through all ten, end to end, down to revoking the actor's access and confirming the action becomes impossible.
-
The authority being delegated to AI agents does not stay confined to the enterprises deploying them. An agent booking travel, moving money, or acting on an account is doing so on behalf of an actual person, and the safeguards that used to sit between a mistake and its consequences, a person who reviewed the request, a human who signed off before funds moved, were built for a pace of decision-making that agentic systems have already outrun, so those safeguards are being removed rather than replaced with anything equivalent. When a delegation chain carries no real limit, or a financial mandate is vague about what counts as one authorized action, the person who discovers it is rarely the vendor; it is the customer whose account moved past what they agreed to, with no one able to fully reconstruct after the fact who was accountable. As agents take on more of the everyday actions and transactions that used to require a person's direct hand, this becomes a question that touches anyone who has trusted a system with something that mattered to them, and could potentially have major ramifications for society as a whole.
The Methodology
Every action an AI agent takes should be able to answer four questions, and we treat an assessment as failed if any one of them can't be answered. We built the Index on to follow a single “instrument”:
Who or what actually did this, proven, not just claimed. Under whose authority, and what exactly were they allowed to do. Was that authority checked before the action happened, or did it just go through. And afterward, can it be shown to someone with no reason to take our word for it, a regulator, an auditor, a customer who got burned, in a form they have no basis to think was altered.
An assessment fails if any one of those four clauses breaks. We evaluate ten domains, set out in full in The Trust Architect's Handbook.
Identity - who or what acted, proven cryptographically, not asserted by the workload itself.
Authority - what the actor is permitted to do, on whose behalf, and within what limit.
Delegation - whether authority narrows every time it passes to another hand, or whether the last link in the chain ends up holding everything.
Secrets - what the actor can reach, with what credential, and for how long.
Policy - whether the stated limit is enforced in code that runs before the action, not just written into a prompt.
Isolation - what the blast radius looks like once something upstream fails.
Provenance - where the model, the data, and anything retrieved into context actually came from.
Supply chain - what the system inherited from packages, base images, and vendors nobody in the room built.
Observability - whether the record of what an actor did could survive being handed to a regulator.
Resilience - whether there's a chosen response when a control fails, and whether the kill switch has ever actually been pulled.
The ATI traces one real action end to end. We call this the Assessment Walk: Ten steps, one action, from identity through to revocation.
Much of what currently passes for AI assurance is a system telling an assessor what it would do under a hypothetical, rather than an assessor watching what actually happens when the system is put under one; a vendor's own documentation can describe a kill switch that has never been built, or a policy limit that exists in a slide deck but not in the code that runs before a payment clears.
The first real test of whether an agent's authority was actually bounded ends up taking place during the incident it was supposed to prevent, with the people affected finding out the limit was never enforced at the exact moment it needed to be. The Assessment Walk exists so that question gets answered before an agent is handling something that matters, not after.
These systems are being embedded into the infrastructure that runs public life: Hospital scheduling and triage support, benefits eligibility, court case management, power grid operations, election logistics, all aspects of finance, and the routine administration of government itself. An agent operating inside any one of those systems without a real, tested limit on what it can do is not only a product risk contained to the company that built it; it is a potentially unaligned and uncontained rogue actor inside a public institution, and the public has no way to see it, audit it, or vote on it.
A society that hands this much operational authority to systems nobody outside a handful of companies can verify is relocating a portion of its own governance to whoever built the model, without ever deciding to do so, and it will not find out where that decision was made until something inside one of those institutions goes wrong - potentially catastrophically.