Methodology · Concept

Agent Success Score.

The Agent Success Score measures how well a brand lets an AI agent complete its human's task. One score per agent lane, computed from two measured gates: AI Visibility (does the agent find and recommend the brand) and AI Usability (once the agent arrives, does the surface let it act). Scored by real agents against frozen tasks, not by a checklist.

Readiness rates potential. Success counts outcomes.

What it is not

A brand metric, not a support metric.

The Agent Success Score is a brand-side metric: it measures how well a company's own surfaces let an incoming AI agent finish a task. It is not the contact-center agent success rate, the support metric that counts how often service agents resolve a customer's issue. Same words, different subject, different agent.

One rates a brand's readiness for the AI agents arriving at its surfaces. The other rates a support team's resolution rate. This page is about the first.

Why this exists

An outcome,
not a checklist.

A readiness check rates what a surface could support: files present, markup valid, boxes ticked. None of that proves an agent gets the job done. The Agent Success Score starts at the other end. Real agents run a real task against the live surface, and the score records whether the brand let them succeed.

The subject is the brand, the agent is the visitor. When a human asks an assistant for a better insurance and the agent cannot identify the insurer or cannot start the switch, that is not the agent failing. That is the brand declining a customer it never saw [S2].

How it is built

Two gates. Both must open.

A brand can be recommended and still un-buyable, or perfectly buyable and never found. The score keeps the two failures apart, because they need different fixes.

Gate 1 · AI Visibility

Does an agent find and recommend the brand for the task?

Measured by a query audit across the AI platforms buyers actually use, on unbranded task questions. A brand that never enters the answer never gets a visitor. This is the classic GEO half, and it is only half.

Gate 2 · AI Usability

Once the agent arrives, does the surface let it act?

Measured by sending an agent fleet against the live surface: from a plain reader to an autonomous operator. The score is derived from the observed per-agent-class access profile and how deep the close goes. The profile is the truth; the number is a reproducible summary of it, never a hand rating [S1].

Evidence · the confidence layer

Is the result provable and consistent across methods?

Not a third sales axis. Evidence weights how consistently independent methods returned the same ground truth. It keeps a lucky single run from reading as a measured capability.

Composition · published

Agent Success Score = (AI Visibility × 0.20) + (AI Usability × 0.70) + (Evidence × 0.10)

The composition is public. How each gate's number is derived from individual runs stays inside the engagement. A live example with the score beside its decomposition: DHL Group, AI Visibility 46.0, derived AI Usability 68, Evidence 85, composite 6.5 of 10 [S3].

Why trust it

Two independent instruments, one finding.

The two gates are measured by two separate rigs that share no machinery. Gate 1 is a query audit across AI platforms. Gate 2 is an agent fleet running the task on the live surface. When both point at the same gap, the finding does not depend on either instrument being right alone. And because every task is frozen and published before a wave runs, the goalposts cannot move between measurements [S2].

The lane model

One score per lane. Never blended.

A brand meets six kinds of agent, each on a different surface with a different job. A shopping agent and a press agent need different things; averaging them would hide the signal [S4].

Commerce buy
Talent apply
After-sales resolve
Procurement source
Investor research
Press research

Each measured lane carries its own Agent Success Score, and every brand record shows a coverage meter for how many of the six lanes are measured. The DAX 40 Index measures the Commerce lane first, across all qualifying brands, before any second lane opens.

On scope

The composition, not the derivation.

This page publishes the metric: what the Agent Success Score measures, the two gates, the composition weights, and the lane doctrine. How individual runs are scored into each gate, how queries are classified, and the task grids behind a wave are proprietary. The metric belongs in the open. The derivation stays inside the engagement.

Questions

What people ask.

What is the Agent Success Score?

The Agent Success Score measures how well a brand lets an AI agent complete its human's task. It is one score per agent lane, built from two measured gates, AI Visibility and AI Usability, and scored by real agents running frozen tasks against the live surface, not by a checklist.

Is the Agent Success Score the same as a contact-center agent success rate?

No. The Agent Success Score is a brand-side metric: how well a company's own surfaces let an incoming AI agent finish a task. A contact-center agent success rate is a support metric: how often service agents resolve a customer's issue. Same words, different subject and different agent.

How is the Agent Success Score calculated?

It is a weighted composite of two measured gates and a confidence layer: AI Visibility (0.20), AI Usability (0.70), and Evidence (0.10), scored per agent lane on a 0 to 100 scale and displayed 0 to 10. AI Usability is derived from the observed per-agent-class access profile, never hand-rated. How individual runs are scored into each gate stays inside the engagement.

Sources

Evidence and provenance.

S1

internal

Agent Success Score methodology (ars-methodology/v1.1) + usability-derivation/v1

Hyperize Internal — Methodology · May 2026

Hyperize HQ/knowledge/strategy.md

Supports: The two-gate model, the derived (never hand-rated) AI Usability rule, and the published composition weights.

S2

internal

Task Selection Doctrine

Hyperize — Methodology · May 2026

https://www.hyperize.ai/en/methodology/task-selection

Supports: Why tasks are frozen and published before a wave runs, so the goalposts cannot move between measurements.

S3

measured

DHL Group BrandScore (Commerce lane)

Hyperize — DAX 40 Agent Success Index · Q2 2026

https://www.hyperize.ai/en/dax40-index/brands/dhl

Supports: A live example of the score beside its decomposition: AI Visibility 46.0, derived AI Usability 68, Evidence 85, composite 6.5 of 10.

S4

measured

DAX 40 Agent Success Index (living dataset)

Hyperize — DAX 40 Agent Success Index · Q2 2026

https://www.hyperize.ai/en/dax40-index

Supports: The lane scaffold and coverage meter on every brand record: scores are per lane, lanes are never blended.

Internal sources reside in the Hyperize project repository. Measured examples link to their DAX 40 Index pages, where the per-brand outcome is published in full, always beside its decomposition.

Last updated
Next review
Evidence tier proprietary
Confidence A