Methodology · Test planning

Which tasks should AI complete?

A customer wants to compare a product, request a quote or track a delivery. Task selection defines which of these journeys an AI agent should test, before measurement starts. It specifies the goal, access and permitted endpoint. Multiple tasks reveal different barriers; a single test remains a limited observation. This shows where Hyperize can improve the customer journey.

Fix the task before the test. Read the result within its tested scope.

The starting point

Start with the customer need.

One failed journey does not describe an entire brand. The task must reflect a real customer need.

The methodology separates selection from measurement: define the goal and limits, then test the journey.

A single website check covers one slice. A broad brand comparison needs multiple comparable tasks.

The rules

Six rules for fair tests

Rules for the test plan, not proof that every historical measurement meets every criterion [S1].

01

Test a real customer need.

02

Test access within the brand’s responsibility.

03

Use multiple tasks for broad comparisons.

04

Keep difficulty comparable within a sector.

05

Separate intermediary routing from technical barriers.

06

Freeze and publish task descriptions before the comparison wave.

Common errors

Avoid five errors

These errors distort a comparison [S1].

FM1
Generalizing one case
One journey represents a whole brand.
Use multiple task classes.
FM2
Unequal comparison
Tasks differ in difficulty.
Use a shared sector grid.
FM3
Misreading mediation
A handoff counts as technical failure.
Clarify responsibility first.
FM4
Task drift
Old and new tests get mixed.
Report versions separately.
FM5
Testing only easy tasks
Measurability displaces customer need.
Require relevance and testability.
Direct or via intermediaries

Test the appropriate endpoint

Direct surface

The journey at the brand

What can the agent complete on the brand’s own pages and systems?

Intermediaries

The journey via third parties

Does the task require an intermediary, or could the agent have stayed with the brand?

Intermediaries provide context, not an extra factor in the published score formula [S2].

Applied

Four task-selection examples

Planning examples, not new brand measurements. Verify the goal and available access for each task.

Insurance

Allianz

Access route · Quote and advice

A suitable inquiry can be the endpoint.

Suitable task

Check coverage; prepare an inquiry.

An unfair requirement

Count an intermediary result as website failure.

Controls

FM1 · FM3

Automotive

Mercedes-Benz

Access route · Configuration and contact

Define the endpoint for each offer.

Suitable task

Check configuration; find a test drive.

An unfair requirement

Require online purchase for every journey.

Controls

FM2 · FM3

Logistics

DHL

Access route · Shipping and tracking

Specify the parcel task and access.

Suitable task

Calculate a price; track a shipment.

An unfair requirement

Treat a price lookup as equivalent to logged-in checkout.

Controls

FM2

Healthcare

Bayer

Access route · Information and supply

Separate product information from purchase.

Suitable task

Find documentation and a supply route.

An unfair requirement

Assume direct manufacturer checkout.

Controls

FM1 · FM3

Comparability

Freeze tasks. Explain changes.

For comparison waves: freeze and publish task descriptions before testing. Retests add to the history [S1].

A changed task gets a new version. Compare trends only under comparable conditions.

Findings can be challenged with evidence. Corrections need a traceable change record.

Scope and limits

What the test establishes.

Rules and score weights are public. Internal prompts, scoring details and raw logs stay within the project. A planned stop before purchase or submission is not a technical failure. Examples do not establish current brand performance; rankings and citations are not guaranteed.

Sources

Evidence and provenance

S1

Internal

Task selection rules

Hyperize · Methodology · May 2026

Internal methodology, not publicly retrievable.

Supports: Six rules, five failure modes and versioned task descriptions.

S2

Methodology

Agent Success Score

Hyperize · October 2026

https://www.hyperize.ai/en/methodology/agent-success-score

Supports: Public 20/70/10 weights and measurement limits.

S3

Evidence

DHL: historical task example

Hyperize · DAX 40 Index · Q2 2026

https://www.hyperize.ai/en/dax40-index/brands/dhl

Supports: One parcel task; checkout behind login untested. Not a complete task mix.

S4

Evidence

DAX 40 Index

Hyperize · Published measurement periods

https://www.hyperize.ai/en/dax40-index

Supports: Results within their stated scope and limits.

Internal methodology is not publicly retrievable. Public examples document their own test scope.

Last updated
Next review
Confidence A
Evidence tierProprietary methodology

From testing to implementation

Foundation

Methodology

How we make your customer journey usable by AI.

Understand the method

Measurement

Agent Success Score

What the score summarizes and where its limits lie.

Understand the score

My next step

My customer task

Agree the goal, access and scope for your project.

Discuss my customer task