How the AgentReadyGo score works
The score is deterministic: identical mission results give an identical score. The model decides what it observes, not what it is graded. This page documents the full calculation.
1. The journey stages
Each mission goes through a subset of stages. The agent declares the result of each one at the moment it clears it.
| Stage | Identifier |
|---|---|
| Site discovery | discovery |
| Product search | search |
| Product understanding | product_understanding |
| Comparison | comparison |
| Variant selection | variant_selection |
| Availability / stock | availability |
| Add to cart | cart |
| Shipping | shipping |
| Returns | returns |
| Checkout | checkout |
Three verdicts are possible, converted into points: pass = 100, warn = 62, fail = 0. A stage outside the mission’s scope is not counted.
2. The dimensions
Stages feed eight dimensions, weighted by their real commercial impact.
| Dimension | Stages aggregated | Weight |
|---|---|---|
| Discoverability Capacité d'un agent à trouver le site, comprendre sa structure et localiser un produit précis. | discovery, search | 14 % |
| Product Understanding Lisibilité machine de la fiche produit : caractéristiques, prix, cohérence avec les données structurées. | product_understanding | 14 % |
| Comparison Possibilité de comparer plusieurs références sur des critères objectifs. | comparison | 8 % |
| Availability Détermination fiable du stock et des variantes réellement achetables. | availability, variant_selection | 12 % |
| Cart Ajout au panier sans ambiguïté et confirmation explicite du contenu. | cart | 14 % |
| Shipping Disponibilité du coût et du délai de livraison avant le checkout. | shipping | 12 % |
| Returns Accessibilité et clarté de la politique de retour. | returns | 8 % |
| Checkout Capacité à atteindre le tunnel de commande et à en comprendre les étapes. | checkout | 18 % |
3. Friction penalties
Each friction removes points from its dimension, capped at 45 points: beyond that, the stage score already reflects the failure.
4. The technical share
Machine-readability checks — robots.txt, sitemap, structured data, internal search — account for 30% of the Discoverability and Product Understanding dimensions, and nothing elsewhere. A technically flawless site on which agents fail gets a bad score, by design.
5. The overall score
A weighted average of the dimensions actually measured, with weights renormalized when a dimension was not tested. Two caps then apply:
- maximum score 58 if no mission reached checkout;
- maximum score 65 if no add-to-cart succeeded.
6. The grades
| 85 – 100 | Excellent | AI agents generally complete the buying journey. |
| 70 – 84 | Solid | The journey completes, with a few significant frictions. |
| 50 – 69 | Risky | Some missions fail or take abnormal effort. |
| 0 – 49 | Critical | Agents hit major blockers. |
7. What the score does not measure
It measures neither human conversion, nor visual design, nor raw technical performance. It measures one thing: whether an autonomous agent can carry a transaction through to the end on your public storefront.