METHODOLOGY·v1.0 · May 2026

How we score every AI token from 0 to 100

A composite score of real AI utility, code activity, token structure, team credibility, and narrative momentum — weighted to separate protocols that build from those that market.

The AI-token sector has a signal problem. Any protocol can publish a roadmap, hire a KOL, and list on a mid-tier exchange. The tokens that are doing actual AI work — routing inference jobs, generating revenue from AI customers, shipping production code — often look identical to the ones that are not, at the surface level of price and market cap.

The AI Score is our answer to that problem. It is a 0-to-100 composite score, calculated across five pillars, that attempts to measure what a protocol is actually doing rather than what it is saying it will do. It does not predict price. It measures the quality of the underlying protocol at a point in time. Higher scores mean more of the measurable evidence points to real AI function. Lower scores mean the opposite.

The methodology below is the complete specification. Every input is sourced from a public API or on-chain data feed. Every weight reflects a judgment call — and we explain those judgment calls here.

Pillar 1 of 5

25% weight

AI Utility

On-chain inference, protocol revenue from AI customers, and active deployments — the closest proxy we have to the question that matters: is someone paying to use this?

What we measure

  • On-chain inference and compute transactions (Dune Analytics, chain explorers)
  • Protocol revenue from AI customers, net of token emissions (DeFiLlama)
  • Active deployments: agents, miners, subnets (Taostats, Flipside)
  • Unique paying users over a 30-day window
  • Whether the token is structurally required for the AI function — or merely adjacent

Why we weight it this way

AI Utility is the highest-weighted pillar because it is the hardest to fake at scale. A team can buy social coverage, seed GitHub stars, and structure a token schedule to look clean. It is substantially harder to fabricate sustained on-chain revenue from actual AI customers, or to maintain consistent inference transaction volume without real demand.

Failure mode we mitigate

Some protocols route payment on-chain but run the actual AI workload on centralized infrastructure (AWS, Azure). On-chain payment volume then looks like AI utility when it is not. We mitigate this by requiring verifiable provider diversity — cross-referencing registered hardware against implied GPU-hours — and by flagging protocols where the ratio looks anomalous.

Pillar 2 of 5

20% weight

Team & Execution

Doxx quality, AI/ML credentials, prior shipped products, and active hiring — credible signals of a team that can deliver, not just announce.

What we measure

  • Doxx quality: verifiable LinkedIn and GitHub identities for key team members
  • AI/ML credentials: prior employment at recognized AI labs, publications on Google Scholar or arXiv
  • Prior shipped products: protocols launched and maintained, not just announced
  • Active ML hiring activity over the prior 90 days (LinkedIn jobs)
  • Peer-reviewed research output attributable to the team

Why we weight it this way

Team is 20% because the protocol layer in AI is early. Many projects are pre-revenue but have credible teams with verifiable track records building toward real products. Dismissing team quality entirely would penalize protocols making legitimate early-stage progress. Overweighting it would let a star-studded founding team paper over absent utility — which is why AI Utility outweighs it by 5 points.

Failure mode we mitigate

Anonymous teams are not automatically penalized to zero, but they receive a material doxx deduction. A team can earn it back through consistent code output and verifiable on-chain delivery — anonymity is a risk factor, not a permanent disqualifier.

Pillar 3 of 5

20% weight

Code Activity

GitHub commits, contributor count, issue ratios, and mainnet deployments — sustained development is necessary for legitimacy, but not sufficient on its own.

What we measure

  • Commit frequency over a trailing 90-day window (GitHub API)
  • Unique contributor count over the same window
  • Open-to-closed issue ratio (a proxy for maintenance responsiveness)
  • Mainnet deployment events: are new contracts shipping to production?
  • Electric Capital developer rank within the broader crypto developer ecosystem

Why we weight it this way

Code activity at 20% reflects that sustained development is necessary but not sufficient. A team can commit-farm trivial changes to inflate metrics. We apply commit quality analysis — weighting commits by lines changed and commit message substance — to reduce the gaming surface.

Failure mode we mitigate

Mature, stable protocols (Numerai, Golem) can appear low-activity because their core code is complete. We apply a maturity adjustment: protocols with four or more years of uninterrupted mainnet operation receive a +10 adjustment to their code activity sub-score to prevent penalizing stability. New protocols do not receive this adjustment.

Pillar 4 of 5

20% weight

Token Model

FDV-to-MCap, emission curves, unlock cliffs, supply concentration, and liquidity quality — a strong protocol with broken tokenomics is still a bad investment.

What we measure

  • FDV-to-market-cap ratio (circulating supply as a percentage of total)
  • Annual emission rate and emission curve shape
  • Supply unlock events in the next 90, 180, and 365 days — weighted by proximity
  • Supply concentration: top 10 non-exchange, non-protocol wallets
  • Token utility strength: required for protocol operation, or governance-only?
  • Liquidity quality: slippage at $100K and $1M trade size, venue concentration

Why we weight it this way

Token Model at 20% reflects that a technically strong protocol with a catastrophic token structure — 95% FDV overhang, heavy insider concentration, imminent cliff unlocks — represents real investment risk that the AI Utility score does not capture. These risks are observable, quantifiable, and priced incorrectly by markets with surprising frequency.

Failure mode we mitigate

Token structure inputs are the most gameable pillar over short windows — teams can extend lock periods right before a scoring event. We calculate unlock risk on a rolling basis and flag anomalous changes to lock schedules within the prior 30 days.

Pillar 5 of 5

15% weight

Narrative Momentum

Google Trends, GitHub stars, quality-filtered social coverage, premium research, and exchange listings — capped at 15% so narrative alone cannot save a low-utility token.

What we measure

  • Google Trends trajectory over 90 days for the protocol name and ticker
  • GitHub star growth over 90 days
  • Quality-filtered forum and social mentions: Hacker News, Kaito-weighted social coverage
  • Premium research desk coverage: Messari, Delphi Digital, independent analyst reports
  • Exchange listing trajectory: new CEX or DEX listings over the prior 90 days

Why we weight it this way

Narrative is real — protocols with strong developer mindshare and research coverage attract talent and capital. But narrative is also the most easily purchased input. Narrative Momentum is hard-capped at 15%. A protocol cannot score above 55/100 on narrative alone, no matter how strong its social coverage. The score exists to separate builders from marketers; the cap enforces that separation mechanically.

Failure mode we mitigate

Coordinated KOL campaigns can spike any social metric in a 30-day window. We flag protocols where Narrative Momentum jumps more than 20 points in 30 days and publish a companion '4-pillar score excluding narrative' metric alongside the primary AI Score during flagged periods.

How we update scores

  • Daily: Price-derived inputs, narrative momentum signals (Google Trends, GitHub stars, social coverage), and on-chain data via streaming APIs.
  • Weekly: Code activity inputs (commit frequency, contributor count, issue ratio). Weekly cadence filters out single-day spikes.
  • Monthly: Tokenomics structure inputs (FDV ratio, emission schedule, unlock cliff calendar). These move slowly; daily updates would create false precision.

Where the score appears

  • This page — the methodology specification you are reading now.
  • Every token detail page — a breakdown card showing each pillar's sub-score, weight, and contribution. Singularity API subscribers can pull the breakdown via the /v1/score endpoint.
  • The token leaderboard — the circular ring in the AI Score column, color-coded by tier: emerald (85–100), blue (70–84), amber (55–69), red (below 55).

How to challenge a score

Scores are wrong sometimes. Data sources have gaps. On-chain activity is misclassified. A team doxes between scoring cycles.

To submit a dispute: email scores@aitokens.app with the token name, the specific pillar you believe is scored incorrectly, and any on-chain or public evidence supporting the correction. We review disputes within five business days. If the correction changes the score by more than 5 points, we publish a changelog entry on the token's detail page.

The methodology itself updates quarterly. Scores recalculate automatically; historical scores are preserved with the methodology version that produced them.

See an example breakdown

Bittensor (TAO) scores 92/100 — Excellent, top 8% of tracked protocols.

The breakdown card on the TAO detail page shows exactly how each pillar contributes.

View TAO AI Score breakdown →