A score is only as good asthe math behind it.
Every Pulse Labs metric is reproducible, versioned and traceable to raw model responses. Here is exactly how we measure what AI thinks about your brand.
Prompt battery design
Each scan dispatches a structured battery of prompts spanning neutral discovery, comparison, recommendation and adversarial framings — across brand themes and audience personas. Prompts are versioned so scores stay comparable over time.
Multi-model dispatch
Prompts run in parallel against the major generative engines — GPT-5, Claude Sonnet 4.6, Gemini 3 and more — with identical phrasing and zero brand priming. Web-grounded variants capture how models behave with live retrieval.
Forensic scoring
Every raw response is scored for sentiment, brand visibility, framing, competitor substitution and factual accuracy. Claims are audited against verifiable sources to flag hallucinations, and red flags are cross-corroborated across models.
Consensus & NCI computation
Per-model scores roll up into a cross-model consensus matrix with divergence weighting. The Narrative Consensus Index (NCI) combines sentiment, visibility and narrative control into one benchmarkable 0–100 figure.
What every number means
How favorably models describe the brand, weighted by framing strength and the prominence of positive vs. negative claims.
How often the brand appears unprompted in category, comparison and recommendation queries — the AI equivalent of share of shelf.
How closely the model's framing matches the brand's intended narrative themes, vs. competitor or third-party framing.
How much models disagree with each other. High divergence means your AI reputation is unstable across engines.
The Narrative Consensus Index — the composite headline score combining the dimensions above with divergence weighting.
The rules we never break
Immutable records
Scan results are write-once. Scores are never retroactively edited — methodology changes only apply forward.
Full provenance
Every score links back to the raw model responses that produced it. Nothing is a black box.
No brand priming
Prompts never hint at a desired answer. We measure what models actually say, not what brands hope they say.
Versioned methodology
Scoring logic carries a version tag on every record, so longitudinal comparisons are always apples-to-apples.
Honest limitations
LLM outputs are stochastic. We report consensus across repeated, multi-model sampling — never a single response.
Versioned, so your trends stay honest
- v2.4Jun 2026
Added truth-audit hallucination forensics, counterfactual crisis simulation and theme-ownership displacement analysis to the scan pipeline.
- v2.3Apr 2026
Introduced persona-resonance profiling and per-model cognitive lens analysis. Expanded model coverage to 8 engines.
- v2.2Feb 2026
NCI composite formula refined with divergence weighting. Theme intelligence expanded to 5-theme batteries.
- v2.0Nov 2025
Cross-model consensus matrix and immutable scan provenance chains introduced.
Want the full technical methodology paper?
We share the complete scoring specification, prompt battery structure and validation studies with prospective enterprise customers.