Skip to content

How Mavzena measures

Every number in Mavzena comes from a recorded audit run: which prompts were asked, to which engines, how many times, and how each answer was read. This page sets out that method, including which numbers are estimates and what the method cannot tell you.

Prompts written for your category

A prompt is the question put to the engines. Mavzena generates prompts from your brand's category, in the brand's language, in four patterns: recommendation, alternative, comparison and use case. You can also write your own or import them from a CSV file.

Where a prompt comes from

Generated
Written by an AI model from your category. Useful for coverage, but not evidence of what customers actually ask.
Added by you
Questions you know your customers ask.
Imported
Brought in from a CSV file and kept apart, because the rows may not have been reviewed one by one.

Each prompt keeps its source permanently, so a generated prompt is never presented as real demand. Prompts can be tied to buyer personas, which are labelled by origin too: generated, added by you, or derived from listening signals. Groups such as branded, competitor and unaided are worked out again each time from the prompt text and your current competitor list, so they stay correct when that list changes.

Commercial intent (buying, comparing options, going to a specific site, looking for information) is assigned by fixed keyword rules. It is a classification, not a measurement, and a value you set by hand is never overwritten.

Six engines, and how each one answers

An audit sends the same prompts to ChatGPT, Gemini, Claude, Perplexity, DeepSeek and Grok. Engines do not all answer the same way: some answer from what the model already knows, others search the web first. The two can give different answers to the same question, so the mode is recorded with every answer and shown next to each engine in reports.

How each engine answers in a standard audit and in web search mode
EngineStandard auditWeb search mode
ChatGPTModel knowledge onlyLive web search
GeminiModel knowledge onlyNot connected yet
ClaudeModel knowledge onlyNot connected yet
PerplexityLive web searchLive web search
DeepSeekModel knowledge onlyNot offered by its API
GrokModel knowledge onlyNot connected yet

The Free Audit asks 15 prompts on ChatGPT only. Paid plans ask 60 prompts on all six engines. Web search mode is available on paid audits only and costs more per answer. The model behind each engine is recorded with the run.

Reading each answer

A second model reads each raw answer and returns a structured verdict: is your brand named, at what position, which competitors appear, and is the tone positive, neutral or negative.

If that verdict cannot be parsed, the answer is not recorded as “not mentioned”. It is left out of every metric and tried again at the next audit. Counting an unreadable answer as absent would quietly lower your score.

The step that parses these verdicts is locked by a fixed set of example verdicts, tested offline.

The visibility score

Three measurements make up the score. Mention rate is the share of scored answers that name your brand. Average rank is your position in the answers where you appear; lower is better. Sentiment is the tone of those same answers.

Visibility score, 0-100

mention rate × rank effect × sentiment weight × 100

Rank effect
1 − (average rank − 1) × 0.1, never below 0.1First place counts in full; each place lower takes off a tenth.
Sentiment weight
positive 1.0 · neutral 0.7 · negative 0.3Averaged over the answers that name you.

The same formula is applied to each engine separately, and to unaided prompts: the ones that name no brand at all. That shows whether engines recommend you without being asked about you.

A brand that is never named scores 0, because that was measured. A brand that was never measured shows no score.

Repeats, stability and confidence

Engines can answer the same question differently from one moment to the next. So an audit can ask each prompt once, twice or three times on each engine. Time and cost multiply with the count; the default is once.

Stability

Compares the repeats of the same prompt on the same engine: was your brand named consistently, how much did its rank move, did the tone hold. A part that cannot be computed is left out, not filled in. With a single sample, stability is shown as not measurable, never as 100.

Confidence level

A label rather than a number, based on two things together: how many repeats were run, and what share of engine requests succeeded.

Low
One repeat, or too many failed requests.
Medium
At least two repeats, and at least 70% of requests succeeded.
High
Five repeats, and at least 90% of requests succeeded. Five repeats are not offered yet, so today an audit reaches Medium at most.

This is why most results read Low. An audit that asks each prompt once is a single reading, and the label says so instead of implying certainty.

Every number traces back to a run

Each audit opens a numbered audit run. It records the prompts, engines, models, answer mode, sample count, start and finish time, and how many requests succeeded or failed. Every answer is stored against that run. A new audit adds new rows and never overwrites earlier ones, so a trend is drawn only from audits that really happened.

  • If an engine fails, the run is marked partial and reports state how many requests failed. The figures then cover only the answers that came back.
  • A brand can have only one audit in progress. A double click or a retried job cannot measure the same thing twice under two run numbers.
  • Not measured is never 0. A value with no measurement behind it is shown as not measured, and a report that cannot be linked to a run says so.
  • Hallucination verdicts are append-only at database level: edits and deletions are rejected. The workspace history is only ever added to.

Sources the engines cite

In web search mode, Mavzena records the pages an answer cites. For ChatGPT it also records the pages the search read but did not cite. Links are stripped of tracking parameters, grouped by site, and sorted by type (review, forum, social, encyclopedia, app store, news, other) using a fixed list of domains rather than a model.

Whether a source is your own site or a third party is worked out when you view it, from your current website address, so changing your site leaves no stale labels. No source is attributed to a competitor by guessing from names. Answers given from model knowledge alone have no sources to record.

Crawler access and hallucination checks

Can AI crawlers read your site?

A rule-based check with no AI model involved. It looks at whether robots.txt blocks GPTBot, ClaudeBot, PerplexityBot or Google-Extended, whether sitemap.xml and llms.txt exist, and whether your pages carry JSON-LD structured data. Only signals with evidence behind them affect the score; llms.txt is reported as experimental and does not lower it. Pages are read as the server sends them, so content that appears only after JavaScript runs is not seen.

Is AI telling the truth about you?

You confirm the facts about your brand in a Truth Profile. A hallucination scan takes the answers from one completed audit and checks each claim against those verified facts, marking it supported, contradicted or uncertain. The scan fixes which audit and which version of the Truth Profile it used. Your own verdicts are added to a history and never overwrite it.

Measurements and estimates are kept apart

Some numbers are measured from answers. Others are estimates built from written rules. Every score on the Overview carries one of three labels: measurement, estimate, or includes an estimate.

Organic Visibilitymeasurement
The visibility score described above.
Paidestimate
How much sense paid AI placement makes for your measured prompts: commercial intent × the organic gap × the reach of AI ad channels in your country. An unconfirmed channel counts as neither available nor absent.
Combined Opportunityestimate
For each prompt, the larger of the organic priority and the paid score, averaged. High means a lot is still on the table.
AI Growthincludes an estimate
70% × organic visibility + 30% × (100 − combined opportunity). Hidden while organic visibility is unmeasured.

The weights are rough and written down, to be calibrated with real data. Opportunity priority (impact × ease × commercial intent) is a ranking rule, not a forecast.

Mavzena does not publish ads or spend budget. Paid AI produces estimates and campaign briefs; launching anything needs a person's approval.

Web listening is measured separately

Public Voice follows what people write about you and your competitors on Turkish sources (Şikayetvar, Ekşi Sözlük, DonanımHaber, Technopat), on Reddit, in Google Maps ratings and in the RSS feeds you add. It is a different measurement from AI visibility and is never blended into it.

  • Each source reports its own status. “Nothing found” and “could not check” are different results and are shown differently.
  • Share of conversation is computed only over sources that worked for every brand being compared, so a competitor without a Şikayetvar page cannot make you look louder.
  • Tone and topics are assigned by rules. An AI classifier exists, but stays off until it has been measured against hand-labelled posts.
  • Ekşi Sözlük text is never sent to an AI model, as the site's robots.txt asks. Collected posts are kept with their authors for 90 days, then deleted.

What this method cannot tell you

Every measurement has edges. This one cannot tell you:

  • How an engine will answer tomorrow. A few repeats show today's spread, not a guarantee.
  • What other model versions would say. When a provider updates a model, the same prompt can get a different answer.
  • Anything about prompts that were not asked. A score covers the prompts and engines in its run, nothing more.
  • How often real people ask a prompt. Generated prompts describe your category; they are not search volume.
  • Exactly what a person sees in the consumer app. Engines are called through their APIs; the apps can add personalisation, location and ads, and ads shown there are not visible to this measurement.
  • Revenue. Being more visible in AI answers is not proof of more sales, and Mavzena does not convert visibility into money.
  • An accuracy rate. The judge that reads answers and the listening rules have not been measured against hand-labelled data, so no accuracy figure is claimed.

Questions about the method?

Write to us. A person on the team reads every message and replies by email.