Field studies

An AI visibility score for a brand that does not exist yet

Nine runs through HubSpot's free AI Search Grader, using a brand whose page had been live for under a week, with zero public mentions. The score moved when the typing did.

Kevin PantanellaFounder, Syndral
Published August 6, 202611 min read
Two identical brand entries in a scoring interface showing two different scores, 7 and 32, on a cream background.

Key takeaways

  • The grader's results page is built entirely from four form fields passed as URL parameters. Load that address directly and it produces a full scored report, with no form and no click.
  • Identical inputs returned identical results across two days, down to the written analysis. The tool serves a stored answer rather than measuring again.
  • Changing one self-declared word moved the Perplexity score by 8 points and a Gemini sentiment total by 13. The company, the country and the brand name never changed.
  • The reports state precise figures with nothing behind them: exactly 42 brand mentions in four different runs, a LinkedIn engagement score for a brand with no LinkedIn page, and an office address belonging to another company.
  • AI visibility can be measured, but it takes repeated prompts in the real chat interfaces, in sessions carrying no history of their own, reported as a frequency with its spread. A free instant score cannot do that, whoever builds it.

Perplexity scored my brand's AI visibility at 7 out of 100. Fifty minutes later, the same AI visibility score for the same brand was 32. Between those two runs, Syndral published nothing, earned no mention, gained no review. The only thing that changed was what I typed into two form fields.

Here is what makes those numbers worth your time. On the day of that test, Syndral was a domain name bought in March 2026, connected on 18 July to a single coming-soon page and a self-assessment form. No launch post, no LinkedIn page, no press, no clients, not even a company registration of its own: Syndral is the new AI visibility brand of Suzaku Productions, the Bangkok web agency I have run since 2014. The brand being scored had, in every public sense, not happened yet. It happens now: the article you are reading is its launch.

I am Kevin Pantanella, and Syndral is my AI visibility practice, built around GEO (generative engine optimization) and AEO (answer engine optimization).

That public silence is what makes this study possible. Any score a tool gives Syndral is a score for nothing. So I put HubSpot's free AI Search Grader through a nine-run protocol over three days, one variable at a time, every input logged, every screen captured. Here is what the score turned out to measure.

The result before the method

Same brand, same country, same tool, 50 minutes apart. I changed the two free-text fields describing what Syndral sells and what industry it is in. Perplexity's overall score went from 7 to 32. Its brand sentiment went from 0 out of 40, with a written note that no public data exists for this entity, to 18 out of 40. The market score doubled across all three engines at once.

21 July 2026Test 1: "None" / "Agency"Test 2: "Ai Visibility Agency" / "Service Agency"
OpenAI3336
Perplexity732
Gemini3040
Market score (all three)3/106/10
Perplexity brand sentiment0/4018/40

Two readings were possible at that point. Either the score depends on what the customer declares, which makes it a mirror rather than a measurement. Or it is run-to-run noise, which does not make it a measurement either. I designed the protocol to separate the two. The answer surprised me: it is not even noise.

Note: this is not a story about HubSpot doing something uniquely wrong. HubSpot builds serious products and, to its credit, the grader's own footnotes name the models it calls and its disclaimer states the results are AI-generated and unreviewed. The question this study asks is what happens when anyone, even a company with those resources, compresses a handful of LLM calls into a single brand score. If they end up here, the direction itself deserves a closer look.

A brand with zero public existence is the perfect test subject

The usual objection to testing a grader is garbage in, garbage out: feed it a fake brand and you get fake results, so what did you prove. Syndral closes that door. It is a real commercial brand, operated by my registered company, with a live page, and its public footprint on test day was verifiably zero. No posts, no reviews, no press, no directory listings, no social accounts. The input was true. It just described nothing the public internet knows about.

That gives me ground truth most testers never have. When a report asserts that this brand enjoys high engagement on LinkedIn, I do not have to estimate how plausible that is. I know it is invented, because I own the LinkedIn silence it claims to have measured.

The protocol: nine runs, one variable at a time

I ran nine graded tests between 21 and 23 July 2026, with inputs recorded to the character. The shape:

  1. Noise baseline. The same inputs as Test 2, submitted three more times over two days. This measures pure run-to-run variance before anything else is claimed.
  2. One variable at a time. Change the capitalization only. Then the industry word only. Never two fields at once.
  3. Replication. Resubmit Test 1's exact inputs two days later and compare.
  4. The URL test. Load a results URL for an input combination never submitted through the form, and see what happens.
RunDateWhat changed
Test 121 Julthe original: product "None", industry "Agency"
Test 221 Julboth free-text fields rewritten
Baseline x323 Julnothing at all, Test 2 resubmitted three times
Capitalisation23 Julone letter, "Ai" to "AI"
Industry word23 Julone word, "Service Agency" to "Professional Services"
Test 1 replay23 Julback to Test 1's exact inputs, two days later
URL only23 Julno form at all, a results address loaded directly

My two original runs were made on 21 July, before this protocol existed. Both were reproduced exactly under it two days later, and it is those reproductions, timestamped and screenshotted, that the evidence pack documents. The full pack, every screenshot, the exact timestamps and the four official PDF reports, is published alongside this article at the evidence annex. The tool is free and public. Every run in this study can be reproduced in minutes, and I would genuinely rather you check than believe me.

What the grader does when you click

The three identical baseline submissions returned identical results. Not similar. Identical to each other, to the sub-score, down to the written analysis, and identical on every figure I had recorded for Test 2 two days earlier. Capitalising the product field changed nothing either: the lookup is case-normalised. The replay of Test 1 reproduced its 7 out of 100 exactly, including the 0/40 sentiment.

Meanwhile the results page's address spells out the mechanism: the four form fields ride along as URL parameters. So I loaded a results address for a combination never submitted through the form. The page displayed "Grading your brand" with a progress bar and produced a complete scored report. No form, no button. The grade is a URL.

No form, no button. The grade is a URL.

New combinations visibly call the engines: when I changed the industry word, the OpenAI column answered first while Perplexity and Gemini sat on loading spinners for half a minute. Known combinations render instantly. This behavior is consistent with results being generated once per input combination and then served from a cache, though I have not seen HubSpot's backend and do not claim to know its internals. What I can say from the outside is narrower and worse: identical inputs never produced a fresh measurement, and the score only ever moved when my typing did.

Per the grader's own footnotes, the OpenAI column runs on GPT-5.4 mini with an August 2025 knowledge cutoff, and the Gemini column on Gemini 3 Flash Preview with a January 2025 cutoff. Both cutoffs predate Syndral's existence. Whatever those two models reported about this brand, they did not learn it from the world.

One more thing changed between the two dates, and by rights it should have mattered. On the evening of the first two runs, syndral.co was returning 403 to the AI crawlers, GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot, because of a legacy Cloudflare rule I had never deliberately set. Which is an embarrassing thing for an AI visibility brand to discover about itself. I switched it off that same night, roughly forty hours before I replayed Test 1. So on the replay, Perplexity's crawler could reach a site that had been refused at the door two days earlier. The score came back 7, sentiment 0 out of 40, unchanged to the point. I cannot tell you whether Perplexity re-crawled the site in those forty hours, and I am not going to pretend otherwise. What I can tell you is that the one thing under my control that could plausibly move a retrieval-based score did change, and the score did not.

The numbers that follow your typing

Once the industry field became the single variable, the pattern was hard to miss. Everything below changed while the company, the country and the brand name stayed fixed.

One field, three answers

Self-declared industry, everything else held fixed.

Perplexity overall and Gemini sentiment total, out of 100, by self-declared industry. Professional Services: Perplexity overall 28, Gemini sentiment 78 out of 100. Service Agency: Perplexity overall 32, Gemini sentiment 65 out of 100. Marketing Agency: Perplexity overall 36, Gemini sentiment 68 out of 100.

Figure 1. Perplexity overall and Gemini sentiment total, by self-declared industry. Bars are drawn on a 0 to 100 scale. The company, the country and the brand name never changed. The third wording, “Marketing Agency”, comes from the URL-only run, never submitted through the form. The word travels as a URL parameter either way, which is what that run demonstrates.

An 8-point spread on Perplexity and a 13-point spread on Gemini's sentiment, from the one word the customer chooses to describe themselves. The brand archetype flipped with the product field too: type "Ai Visibility Agency" and all three engines call Syndral an Innovator, type "None" and all three call it a Traditionalist. Same brand. My typing has a brand archetype, apparently.

The report's detail pages moved the same way. With one industry wording, Gemini flagged its own analysis "Reliable Data: No". One word later it declared "Reliable Data: Yes", sourced from professional networking sites and regional tech news, and scored Syndral's LinkedIn presence 75 out of 100 for high engagement and well-received AI thought leadership. Syndral has never had a LinkedIn page. There is no engagement. There are no testimonials, no directory listings, no coverage. Every one of those sources was named by a system that had no data and said so one run earlier.

The numbers that never move

Some figures did the opposite: they refused to change at all. The OpenAI column reported exactly 42 brand mentions in all four generated reports, across four different input sets describing four differently framed companies. Gemini reported 1,250 mentions twice and 450 twice. Numbers that precise, that stable, across inputs that different, are not observations. They behave like filler constants in a template.

The numbers that never move

OpenAI, brand mentions reported: 42, 42, 42, 42. Four generated reports, four differently framed companies. Gemini, brand mentions reported: 1,250, 1,250, 450, 450. Two values, four reports. Nothing in between.

Figure 2. Stability where variation was expected. Every other figure in the report moved with the form; these did not move at all.

The competitor charts follow the industry word as faithfully as everything else. Declare yourself an "Agency" and OpenAI's chart gives Syndral a 12% share of voice against BBDO Bangkok, Ogilvy Thailand and Leo Burnett. Declare "Professional Services" and the rivals become AI visibility tools like Profound and Semrush. A five-day-old site with zero mentions, out-shouting Ogilvy in one chart and Semrush in the next, depending on a form field.

Three verdicts, one form field apart

Same brand, same country, same day.

Gemini data reliability. Reliable Data: No Then: Reliable Data: Yes LinkedIn presence scored 75 out of 100 for high engagement. Syndral has never had a LinkedIn page. Competitor set. Declared “Agency”: BBDO Bangkok, Ogilvy Thailand, Leo Burnett. Syndral at 12% share of voice. Then: Declared “Professional Services”: Profound and Semrush. Brand archetype, all three engines. Product field “None”: Traditionalist Then: Product field “Ai Visibility Agency”: Innovator

Figure 3. Three verdicts that flip on one form field. Same brand, same country, same day.

The one number someone actually measured

Buried in the Test 1 report, and captured again in its replication two days later, is my favorite result of the whole study. Perplexity, the only engine of the three whose report showed any sign of a live lookup, went looking for Syndral and wrote: zero public mentions detected, no review platform presence, the brand has no public digital presence. Sentiment: 0 out of 100 in every category, which is the 0 out of 40 sub-score you saw earlier.

Two engines, one run, opposite verdicts

Perplexity, same run: 0/100. Sentiment, in every category. Zero public mentions detected, no review platform presence, no public digital presence. Honesty scored 7. Gemini, same run, same inputs: 30. An agile delivery model and sustained expansion, described for a brand with no public footprint. Nothing was found because there was nothing to find.

Figure 4. Two engines, one run, opposite verdicts. The correct answer is the one framed as a failing grade.

That is the correct answer. It is the only exact measurement in nine runs, and the report frames it as a failing grade to be improved rather than a fact about a five-day-old brand. In that same run, on those same inputs, Perplexity's honesty scored 7. Gemini described an agile delivery model and sustained expansion it could not have seen anywhere, and scored 30.

When the machine merges two companies

The same report, again captured in the replication, also states that Syndral is headquartered at Asoke Towers in Bangkok. Syndral has never had an office at Asoke Towers. The building is real and the address is real. It belongs to somebody else.

Asoke Towers is the Bangkok head office of Syndacast, a real and established digital marketing agency whose own site says it drives digital performance through audience targeting, big data and machine learning. The same report lists Syndacast among Syndral's competitors, hands Syndral its address, and a few lines later describes Syndral as delivering ROI-driven performance digital marketing in Thailand. Syndral, Syndacast: two names close enough that a retrieval system merged the entities and gave one company's facts to the other.

To be plain about it: Syndacast has done nothing wrong here and is not part of this story except as collateral. That is what makes entity confusion worth your attention. It does not need a fake brand or a trick prompt. It happens between two legitimate companies whose names share a syllable, inside a tool that presents the result with confidence scores attached.

What this means for measuring AI visibility

None of this means AI visibility is unmeasurable, and none of it retires the fundamentals. Ahrefs' December 2025 study of 75,000 brands found branded web mentions correlating with AI visibility at 0.66 to 0.71 while raw backlink counts barely registered, with the usual caveat, theirs and mine, that correlation is not causation. The classic work of earning mentions, building a clear site and being a real entity still sits under everything. SEO practitioners who say the levers have not changed overnight are right.

What the field lacks is honest measurement, and the shape of honest measurement is known. When SparkToro and Gumshoe studied AI recommendation consistency in January 2026, they ran twelve prompts 2,961 times through ChatGPT, Claude and Google's AI, using 600 volunteers, modeled on Carnegie Mellon's LLM consistency research. They deliberately left conditions uncontrolled, no instruction on chat history, device or country, because they wanted the variety a normal buyer actually meets. Some volunteers were longtime users with years of history, others had never opened the tool. Their finding, for ChatGPT and Google's AI: under a 1 in 100 chance that two identical prompts return the same brand list, closer to 1 in 1,000 for the same order. Their conclusion: single answers cannot be tracked, but the frequency a brand appears across many repeated runs can. Worth knowing when you weigh that: Gumshoe sells AI visibility monitoring. The finding still cuts against the easy version of its own product, and the full method is published, which is more than most vendor research offers.

Hold the two studies side by side and the requirements write themselves. A real AI visibility measurement runs many repetitions, in the real interfaces buyers use, in sessions carrying no history of their own, so that what moves is the brand signal and not the tester's past conversations, and reports frequency with its variance. That takes hours per brand, every time. A free instant grader cannot do it, whoever builds it, because the instant is the problem: one sample from a system this inconsistent measures nothing, so the tool must fill the gap with whatever it has. What it has is your form.

One sample from a system this inconsistent measures nothing, so the tool must fill the gap with whatever it has. What it has is your form.

Where Syndral stands on scoring

Syndral publishes a free AI Visibility ScoreCard, and after three days inside a grader I want to be exact about what it is. It is a declared self-assessment. It estimates your readiness from your own answers, it costs nothing per run, and it never claims to have measured your brand in the wild, because a free instant tool cannot do that, and everything above is why.

The measured version is different work. Syndral's paid Diagnostic runs repeated buyer-style prompts by hand, in the real chat interfaces, in temporary sessions with personalization cleared, and reports the median with the share of runs your brand actually appeared in. Slow, and honestly priced for it. The alternative, as you have now seen, is a fast number that follows your typing. If you want a two-minute read on how ready you are, take the ScoreCard, and know it for what it is: your own answers, scored. If you want to know where you actually stand, that is the Diagnostic, and it takes hours because it has to. Either way, the next field studies will land in our resources, and the method behind both is on the about page.

I have not contacted HubSpot yet. I will within a few days of publication, offering them a right of reply, and I will update this piece with anything they say. The complete run log, screenshots and PDF reports are open in the evidence annex. Check my work. That is rather the point.

Sources

FAQ

Frequently asked questions

Is HubSpot's AI Search Grader accurate?

No, not as a measurement of your brand. In our nine-run test it returned identical stored results for identical inputs, and moved 8 points when we changed one word in the form. It grades what you type, not what AI systems say about you. As a demo of what an AEO report looks like, it works.

Why did my AI Search Grader score change when I edited the form?

The results page is built from the four form fields, passed as URL parameters. Change the industry or product wording and the engines are prompted about a differently described brand. In our tests one changed word moved a score by 8 points.

How do I actually measure my brand's AI visibility?

Run the same buyer-style prompts repeatedly in the real chat interfaces and record how often your brand appears. Five runs per prompt per platform is enough for a first read. Use temporary or incognito chats with account personalization cleared, because your own history changes what you are shown. Frequency is the metric, never one answer.

Does a low AI visibility score mean AI never recommends my brand?

No. A single automated check is one sample from a noisy system, and free graders only see how you filled their form. Repeated manual runs can still show real visibility, or real invisibility, that a one-shot score misses in either direction.

What is a good AI visibility score?

There is no established benchmark, and a score built from your own form answers has no scale to be good on. The measurable version is the share of relevant buyer prompts where your brand appears, tracked across repeated runs. That number is comparable to itself over time, which is what you need.

How can I tell if an AI visibility tool is measuring or guessing?

Run it twice with identical inputs, then change one word in a form field and run it again. If identical inputs never vary and one word moves the score, it is reading your form. Then ask whether it runs repeated prompts in the real chat interfaces, in sessions with no history, and reports how often you appeared.

Written by
Kevin Pantanella
Founder, Syndral

Syndral is my AI visibility practice, the new brand of Suzaku Productions, the Bangkok web agency I have run since 2014.

Where this leads

See where AI names your competitors and skips you.

The first step is the same one behind every study we publish: a clear, honest look at where AI names you today.

No call required. We show you the questions, the names, and your gap.