Key takeaways
- Bluefish AI raised a $43M Series B in April 2026 and positions itself as an "Agentic Marketing Platform" for the Fortune 500, but pricing is quote-only and enterprise contracts reportedly run well into six figures. Due diligence matters more here than with a $99/month tool.
- The biggest gaps reported by independent reviewers: opaque pricing, no GA4 or conversion attribution, restricted API access on lower tiers, and domain-level (not URL-level) citation data.
- AI platforms change citation behavior overnight with model updates. Around the GPT-5.3 rollout in March 2026, ChatGPT's average citations per response dropped roughly 27% in a single day. Ask every vendor how they handle this before you sign.
- "SOC 2-aligned" is not SOC 2 certified. Get the actual report, check the type (Type II, not Type I), and check the coverage period.
- Run your own fixed-query trial before signing anything. A decision this size should rest on evidence from your own test, not a vendor demo.
Why this matters more than usual in 2026
Enterprise GEO has moved from "interesting experiment" to "line item in the budget" fast. Bluefish AI closed a $43 million Series B in April 2026, bringing total funding to $68 million, with backing from Threshold Ventures, NEA, Amex Ventures, Salesforce Ventures, and Bloomberg Beta. Their named customers include Adidas, American Express, Hearst, Ulta Beauty, and Tishman Speyer. That's a serious company.
It's also a two-year-old company selling multi-year contracts at price points that, per third-party reporting, can reach $100K–$500K+ annually for large deployments. And the category itself is young enough that buyers are, as one industry critic bluntly put it, "choosing among them blind."
So this guide isn't an anti-Bluefish piece. It's a due-diligence checklist that happens to use Bluefish as the case study, because they're the vendor most enterprise teams are evaluating right now. Every question below applies equally to Profound, Scrunch, or any other platform you're putting through an RFP.
One more piece of context before the questions. A standard IT RFP template misses up to 60% of the risk-relevant questions specific to AI vendors, according to enterprise RFP research from Worqlo. AI vendors route your data to third-party model APIs, produce non-deterministic outputs, and sit inside a compliance landscape (EU AI Act, new liability frameworks) that older contract templates simply don't cover. The ten questions below close that gap for GEO specifically.
The 10 questions
1. Where does your data actually come from: live UI, API calls, or a synthetic panel?
This is the foundation question, because everything downstream depends on it. AI platforms can show different answers in their user-facing interfaces than what their APIs return. A vendor monitoring only APIs may report visibility scores that don't match what real users see in ChatGPT or Google AI Overviews.
Ask specifically:
- Do you query the actual user interfaces of ChatGPT, Gemini, Perplexity, Claude, and AI Overviews, or just the APIs?
- How many responses do you sample per prompt, per platform, per week?
- Is the methodology documented anywhere I can read it?
A good answer names the method and the sampling cadence without hesitation. A bad answer is vague deference to "proprietary technology." Platforms like Promptwatch publish that they monitor the real UIs of the major platforms precisely because API-only data diverges from what users see, and their methodology notes are public. If a vendor can't or won't explain their collection method at the same level of detail, that tells you something.

2. Is your SOC 2 Type II certified, or just "aligned"?
Here's a claim worth verifying directly. Bluefish's marketing materials reference "SOC 2-aligned controls," and a competitor-published comparison (Profound's, so treat it as a claim to verify, not neutral fact) states that Bluefish's SOC 2 audit was "in progress, not yet certified" as of their review. "Aligned" means the vendor follows practices similar to the standard. Certified means an auditor actually tested the controls over time.
Ask for:
- The most recent SOC 2 report, including type (Type I is a point-in-time snapshot; Type II covers a period and is the enterprise baseline) and the coverage dates
- A signed GDPR Data Processing Agreement if you process EU personal data
- The full sub-processor list, which GDPR requires anyway
- Breach notification SLAs (72 hours is the GDPR minimum)
- Penetration test frequency and the most recent summary
If a vendor stalls on handing over the actual report during a sales process, they will not get faster about it after you've paid.
3. How do you handle model updates that change citation behavior overnight?
This is the question almost nobody asks, and it's the one that will save you from a painful Q3 conversation with your CFO.
Here's what happens in the real world. When GPT-5.3 rolled out on March 4, 2026, ChatGPT's average citations per response dropped from roughly 6.4 to around 4.7–4.9 across all models simultaneously, and stayed there a month later, according to Promptwatch's data on the post-GPT-5.3 citation drop. That's about 27% fewer citation slots available to your brand overnight, with nothing your team did wrong and nothing your vendor could have prevented.

If your vendor's visibility score drops 20% the week after a model update, you need to know in advance whether their reporting separates "platform-side shift" from "your content lost ground." Ask:
- How do you annotate model updates in our reporting?
- Will you proactively tell us when a citation drop correlates with a platform change rather than our performance?
- How do you normalize scores across model versions?
Also worth knowing: citation behavior varies wildly by engine. Per Promptwatch's average sources per response data, ChatGPT cites around 5 sources per web-search response, Google AI Overviews around 10, Perplexity is remarkably stable at almost exactly ten, and Microsoft Copilot has swung from under 2 to nearly 17 within weeks. A vendor quoting a single "visibility score" across all engines without explaining these baselines is flattening something that shouldn't be flat.
4. Do you give me URL-level citation data, or just domain-level?
This one decides whether the platform can actually drive content work. Competitor-published comparisons claim Bluefish provides domain-level citation data only, with no individual URL breakdown. Independent hands-on reviews corroborate that Bluefish's strength is monitoring and brand governance rather than page-level content optimization.
Why it matters: "your domain got cited 40 times" is a vanity metric. "This specific comparison page got cited for these three prompts, and here's the wording the AI used" is an optimization roadmap. If the vendor only offers domain-level data, ask how their content recommendations are supposed to work without knowing which pages earn citations.
5. What happens after the dashboard shows a problem?
GEO platforms split into two species: trackers and optimizers. Trackers tell you that you're invisible. Optimizers help you do something about it, with content gap analysis, briefs, generation, CMS publishing, and prioritized action lists.
Ask the vendor to walk you through the full loop: we detect a gap, then what? Who writes the content? Where does it get published? How do we know it worked? A hands-on review of Bluefish found they provide recommendations and content briefs but no content generation, optimization, or publishing pipeline. That may be fine if you have a 50-person marketing team (which, to be fair, is exactly who Bluefish is built for). It's not fine if you expect the platform to close the loop itself.
If you want the loop closed, say so in the RFP and make the vendor demo it with your data, not theirs.
6. Can we see and control the prompt methodology?
Your visibility score is only as honest as the prompt set behind it. Two vendors can report wildly different scores for the same brand purely because of which prompts they track and how they weight them.
Ask:
- Who builds the prompt set, us or your services team?
- Is there a self-serve interface to add, remove, and weight prompts?
- Do you have real user prompt volume data, or are prompt suggestions heuristic?
- Do you expose query fan-outs (the sub-queries AI engines expand a prompt into)?
Competitor comparisons claim Bluefish's prompt tracking runs as a "managed engagement" with no self-serve prompt configuration and an undisclosed execution methodology. Again: verify directly. But the underlying question is universal. If you can't audit the prompt set, you can't audit the score, and you'll spend every quarterly review arguing about the measurement instead of acting on it.
7. What's the actual price, and what triggers overage fees?
Bluefish has no public pricing. Contracts are delivered via Order Forms, invoiced annually in advance, with a hybrid model that includes potential overage fees for exceeding usage limits. Third-party reporting puts entry tiers somewhere in the $99–$799/month range with enterprise contracts running to five and six figures annually, but these numbers are anecdotal, not official.
The questions that matter:
- What exactly counts as "usage"? Prompts tracked, responses sampled, sites monitored, seats, or some combination?
- What happens when we hit the limit: hard stop, or overage billing? At what rate?
- Is there an annual cap on overage exposure?
- What does renewal pricing look like, in writing?
- What's the exit clause if we're not seeing results in six months?
Because both Bluefish and Profound quote "custom" at the enterprise tier, competitive bids are the only real pricing benchmark you have. Get at least two quotes before you sign anything, and make sure the vendors know they're competing.
8. How does this connect to revenue?
An independent 30-day review of Bluefish flagged the absence of GA4 or conversion/attribution integrations as a real limitation. That's a big deal for an enterprise purchase, because the eventual question from your CFO will not be "what's our visibility score," it will be "what did this do to pipeline."
Ask:
- Can you attribute AI-driven traffic and conversions on our website, not just count mentions?
- Do you integrate with GA4, our BI stack, or anything downstream of the dashboard?
- Can we export raw data via API or CSV without paying for a higher tier?
Mentions are interesting. Traffic and conversions are the business case. If the platform can't connect the two, budget for the analytics work you'll have to build yourself.
9. What happens to our data?
Three AI-specific data questions that standard SaaS contracts miss:
- Model training. Does the vendor use our prompts, inputs, outputs, or brand data to train or fine-tune their models? Get the restriction in writing in the contract, not in a marketing FAQ. Legal guidance on AI vendor agreements is unambiguous on this point: contracts should explicitly restrict vendor use of customer data for model training unless separately negotiated.
- Third-party routing. Which LLM APIs (OpenAI, Anthropic, Google) receive our data when the platform runs its queries, and are they covered by the DPA?
- Deletion. What's the data retention policy, and what's the deletion timeline (including backups) at contract end?
10. Can we run our own fixed-query trial before signing?
This is the closer, and in my opinion the single best predictor of whether you'll be happy with the vendor.
The method, adapted from GEO platform validation research: build a fixed set of 20–50 queries mixing head terms, long-tail terms, and conversion-tied questions. Run controlled trials varying one variable at a time: sampling time, model version, geographic region, login state. If the platform shows no variation when you change region, that's a red flag that it isn't actually using regional data. If the top five cited sources shift more than about 10% between identical runs, the sampling is too thin to trust.
Then compare the vendor's numbers against what you see when you run the same queries yourself in ChatGPT and Perplexity. If they diverge materially, ask why before you sign, not after.
A vendor confident in their data will hand you a sandbox and cheer you on. A vendor who insists you trust their demo numbers instead is telling you what the next three years of the relationship will feel like.
The questions at a glance
| # | Question | Good answer sounds like | Red flag |
|---|---|---|---|
| 1 | Where does your data come from? | Named method, sampling cadence, UI vs API explained | "Proprietary methodology" |
| 2 | SOC 2 status? | Type II report handed over, with coverage dates | "Aligned" or "in progress" |
| 3 | How do you handle model updates? | Updates annotated in reporting, platform shifts separated from your performance | Scores move with no explanation |
| 4 | Citation granularity? | URL-level data with the AI's actual wording | Domain-level counts only |
| 5 | What happens after the dashboard? | Gap detection through content creation and publishing, demoed on your data | "We provide recommendations" |
| 6 | Prompt methodology? | Self-serve prompt control, real volume data, fan-outs exposed | Managed black box you can't audit |
| 7 | Real price and overages? | Usage limits, overage rates, caps, renewal terms in writing | Quote-only, annual advance, vague overages |
| 8 | Revenue connection? | GA4/attribution integrations, raw data exports | Mentions only, no traffic data |
| 9 | Data handling? | Written no-training clause, full sub-processor list, deletion timeline | Verbal assurances |
| 10 | Our own trial? | Sandbox access, encouragement to test | Pressure to close before you test |
If you're comparing vendors
Don't run this evaluation with a single vendor. Bluefish, Profound, Scrunch, and the broader field each make different trade-offs between monitoring depth, execution help, and price, and the only way to see those trade-offs clearly is side by side. The GEO software directory at bestgeosoftware.com is a good starting point for building a shortlist beyond the two or three names your sales reps keep mentioning.
A few honest notes on the sourcing above: several of the sharpest claims about Bluefish's limitations (domain-level citations only, managed prompt methodology, no content pipeline) come from a comparison page published by Profound, a direct competitor. I've flagged them as claims to verify rather than established fact, and the independent tryanalyze.ai review corroborates the pricing opacity and missing attribution pieces. That's exactly why question 10 exists. Don't take my word for it, don't take Profound's word for it, and don't take Bluefish's word for it. Run the trial.
Bottom line
Bluefish has real, differentiated strengths. The AI Brand Vault for brand fact governance is genuinely useful for large orgs where legal, PR, and marketing all touch brand claims, and their enterprise benchmark research on how 194 large brands appear in AI is worth reading regardless of which platform you buy. But a two-year-old company with quote-only pricing, no public methodology, and annual-in-advance contracts is a vendor you negotiate with from a position of information, not hope. These ten questions get you that information. Any vendor worth six figures of your budget should answer all of them without flinching.

