Summary
This is not an owner survey. We tested 24 realistic growth and project-acquisition questions that architecture and design-build firm owners could plausibly ask, across five AI systems (ChatGPT, Claude, Gemini, Google AI Overview, Perplexity), and manually classified the answers, citations and recommended tactics.
Strongest findings:
- Referrals, portfolio/case studies, CRM/pipeline management and networking were recurring recommendations across nearly every system.
- A website was positioned far more often as a post-referral trust-verification layer than as a direct client-acquisition source.
- Of roughly 90 citation instances reviewed, only about 2% qualified as primary research; the rest was largely marketing-agency, SEO-agency and vendor content.
- GEO/AEO was never spontaneously recommended across 72 answers where it wasn't directly asked about.
- Developer/hospitality/institutional large-project transition and Turkey-specific tender/RFP context were the weakest answer clusters; one ChatGPT answer imported US federal procurement rules directly into a Turkish-context question.
- A significant share of Turkish-language answers cited nothing at all; the ones that did leaned heavily on generic SEO/marketing content.
This page measures AI systems' answers, not owners' actual behavior. Reading it as "96 architecture firm owners were studied" would be incorrect: what was tested is the roughly 96 answers five AI systems produced to 24 questions.
What we tested
The 24 questions span the commercial territory architecture and design-build firm owners actually face:
- referral dependence and predictable pipeline
- winning larger, higher-budget projects
- attracting higher-budget, better-fit clients
- evaluating marketing and channel investment
- SEO, digital authority and AI visibility
We're not publishing the exact question wording or how it was selected, only the category-level scope, so the measurement stays a clean read on the current AI answer ecosystem rather than a template for gaming a future rerun.
The language mix was deliberately split between Turkish and English; some questions were tested as matched Turkish/English pairs, specifically to check whether language itself changed answer quality.
An "answer" was counted as one complete model response, generated in a fresh, history-free session for each question. A "citation instance" was counted as any domain, organization or vendor named within an answer (an inline citation chip, a footnote-style link, or a source named in prose).
Platform coverage
| Platform | Questions completed | Citation behavior | Important limitation |
|---|---|---|---|
| ChatGPT | 24/24 | Cited sources on roughly half the answers; AIA (American Institute of Architects) was cited repeatedly | Model version wasn't visible in the anonymous interface |
| Claude | 24/24 | Zero citations across all 24 answers; answered purely from model knowledge | Tested as a context-free, blind instance; no web search was used |
| Gemini | 24/24 | Web-search grounding fired on only 3 of 24 answers | The session showed the model as "3.5 Flash-Lite"; one answer required an in-app retry after a generation error |
| Google AI Overview | 23/24 | Cited sources on most answers, weighted heavily toward marketing-vendor content | 1 answer returned a Local Results pack instead of an AI Overview; no model version is ever exposed |
| Perplexity | 1/24 | Cited 15 sources on the single completed answer, none architecture-industry-specific | Hit a Cloudflare bot-verification challenge after the first query; per instructions, not bypassed, so the remaining questions couldn't be tested |
What AI systems recommended most often
The frequencies below come from manually coding roughly 96 real answers: not automated text analysis, but human judgment, approximate, and subject to coder interpretation.
| Tactic | Approximate frequency |
|---|---|
| Referrals | ~85/96 |
| Portfolio/case studies | ~80/96 |
| CRM/pipeline management | ~70/96 |
| Networking | ~65/96 |
| Website | ~60/96 |
| Specialization/niche positioning | ~55/96 |
| SEO/local SEO | ~50/96 |
| Developer relationships | ~30/96 |
| Outbound/cold sales | ~30/96 |
| Content marketing | ~30/96 |
| Social media | ~28/96 |
| LinkedIn (as a distinct channel) | ~20/96 |
| Competitions | ~15/96 |
| PR/awards | ~15/96 |
| Paid advertising | ~15/96 |
| Business-development hire | ~12/96 |
| Pricing/fee positioning | ~10/96 |
| RFP/tender strategy | ~8/96 |
| GEO/AEO/AI visibility (spontaneous) | 0/72 |
The GEO/AEO row is worth flagging separately: it wasn't volunteered even once across 72 answers where it wasn't directly asked about. The SEO row has a nuance too: architecture-specific SEO provider supply is quite dense, but direct buyer-demand evidence isn't as strong (see below).
Finding: referral still sits at the center
Referrals, portfolio/case studies and networking were the most consistently recommended tactics across nearly every platform, not a surprising result on its own.
The more useful question for an owner isn't whether referral works. It's: even where referral is strong, does the firm have a second, more predictable channel underneath it, or does growth simply stop when the referral pipeline slows down?
We go deeper on that diagnosis in a separate article on winning larger projects beyond referrals.
Citation landscape
Citation quality varied significantly across the ~90 instances reviewed. Genuinely authoritative, architecture-specific sources (professional associations, verified primary research) were rare; the large majority of citations were marketing-agency, SEO-agency or generic vendor content rather than primary evidence. Cross-platform agreement on which sources to cite was weak: the same source was rarely cited by more than one platform for the same question.
We're not publishing a source-by-source breakdown here. Naming which vendors or agencies were cited more or less often would say more about the current content landscape than about an owner's actual growth problem, which is what this page is for.
The Turkey evidence gap
This finding earns its own section because it directly affects the real quality of answers aimed at the Turkish market.
- Roughly half of Turkish-language answers returned no citations at all: purely model-generated, unverifiable content.
- Where Turkish-language answers were cited, sources skewed heavily toward generic SEO/marketing content rather than architecture-specific or primary sources.
- Independent Turkish primary data was extremely scarce; genuinely architecture-specific professional-association sources (sector bodies such as GYODER, konutder and İNDER) appeared only once, on a single question.
- One answer referenced a source we could not independently verify. We're not repeating it here, only noting it as a methodological-integrity example: AI systems can sometimes cite sources that are unverifiable or appear not to exist.
We apply the same rigor to foreign sources: several ChatGPT answers referenced AIA survey data, but we couldn't independently verify the full underlying document within this study. So we're not presenting those specific figures as confirmed fact here, only the observation that the AI system cited that source.
Where AI answers were strongest
Website = verification-infrastructure framing: the strongest cross-platform convergence in the entire benchmark. Asked whether a referral-driven firm's website is even necessary, all four fully-tested platforms (ChatGPT, Claude, Gemini, Google AI Overview) independently arrived at the same framing: the site functions as a post-referral trust-verification layer (covering credibility, past work, specialization and capacity) rather than a direct lead source.
This lines up with a simple way to think about a firm's digital presence: Discovery (how did a prospect first hear about you), Verification (what did they check once they were seriously considering you), and Decision (why did they choose you). For a referral-heavy firm, the website's real job often sits in the middle stage.
We cover what this finding means for architecture firm owners in a separate editorial piece.
Even the strongest answers here remained limited by US-to-Turkey evidence transfer: no platform had Turkish primary data backing the same claims.
Demand validation showed that the website question isn't only a theme AI systems recommend often; it's a real decision problem in the open web/search environment too. This is the strongest intersection of demand and answer quality in the whole benchmark (see below).
Where AI answers were weakest
- Developer/institutional/hospitality transition: questions in this cluster (reaching developers, winning hospitality clients, entering RFPs, diagnosing why a firm can't win larger projects) got generic, unsourced, or wrong-jurisdiction answers.
- Turkish RFP/procurement context: one ChatGPT answer imported US Federal Acquisition Regulation guidance directly into a Turkish-context question, a clear jurisdiction mismatch. We're reporting this as one concrete, illustrative example, not sensationalizing it.
- Channel effectiveness for high-value projects: no platform could quantify which channel actually produces large, as opposed to merely frequent, projects.
- Turkey-specific acquisition-channel evidence: essentially absent across the whole panel.
- Premium pricing and fee-pressure evidence: thin across the board; most fee-related content came from generic marketing-agency blogs, not primary data.
There's a notable overlap worth naming here: external demand signal for winning larger projects isn't entirely absent: a small but real ecosystem exists, particularly in the English-language market. At the same time, this is one of the clusters where AI answers were weakest in the benchmark itself (see below).
For an owner, the useful diagnostic isn't "which channel should I use more." It's narrower: is the real bottleneck access (reaching the right decision-maker), track record (relevant project history), verification (can that history be confirmed externally), intent (are the right people even looking), or capacity (could the firm actually deliver the next tier up)? We break this down in a separate decision framework.
The GEO/AEO finding
Observed: GEO/AEO was never spontaneously recommended across 72 answers where it wasn't directly asked about.
Observed: When asked directly (the one control question), platforms disagreed with each other, and the answers arguing it's important leaned heavily on sources with a direct financial interest in that being the answer.
What this means: an architecture firm owner's natural growth problem, at least as reflected in current AI answers, isn't "do GEO." AI visibility may be a real technical distribution question, but it didn't surface here as the owner's own bottleneck.
Demand validation points the same direction. Across the Turkish and English queries tested, we found no strong evidence that AI/GEO/AEO visibility is a natural, recurring commercial problem for architecture firms. So we're not drawing a new "AI visibility opportunity" conclusion here (see below).
Is there real demand behind these questions?
An AI system giving a weak answer to a question doesn't mean that question is commonly experienced by architecture firms. So we separately tested the benchmark's main problem families through search behavior, open-web discussion and provider/source supply density.
The core discipline of this section: SOURCE_GAP ≠ DEMAND_FREQUENCY. An AI system answering a question poorly (a source gap) is not evidence that architecture firm owners actually encounter that question often (demand frequency); we measured the two separately.
Direct ChatGPT query volume isn't available, so the results rest on proxy evidence: Google Ads search volume, SERP/PAA/related-search behavior, social surfaces and provider/source supply were considered together.
| Problem family | Demand frequency | AI answer gap | How well does the page answer it? |
|---|---|---|---|
| Referral / pipeline (REFERRAL_PIPELINE) | MEDIUM | LOW | Partial, a supporting article is direct |
| Larger projects (LARGER_PROJECTS) | MEDIUM | HIGH | Partial |
| Premium clients (PREMIUM_CLIENTS) | MEDIUM, market split | HIGH | Mentions only |
| Marketing / channel choice (MARKETING_CHANNEL_CHOICE) | MEDIUM-HIGH | MEDIUM | Partial |
| Website (WEBSITE) | MEDIUM-HIGH | LOW | Direct |
| SEO / discovery (SEO_DISCOVERY) | MEDIUM | MEDIUM-HIGH | Mentions only |
| AI / GEO / AEO (AI_GEO_AEO) | LOW / UNPROVEN | — | Direct, a correct negative conclusion |
The kind of evidence behind each family isn't the same:
- Referral / pipeline: Search volume isn't measurable, but hiring listings and practitioner discussion show the problem is real.
- Larger projects: A niche but genuine ecosystem exists in the English-language market; search behavior is more scattered on the Turkish side. At the same time, this is one of the clusters where the benchmark's own AI answers were weakest.
- Premium clients: Signal is strong in the English-language market; the Turkish test query drifted into unrelated SERP behavior, so Turkish-market demand isn't validated yet. The current Turkish-side test shows a query-language measurement failure, not an absence of demand.
- Marketing / channel choice: Real volume and dense provider supply are visible in the English-language market; the benchmark touches the topic without fully resolving the channel-choice decision.
- Website: This is the benchmark's strongest intersection: real demand exists, and the existing page answers this question directly and well.
- SEO / discovery: Direct buyer-demand evidence is thin, but architecture-specific SEO provider supply is one of the densest clusters in the study. The provider market's conviction here is stronger than the direct search evidence.
- AI / GEO / AEO: No sufficient evidence was found that AI/GEO/AEO visibility is a natural, recurring commercial problem for architecture firms. Four independent evidence lines converged on the same result: no search volume, Turkish search intent drifting toward AI rendering/design-tool use, no social evidence in English, and CR-2026-003's own finding that GEO/AEO was never spontaneously recommended across 72 non-direct answers.
Search volume alone is not demand. Some problems, like referral dependence, can be confirmed through behavioral signals (hiring listings, peer discussion) despite low or zero keyword volume. Conversely, dense provider supply on its own is not treated as strong evidence of buyer demand.
Demand validation shows that CR-2026-003 isn't entirely disconnected from real buyer problems, but not every problem family is equally well supported. The website question is the strongest area on both demand and answer quality. Referral/pipeline, larger projects and marketing/channel problems reflect real demand, but the page mostly answers them only partially. Premium clients and SEO show information gaps where demand evidence isn't yet equally strong. AI/GEO/AEO was not confirmed in this research as a natural architecture-firm problem.
Full methodology, the keyword set used, and the falsification tests are kept in ClassyDesign's internal research archive; only a category-level evidence summary is shared here.
What this benchmark does not prove
- Doesn't prove owner behavior (this measures AI answers, not real firm behavior).
- Doesn't prove channel ROI.
- Doesn't prove SEO generates high-value projects.
- Doesn't prove a website increases project win rate.
- Doesn't prove GEO/AEO affects revenue.
- Doesn't prove AI platform behavior is stable over time (single-run snapshot).
- The Perplexity sample (1/24) is insufficient for any conclusion about that platform.
- This is a single-run measurement; no repeated sampling was done.
- AI outputs are stochastic; rerunning comparable questions on a different date may produce different results.
- Model-version visibility was incomplete (not shown for ChatGPT/Claude in this run; Gemini showed "3.5 Flash-Lite"; Google AI Overview never exposes one).
- Location/login/personalization context can influence outputs; this run used anonymous, logged-out sessions with unknown/default location state.
- Manual coding of tactics and citations is subject to coder judgment, not automated NLP extraction.
What this means for an architecture firm owner
If you want more, or larger, projects, the first question usually shouldn't be "do we need more visibility." It's worth separating a few things first:
- Are we actually reaching the right decision-maker, or only the people already inclined to refer us?
- Do we have relevant project history for the specific type of client we want next?
- Can that history be verified by someone outside the firm, not just claimed?
- Are the enquiries we get simply the wrong scale for where we want to go?
- Could the firm actually deliver the next tier of project if it landed tomorrow?
None of this is answered by a single AI benchmark, and it shouldn't be. It's a diagnostic starting point, not a prescription.
Where a firm's real bottleneck sits (traffic, referral dependence, positioning, verification, or the acquisition structure underneath all of it) determines what's worth investing in next. Starting with digital spend before that's clear tends to grow the wrong problem faster.
That diagnostic is what ClassyDesign's customer-acquisition framework looks at. If the AI-visibility layer itself is the open question, that's covered separately in ClassyDesign's approach to AI discovery.
Where do these findings sit in a found-understood-trusted-chosen chain?
This section introduces no new measurement; it re-reads the findings above through ClassyDesign's framework for how a brand gets found, understood, trusted, and chosen. The page's own "Discovery → Verification → Decision" split (above, in Strongest answers) is a simpler version of that same four-stage model; below we map the same findings onto the full framework. This benchmark produced strong evidence for two of the four stages and partial evidence for one; it didn't directly test the understood stage, so that one doesn't appear here.
Found
In this benchmark, being found runs through referral networks, not digital search. Referrals, portfolio/case studies, and networking were the most consistently recommended tactics across nearly every AI system (roughly 85/96, 80/96, and 65/96 respectively); the website appeared in only about 60/96 answers and mostly in a secondary role. GEO/AEO never being volunteered across all 72 not-directly-asked answers points the same way: on the current evidence, digital discoverability isn't the architecture firm owner's natural growth bottleneck; the referral network is.
Trusted
This benchmark's single strongest finding lives here: four platforms (ChatGPT, Claude, Gemini, Google AI Overview) independently converged on the same framing: a website isn't a direct lead source, it's a post-referral verification layer. This finding is also supported in the open web/search environment, not just repeated across AI systems.
Chosen
This benchmark didn't directly measure the selection moment itself. Still, the five questions that diagnose the growth bottleneck (access, project history, verification, demand fit, capacity) point to the factors that shape being chosen. The capacity question in particular ("could the firm actually carry the next tier of project") relates directly to executability once chosen.
This research doesn't just measure AI answer quality; it helps separate which stage an architecture firm's growth question is actually stuck at. ClassyDesign produces benchmarks like this not to sell digital visibility as an automatic fix, but to diagnose the real bottleneck at the right stage. If you'd like to talk through which stage your firm is stuck at, get in touch.
Methodology
Question selection: 24 questions spanning referral dependence, larger-project transition, developer/commercial-buyer access, website credibility, pipeline predictability, and SEO/AI visibility. Exact question wording and selection methodology are kept unpublished (see FAQ).
Platform coverage: ChatGPT, Claude, Gemini, Google AI Overview, Perplexity (see table above).
Date: 22 September 2026.
Manual citation classification: each citation instance was manually assigned to one of: professional association, architecture business-development consultancy, generic marketing agency, SEO agency, Reddit/forum, academic, survey/primary research, platform/vendor blog, other.
Manual tactic coding: an 18-item tactic list was manually coded from roughly 96 real answer texts.
Incomplete Perplexity coverage: a Cloudflare bot-verification challenge on the second query; the remaining questions couldn't be tested.
Single-sample nature: each question was run once per platform; there is no automated reproducibility guarantee.
Language/locale limitations: anonymous, default-locale sessions were used; no city-level Turkey targeting was applied.
Model-version limitations: see table above.
FAQ
Did this research study architecture firm owners?
No. This is not an owner survey; it's an AI answer benchmark. We compiled realistic questions real owners could plausibly ask, then tested those questions across five AI systems. What we measured is AI answer quality and citation behavior, not owners' actual behavior.
Is GEO/AEO necessary for architecture firms?
The current evidence doesn't support that. GEO/AEO was never spontaneously recommended across 72 answers where it wasn't directly asked about; when asked directly, platforms disagreed, and the affirmative answers leaned heavily on sources selling that exact service.
Is this benchmark reproducible?
Partially. The question categories and full methodology are published; the exact question wording is kept unpublished so the measurement can't be gamed by future benchmark runs. AI outputs are also non-deterministic; rerunning comparable questions on a different date may produce different results. This is a single-run measurement from 22 September 2026.
Why was Perplexity only tested on 1 prompt?
The second query hit a Cloudflare bot-verification challenge. Per our own instructions, we did not attempt to bypass it; the remaining questions couldn't be tested and aren't part of the dataset.
Sources
A. ClassyDesign first-party findings
All benchmark findings on this page are ClassyDesign's own measurement (22 September 2026). None depend on an external source.
B. External contextual sources
- The American Institute of Architects, 5 ways to level up your firm marketing, independently verified (22 September 2026): states that referrals are among the highest-quality lead sources for architecture firms, though they can dry up during economic downturns. The specific percentage statistics some AI answers referenced (e.g., repeat-client share) are not presented as confirmed fact on this page, since the underlying source document couldn't be fully verified in this study; only the observation that the AI system cited that source is reported (see the Turkey evidence gap section).
This benchmark is a snapshot. AI output and citation behavior can change over time; future reruns will be published as a separate, dated version rather than silently overwriting these findings.