Summary
We crawled 17 firm websites that appear on Google's first pages for Istanbul architecture-office searches, without looking at their source code or admin panels, over HTTP only. The question: on sites that already rank, are the "found" and "understood" layers in place?
Main findings (n = 17, at most 30 URLs per site, 26 September 2026):
- Meta description is missing on at least one crawled page on 13 sites, and on more than half of the crawled pages on 3. Duplicate descriptions on 9.
- Canonical points at itself on every crawled page on only 11 sites. 1 site has no canonical at all; on 1 the canonical points at
.htmlvariants of the pages, on 1 at a different domain. - H1 is missing on at least one page on 10 sites; 8 sites have multiple H1s. Over-long titles on 13.
- robots.txt found on 14 sites, an XML sitemap on 13.
- Structured data on at least one page on 15 sites; a business-describing type (Organization, LocalBusiness, Architect and so on) on 13. This is a presence count, not a correctness measurement.
- Bot access: Googlebot received 200 on all 17. One site returns 406 to bingbot; another returns 403 to GPTBot and ClaudeBot.
- Image alt text missing on at least a quarter of images on 9 sites (range 0–83%).
- Median response time as seen by the crawler above 1 second on 4 sites.
- Independent audit score: median 85/100 (range 68–99).
This is not a ranking or traffic study. Ranking in Google's top 30 is the sample rule; what is measured is the technical state of those sites as seen from outside. The crawl is capped at 30 URLs per site, so measurements that need the whole site, such as orphan pages, are not reported.
What we measured
Every site was crawled with the same tool and settings: our source-blind crawler (site-auditor 0.2), without executing JavaScript, at most 30 URLs, 3 concurrent requests. For every page the crawler records status code, canonical, robots meta, title and description lengths, H1 count, hreflang, structured-data types, image alt text, internal links and a sitemap comparison. We then sent one request to each homepage with five user agents (browser, Googlebot, bingbot, GPTBot, ClaudeBot) and recorded the response code.
This is the same order as in our article on how to read an external audit: scope, indexability, canonical, titles and descriptions, schema and sitemap, bot access.
Sample
Three queries: "istanbul mimarlık ofisi", "istanbul mimarlık firması", "istanbul iç mimarlık ofisi". For each, Google's top 30 organic results (DataForSEO, Türkiye, Turkish, desktop, 26 September 2026). Social networks, directories, news and list pages and deep article URLs were removed by rule; each domain counted once. Of the remaining 18 domains one was a professional chamber and was excluded as not a firm. Result: 17 firm sites.
Sites are anonymised as site_01 … site_17. The aim is not to grade individual firms but to see how far the technical layer is in place on sites that rank.
Found layer
| Measurement | Sites (n = 17) |
|---|---|
| robots.txt present | 14 |
| XML sitemap found | 13 |
| Every crawled URL on HTTPS | 16 |
| Self-referencing canonical on every crawled page | 11 |
| No canonical on any page | 1 |
Canonical pointing at another version (.html variant or another domain) | 2 |
| Meta description missing on at least one page | 13 |
| Meta description missing on more than half of crawled pages | 3 |
| Duplicate meta descriptions | 9 |
| Title over the recommended length (at least one page) | 13 |
| Duplicate titles | 5 |
| H1 missing on at least one page | 10 |
| Multiple H1s (at least one page) | 8 |
| Internal links to redirecting URLs | 8 |
| Broken internal links (4xx) | 3 |
Title and description findings are the most common; canonical findings are the most serious. On the two sites whose canonical points at another version, part of the crawled pages fold into another address in the search engine's eyes: on one the .com.tr address points at the .com domain, on the other clean URLs point at .html copies. This kind of error has nothing to do with design, is visible from outside, and can be fixed on most platforms.
Understood layer
| Measurement | Sites (n = 17) |
|---|---|
| Structured data on at least one page | 15 |
| Structured data on every crawled 200 page | 12 |
| A business-describing type (Organization, LocalBusiness, Architect, HomeAndConstructionBusiness, Corporation) | 13 |
| hreflang present (multilingual site) | 4 |
| Social sharing meta tags missing | 4 |
The types seen look mostly like the default output of content-management plugins: WebSite, WebPage, ImageObject, BreadcrumbList and ListItem are frequent; a Service type appears on few sites. We counted presence only; we did not measure whether the blocks describe the business correctly or whether address and phone match the text on the site. So "13 sites have a business type" does not mean "13 sites are well described for AI systems".
hreflang on 4 sites is not a defect; the other 13 appear to be single-language. For a firm targeting international clients it is part of the verification layer described on the architecture firms page.
Bot access
| User agent | Homepages returning 200 | Exception |
|---|---|---|
| Browser (Chrome) | 17 | — |
| Googlebot | 17 | — |
| bingbot | 16 | 1 site: 406 |
| GPTBot | 16 | 1 site: 403 |
| ClaudeBot | 16 | same site: 403 |
We do not think the site returning 406 to bingbot does so on purpose; it looks like a hosting or security-layer filter, and the result is that the site is absent from Bing and from assistants that use Bing as a source. The site returning 403 to AI crawlers may be making a deliberate choice; that is a decision, not a defect. But if the decision is not deliberate, that site is not among the sources ChatGPT's and Claude's web crawlers can read. Both situations can only be found this way: from outside, with different user agents.
Speed and content basics
- Median response time as seen by the crawler is above 1 second on 4 sites; the median across the 17 is 329 ms. This is a server response measured on one day from one location; it is not Core Web Vitals.
- On 9 sites at least a quarter of images have no alt text; range 0–83%. On architecture sites images are the main content, so this is a direct gap for accessibility and image search alike.
- On 4 sites the median word count of crawled pages is under 200. On an image-led portfolio site that is not a defect by itself, but search engines and AI systems learn what a page is about from its text.
What we did not measure
- Orphan pages and full internal-linking architecture. The crawl is capped at 30 URLs; 14 sites hit the cap. On sites whose sitemap lists more than 1,000 URLs, a "in sitemap but not crawled" finding is an artefact of the cap and was not reported.
- Rankings, impressions, clicks. That is the sites' own Search Console data; we do not have it.
- Core Web Vitals. Needs field data; an external crawl does not provide it.
- JavaScript-loaded content. The crawler did not execute JS; a site that loads content only through JS may show low word counts and schema.
- Schema correctness, content quality, mentions in AI answers. None of these are within the scope of this research.
What it means for a firm owner
These 17 sites rank on Google, so the "found" layer is not entirely broken. But most findings show where even a ranking site can lose the next client: if the description shown in the search result is missing, clicks drop; if canonical points at the wrong version, the page folds in the search engine's eyes; if a bot filter cuts off Bing or AI crawlers, the site is absent from those channels.
None of the findings calls for a new website. All of them are the kind that can be fixed with the technical layer added under the existing design; which item can be done on which platform is a separate question, answered with two boxes on the service page for existing websites. If you want to see where your own site sits in this table with the same measurement, the Check-up report repeats this crawl for your site and your competitors.
Methodology
Sample: DataForSEO SERP API, Google organic, location_code 2792 (Türkiye), language_code tr, desktop, depth 30; three queries; 26 September 2026. Exclusion rules (on domain or URL): social networks, directories, news sites, list and guide pages, URLs more than two levels deep. One record per domain. 18 domains; one excluded as a professional chamber; 17 firm sites.
Crawl: site-auditor 0.2 (ClassyDesign's source-blind crawler), --max-urls 30 --concurrency 3 --render-js false --external-link-sample 0 --timeout 15. 3–51 seconds per site. Date: 26 September 2026.
Bot probe: One curl request per homepage, following redirects, with five user-agent strings; response code recorded. robots.txt rules were not applied in this probe; what is measured is the server's response to the user agent.
Counting rules: "On at least one page" refers to crawled HTML pages returning 200 on that site. The "business-describing type" list: Organization, LocalBusiness, Architect, HomeAndConstructionBusiness, Corporation. Image alt text: images with no or empty alt attribute / all images found.
Single-sample nature: One day, one location, one crawl. Results change as sites change; future repeats will be published as separate dated versions, not written over this page.
FAQ
01Which sites did this research examine?
17 firm websites: the 18 non-directory, non-social domains appearing in Google's top 30 organic results for three Istanbul architecture-office queries, minus one that is a professional chamber rather than a firm. Sites are anonymised in the report; names are not published.
02Do the findings represent all architecture firms in Türkiye?
No. This is a single-day snapshot of 17 sites ranking for three queries. It does not cover firms that do not rank, other cities or other dates. The proportions describe the sample, not the sector.
03Does having structured data mean the schema is correct?
No. We only counted the presence of a parseable JSON-LD or microdata block on a page and the types inside it. We did not measure whether the blocks are correct, complete or actually describe the business.
04Can this measurement be repeated?
Yes, with the same method: the sample rule, crawler settings and user agents are in the methodology; per-site raw rows are attached as CSV. Because sites change, results on another date may differ.
Raw data
Per-site raw rows (crawled URLs, indexable pages, canonical, meta description, H1, hreflang, image alt text, response time, word count, four bot responses, audit score), anonymised: CR-2026-004 per-site data (CSV). Every number on this page can be derived from that file. All findings are ClassyDesign's own measurement; none rests on an external source.