travel_exploreSite Crawler

A crawler that keeps running after you close your laptop.

PerfBee crawls your whole site on a schedule, gets through the bot walls that stop a desktop crawler, and traces every broken link back to the pages and the exact element that point at it.

  • check_circleNothing to install
  • check_circleCredits by the page, not the scan
  • check_circleCost estimate before you start
Broken URLs · example.comdetailed crawl
404/images/hero-banner.webpimage
arrow_outwardfound on 14 pages
404/collections/ss24-archivelink
arrow_outwardfound on 6 pages
503https://partner.example.com/specexternal
arrow_outwardfound on 3 pages
404/assets/legacy/theme.csscss
arrow_outwardfound on 41 pages

//*[@id="product-gallery"]/div[2]/img[1]

the element carrying the first broken image

Why not the desktop tool

A desktop crawler tells you about the day you ran it

It is a good tool. It also needs your machine awake, crawls from an office IP that large sites increasingly refuse, and forgets everything between runs — so a link that broke on Tuesday is only news if you happened to crawl on Wednesday.

schedule

It runs without you

Daily, weekly or monthly crawls on our workers. Nothing installed, nothing to keep open, no laptop that has to stay awake for six hours.

shield

It gets through

Three fetch strategies, escalating to a real headless browser when a page is refused — instead of a report full of 403s.

history

It remembers

Every crawl is kept and compared with the last, so you see what is new, what got fixed and what has been broken all along.

Findings you can act on

Every broken URL, traced back to what points at it

A list of 404s is not a task list. PerfBee records every (broken URL → page that links to it) pair it sees, so one dead asset referenced by four hundred product pages reads as one template fix, not four hundred tickets.

  • check_circleThe pages that reference itA drill-down listing the source pages, with the exact total and each URL’s status across recent scans.
  • check_circleThe element that carries itA DevTools-style XPath and the anchor text, so you can find the markup without hunting for it.
  • check_circleSorted by how much it mattersA priority list that ranks issues by impact rather than dumping them in crawl order.

/images/hero-banner.webp

404referenced from 14 pages
/products/dishwasher-9000
/products/dishwasher-9100
/collections/kitchen
/en/products/dishwasher-9000

+ 10 more source pages

Coverage

What a single crawl gives you

A fast scan checks status codes across the site. A detailed audit additionally writes a full technical-SEO record for every HTML page and computes the site-wide patterns.

link_off

Links and resources

Internal links, external links, images, stylesheets, scripts, and embedded video, audio and iframes — each checked and recorded with its HTTP status.

  • chevron_rightBroken internal + external links
  • chevron_rightImages, CSS and JS
  • chevron_rightVideo, audio, iframe and embed URLs
search

Technical SEO, per page

A detailed crawl writes roughly 45 fields for every HTML page — the crawl table an SEO team actually works from.

  • chevron_rightTitles, meta descriptions, H1–H6
  • chevron_rightCanonicals and robots directives
  • chevron_rightOpen Graph, Twitter cards, favicon
  • chevron_rightWord count, image alt coverage, page size
content_copy

Site-wide patterns

Problems that only exist across pages, not on any single one, computed once the crawl has the whole site.

  • chevron_rightDuplicate titles and meta descriptions
  • chevron_rightOrphan pages nothing links to
  • chevron_rightMixed content (HTTPS page, HTTP resource)
  • chevron_rightSchema.org structured-data inventory
difference

Change between crawls

Each recurring crawl is compared with the one before it, so you read the delta instead of re-reading the whole report.

  • chevron_rightNew, fixed and still-broken URLs
  • chevron_rightA Changes tab per scan
  • chevron_rightPer-URL status history across recent scans

Checking that required content is still present on every page — labels, manuals, spec tables, schema — is a separate job. See Content Monitor →

AI search readiness

Can ChatGPT, Claude and Perplexity actually read your site?

AI answers cite pages their crawlers could fetch and are allowed to quote. Every detailed crawl checks both, on every page — so a firewall rule or a stray noindex does not quietly take you out of AI answers. Included in the crawl, no extra credits.

  • check_circleEvery AI crawler, checked two waysrobots.txt is evaluated for 14 AI crawlers against every crawled URL, and sample pages are fetched as each crawler next to a normal browser — so a firewall rule that answers ChatGPT with a 403 shows up.
  • check_circlePages an AI answer is allowed to quotenoindex and nosnippet in the meta tag or the X-Robots-Tag header, max-snippet limits, and pages whose content only appears after JavaScript runs — most AI crawlers never run it.
  • check_circleStructured data that agrees with itselfDuplicate FAQPage blocks and JSON-LD that fails to parse, plus Organization markup and sameAs profiles on the homepage — what lets an AI connect your brand name to your domain.
  • check_circleAn alert when access is lostIf an AI search crawler reached the site on the last crawl and cannot now, you are notified. Blocking training crawlers is treated as your choice, not an error.

What it cannot tell you: whether an AI answer actually cites you — it checks the conditions for being cited. A block tied to a crawler's verified IP addresses, which Cloudflare can apply, cannot be reproduced from outside; when your site is behind Cloudflare, the report says so.

AI crawler access · example.comdetailed crawl

OAI-SearchBot

ChatGPT search

Blocked · 403

Claude-SearchBot

Claude search

Reachable

PerplexityBot

Perplexity answers

Challenge page

Googlebot

AI Overviews, AI Mode

Allowed

GPTBot

Model training

Opted out

18 pages need JavaScript to show content

25% of the homepage text is in the raw HTML — the part most AI crawlers see

1

TLS-impersonating client

Requests carry a real Chrome fingerprint, which is what most filtering actually inspects.

2

Plain HTTP client

A straightforward retry for hosts that dislike the first client.

3

Headless Chromium

On a 403, 429 or 503 the request is escalated to a managed real browser.

Challenge and interstitial pages are recognised from the response body and kept separate from healthy pages in the report.

Getting through

Large sites block crawlers. This one escalates.

The reason a desktop crawl of a big commercial site comes back full of 403s is that the site is refusing it. PerfBee tries progressively more expensive strategies per request, so the crawl finishes with real status codes instead of a wall of blocks.

It stays polite while doing it: robots.txt is obeyed on your own domain, request rates adapt to how fast the server responds, and you can verify a site you own to unlock a faster lane on it.

verified_user

Checking your media never counts as a view

A crawler that fetches every embedded video to see whether it still exists will quietly add thousands of plays to your analytics and your vendor’s billing. PerfBee checks video, audio, iframe and embed URLs with a HEAD request only, from a bot user agent, with no Referer and no credentials attached.

The body is never downloaded, so nothing registers as a view, a play or an impression — and no tracking parameter from your page travels with the request.

Scope

Decide what gets crawled before it costs you anything

On a large catalogue the difference between a useful crawl and an expensive one is scope. PerfBee reads your sitemap first and shows the page count and the credit cost for your plan, before the crawl starts.

account_tree

Seed from your sitemap

Reads /sitemap.xml, follows sitemap-index files recursively, and queues every URL it lists — including pages nothing links to.

filter_alt

Exclude what you do not want

Wildcard patterns like /tag/*, *utm_source* or /blog/* are dropped before a single request is made, so faceted and paginated noise never costs you credits.

checklist

Or crawl an exact list

Paste a set of URLs, or a wildcard pattern to match against your sitemap, and PerfBee checks only those — no crawling at all.

tune

Save it as a preset

Scope, depth, page cap and which resource types to check, stored as a reusable configuration and applied to manual and scheduled crawls alike.

After the crawl

A report you work through, not just read

Findings on a big site arrive in the thousands. The results view is built for triaging that down to the handful of template fixes underneath.

filter_list

Filter and save the filter

Status chips, full-text search and an include/exclude sidebar. Save a combination you use often as a preset.

workspaces

Group the noise

Collapse the table by source page, by external domain or by status code to see the shape of the problem.

done_all

Mark things resolved

Select in bulk, acknowledge with a note, and keep the rest of the list as your queue.

refresh

Recheck a single URL live

Re-request one URL on the spot with fresh SEO and technical analysis, without re-running the crawl.

download

Export what is on screen

CSV from every tab, honouring the filters you applied — plus a 22-column SEO crawl export.

share

Share without an account

A public link to a completed report that expires after 7, 30 or 90 days and can be revoked.

Want to check one page first?

These run instantly, with no account. They check a single URL — the crawler is what covers the whole site.

Site Crawler FAQ

The questions people actually ask before moving a crawl into production.

How is this different from Screaming Frog?expand_more

Screaming Frog runs on your machine, on demand, from your IP. PerfBee runs in our infrastructure on a schedule, keeps every crawl so it can show you what changed, and escalates through three different fetch strategies when a site blocks it. The honest trade: Screaming Frog is still the better tool for a deep one-off audit on your desk — it has log-file analysis, GSC and GA integration, JavaScript-render configuration and custom extraction that we do not match. PerfBee is the better tool when the crawl has to keep happening without you.

What does it cost to crawl a large site?expand_more

Crawling is metered per page, not per scan: 1 credit per page for a fast scan and 4 per page for a detailed audit. JS rendering is charged at the detailed rate. Verified sites you own are cheaper. Before you start, PerfBee reads your sitemap and shows the estimated page count, the cost for your plan and whether your remaining credits cover it. The free plan includes 500 credits a month, which is one 500-page fast scan.

What happens when a site blocks the crawler?expand_more

PerfBee first requests pages with a client that reproduces a real Chrome TLS fingerprint. If that is refused it retries over plain HTTP, and on a 403, 429 or 503 it escalates that request to a managed headless Chromium. Challenge and interstitial pages are recognised from the response body and reported separately from healthy pages, so a wall of Cloudflare pages does not show up as a wall of 200s.

Will checking my video and audio links inflate their view counts?expand_more

No, and this is deliberate. Media URLs — video, audio, iframe and embed sources — are checked with a HEAD request only, from a bot user agent, with no Referer and no credentials. The body is never fetched, so no view, play or impression is recorded and nothing reaches your analytics.

Can it crawl pages behind a login?expand_more

No. Cookies are disabled for the whole crawl, so staging environments and member areas that require a session cannot be reached today. Sites protected only at the network level can be crawled if you allowlist us.

Does it render JavaScript?expand_more

Yes, as an option. With JS rendering on, every page is loaded in a headless Chromium and the rendered DOM is what gets parsed, so content injected by a framework is seen. It is significantly slower and charged at the detailed rate, so it is normally used for a targeted crawl rather than a nightly one.

Does it check whether AI search can read my site?expand_more

Yes, on every detailed crawl, at no extra cost. robots.txt is evaluated for 14 AI crawlers — OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Meta and Common Crawl — against every crawled URL, and sample pages are fetched as each crawler next to a normal browser to catch firewall blocks. Every page is checked for noindex, nosnippet and max-snippet in the meta tag or the X-Robots-Tag header, content that only appears after JavaScript runs, and broken or duplicated structured data. It does not tell you whether an AI answer actually cites you, and a block tied to a crawler's verified IP addresses cannot be reproduced from outside.

How often will I hear from it?expand_more

Crawls can be scheduled daily, weekly or monthly, and each run is diffed against the previous one so the Changes tab is ready when you open it. When a crawl finishes you get an in-app notification, and an email when it found broken links or failed — set per event, on plans that include email notifications. You are also told when an AI search crawler loses access to the site. Emailed summaries are still available as a digest — weekly by default, monthly optional, daily if you turn it on. There is no Slack integration yet.

Can I share a report with someone who has no account?expand_more

Yes. Any completed scan can be given a public link that expires after 7, 30 or 90 days, viewable without logging in, and revoked at any time. Results also export to CSV from every tab, filtered exactly as you left them on screen.

How large a site can it handle?expand_more

Crawls run on our workers with a six-hour ceiling per run and a page cap you set yourself, and tens of thousands of URLs per crawl is normal. For a very large catalogue the practical approach is several scoped schedules — one per section — rather than one crawl of everything.

Your website has something broken — or missing — that you don't know about yet.

Broken links, slow pages, and the labels, manuals and banners that silently vanish when a template changes — PerfBee finds them all and tells you exactly how to fix them.

  • Free forever, no credit card
  • Your first audit in minutes
  • Built for large catalog sites