Advanced Scan Settings
The PerfBee Site Crawler offers a range of advanced settings to give you granular control over your website scans. These options allow you to tailor the crawl to your specific needs, whether you're focused on internal links, external validation, or optimizing crawl speed. These settings are configured inside a preset or directly in the scan form.
Excluding URLs
Sometimes you need to exclude certain sections of your site from the crawl. This is easily done using the exclude_urls setting. Provide a list of URL patterns to skip, such as /admin/ or /cart/. This helps to avoid crawling unnecessary areas and can save on credit usage.
Scanning Specific URLs
Instead of crawling your entire site, you might want to focus on a specific set of pages. Use the specific_urls setting to define a list of URLs to crawl. The crawler will only visit the URLs you specify, ignoring any links it finds on those pages.
Crawl Scope: Internal vs. External Links
By default, the crawler focuses on internal links. However, you can control its behavior regarding external links:
- Internal Links Only: The crawler will only follow links within your domain.
- Check External Links Toggle: This option, when enabled, validates external URLs, ensuring they're reachable and returning the correct status codes. Note that checking external links consumes credits for each external URL checked.
Crawl Speed Modes
PerfBee offers different crawl speed modes to accommodate various website structures and server configurations. The crawl_mode setting allows you to select from the following options:
- Auto: PerfBee automatically determines the optimal crawl speed.
- Polite: A slower, more respectful crawl, suitable for sites with strict rate limits.
- Balanced: A good balance between speed and resource usage.
- Fast: Crawls as quickly as possible, potentially consuming more server resources.
Ignoring Query Parameters
To avoid crawling duplicate content caused by query parameters (e.g., /page?color=red and /page?color=blue), enable the ignore_query_params option. This treats URLs with different query parameters as a single URL, preventing redundant crawls and saving credits.
Sitemap Crawling
Leverage your existing sitemap with the crawl_sitemap setting. When enabled, the crawler reads your sitemap.xml file to discover URLs, in addition to following links found on your site. This ensures comprehensive coverage of all your important pages.
Checking Images and CSS Resources
While not a direct setting, the crawler automatically checks the status of images and CSS resources linked on your pages. This is a default behavior and doesn't require any specific configuration. Ensure all your assets are loading correctly to avoid performance issues.
By using these advanced settings, you can fine-tune your Site Crawler scans to be more efficient, accurate, and tailored to your specific SEO and performance needs. Remember to monitor your credit usage, especially when enabling features that consume more credits, like checking external links.