Website check bot
Crawlspace checks a website when someone submits its URL. Its requests identify themselves
with crawlspace-sg in the user agent.
What it fetches
Each website check is short and bounded. It fetches the submitted page,
robots.txt, sitemaps, and a small sample of pages and images from the same
website. The bot does not run JavaScript, submit forms or log in.
Allow the Crawlspace crawler
An HTTP 403 response means the website refused our request. Security rules or bot protection can cause this even when the page opens normally in your browser. Crawlspace cannot complete browser challenges.
-
Check your hosting, firewall or CDN security logs at the time of the audit. Look for
requests whose user agent contains
crawlspace-sgand identify the rule that blocked or challenged them. - Ask your hosting or security provider to create a narrow exception for the crawler on the public pages you want checked. Keep protection enabled for other traffic and private pages. A user agent can be copied; it is an identifier, not proof of identity.
- Run the audit again and check that those requests receive the page content without a block or browser challenge.
Changing robots.txt alone will not remove a firewall block. If robots rules
also restrict the pages you want checked, review the crawlspace-sg group below
with your provider.
Robots rules
Pages and images the bot discovers follow the robots.txt group for
crawlspace-sg, or the * group when none names it, including
Crawl-delay. The submitted page is always fetched because a person asked for
it, as with other user-requested tools.
To stop the bot crawling beyond the submitted page:
User-agent: crawlspace-sg Disallow: /