Website check bot

Crawlspace checks a website when someone submits its URL. Its requests identify themselves with crawlspace-sg in the user agent.

What it fetches

Each website check is short and bounded. It fetches the submitted page, robots.txt, sitemaps, and a small sample of pages and images from the same website. The bot does not run JavaScript, submit forms or log in.

Allow the Crawlspace crawler

An HTTP 403 response means the website refused our request. Security rules or bot protection can cause this even when the page opens normally in your browser. Crawlspace cannot complete browser challenges.

  1. Check your hosting, firewall or CDN security logs at the time of the audit. Look for requests whose user agent contains crawlspace-sg and identify the rule that blocked or challenged them.
  2. Ask your hosting or security provider to create a narrow exception for the crawler on the public pages you want checked. Keep protection enabled for other traffic and private pages. A user agent can be copied; it is an identifier, not proof of identity.
  3. Run the audit again and check that those requests receive the page content without a block or browser challenge.

Changing robots.txt alone will not remove a firewall block. If robots rules also restrict the pages you want checked, review the crawlspace-sg group below with your provider.

Robots rules

Pages and images the bot discovers follow the robots.txt group for crawlspace-sg, or the * group when none names it, including Crawl-delay. The submitted page is always fetched because a person asked for it, as with other user-requested tools.

To stop the bot crawling beyond the submitted page:

User-agent: crawlspace-sg
Disallow: /