Skip to content

For site owners

EnrichedYakBot

EnrichedYakBot is the crawler behind EnrichedYak. It reads a small number of public pages on a company's website — contact, team, about, locations, careers — to build a business profile with the source of every fact.

How to recognise it

Every request sends this User-Agent:

EnrichedYakBot/1.0 (+https://enriched.yakware.com/bot)

It never changes its User-Agent, never uses residential proxies, and never renders pages in a disguised browser. If a request does not carry this header, it is not us.

What it does

  • Reads robots.txt first (cached for up to 24 hours) and follows it for the token EnrichedYakBot, falling back to *.
  • Fetches at most 12 pages per company per visit, plus robots.txt and one sitemap.
  • Waits at least one second between requests to the same host, or longer if you set Crawl-delay.
  • Sends conditional requests (If-None-Match, If-Modified-Since) on return visits so unchanged pages cost you almost nothing.
  • Treats a server error on robots.txt as “do not crawl” until it can read it.
  • Stops at the first 401, 403 or 429 from your homepage.

How to limit or block it

Block it entirely:

User-agent: EnrichedYakBot
Disallow: /

Or slow it down and keep it out of a section:

User-agent: EnrichedYakBot
Crawl-delay: 10
Disallow: /internal/

People listed on your site

We store names and titles only as your site publishes them, and email addresses only when they are on your own domain. Anyone can ask us to remove or correct their details at /privacy/request; once confirmed, the value is suppressed and never collected again.

Contact

Questions or a problem with the crawler: care@yakware.com.