Baiduspider
The crawler for Baidu, the dominant search engine in China.
Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)Official documentation: help.baidu.com
What Baiduspider does
Baiduspider indexes the web for Baidu Search, which handles the majority of search traffic in mainland China. Baidu documents a family of sub-crawlers (Baiduspider-image, Baiduspider-video, Baiduspider-news and others) that all match the Baiduspider token in robots.txt, with a couple of ad-related exceptions Baidu itself admits do not follow robots rules. It crawls plain HTML and is not documented as rendering JavaScript, so heavily client-rendered pages may index poorly on Baidu regardless of what you do. Crawl volume on sites outside China ranges from negligible to surprisingly persistent.
Why it visits your site
Standard search indexing for Baidu's results pages.
The bigger picture: search engine crawlers
Search crawlers remain the one category with a decades-old, well-understood bargain: they consume crawl budget and return ranked visibility. Everything about them is comparatively mature - published IP ranges, verification methods, granular directives, webmaster consoles. Blocking them is almost always self-harm, which is precisely why they need to be cleanly separated from AI crawlers in your data: a blanket bot block that catches Googlebot costs you your organic channel.
The complication the AI era added is that search indexes now feed AI features too - AI Overviews draw on Google's index, Copilot on Bing's. Operators have answered with control tokens that split AI use from search use of the same crawl. The practical upshot: your lever for AI concerns is usually a token, not a block on the search crawler itself.
Should you block it?
If any of your audience is in China, keep it - Baidu is how they will find you, and blocking it removes you from the biggest Chinese search engine. If you have no Chinese-market interest and the crawl volume bothers you, blocking it costs you little. Baidu documents robots.txt compliance for its main crawlers, so a robots.txt rule is the sensible first step.
Block via robots.txt
Add these lines to the robots.txt at your site root to ask Baiduspider to stay away:
User-agent: Baiduspider Disallow: /
Remember that robots.txt is a request, not a lock - compliance is voluntary, and enforcement happens at the server. Viz's AI Control pairs robots.txt rules with server-level 403 responses and then probes your site to verify the block is actually holding.
How to see Baiduspider traffic on your site
You have three windows onto it, each with a catch. Raw hosting access logs contain every Baiduspider request, but many managed WordPress hosts don't expose them, and when they do you are grepping text files by hand. A CDN dashboard sees the traffic at the edge, but it lives outside WordPress and speaks in totals, not in your site's terms. And JavaScript analytics - GA4 and friends - will never show it at all, because Baiduspider doesn't run scripts.
Viz's answer is the request log: filter to Baiduspider and read exactly which URLs it fetched and when, right inside wp-admin - then watch the same filter after any blocking decision to see whether the visits stopped, slowed, or kept coming.