User-agent: Googlebot Disallow: / User-agent: * # MediaWiki's Special: namespace is machine plumbing, not content. It cannot # rank, and on this fleet it was consuming the majority of the crawl budget that # should be reaching articles: on the SG server Baiduspider spent 1142 requests # on Special:CentralAutoLogin against 347 real article fetches in a 12h window, # with Special:RecentChanges, Special:Nearby, Special:NewPages and # Special:BookSources behind it. CentralAutoLogin is also answered with a 204 in # smart-mirrors.conf and stripped from the HTML, but robots.txt is what stops a # crawler asking in the first place. # # Deliberately scoped to Special: alone. The Wikipedia: namespace is NOT listed: # /wikipedia/zh-cn/Wikipedia:首页 is the mirror homepage, the page that actually # holds a Baidu ranking for 维基百科, and it is the single most valuable URL here. # # Both the wildcard and the explicit variant prefixes are given: Baidu documents # support for * in Disallow, but the literal prefixes cost nothing and mean the # rule still bites if a crawler's wildcard handling disappoints. Disallow: /wikipedia/*/Special: Disallow: /wikipedia/wiki/Special: Disallow: /wikipedia/zh-cn/Special: Disallow: /wikipedia/zh-hans/Special: Disallow: /wikipedia/zh-hant/Special: Disallow: /wikipedia/zh-tw/Special: Disallow: /wikipedia/zh-hk/Special: Disallow: /wikipedia/zh-mo/Special: Disallow: /wikipedia/zh-sg/Special: Disallow: /wikipedia/zh-my/Special: Sitemap: https://torrent-yopta.ru/sitemap.xml