Robots.TXT

WordPress implements virtual Robots.txt file used by search engines and other bots to understand website limits when it comes to accessing and crawling website content. For each user agent (used to identify different bots), you can add set of rules to control what is allowed and what it disallowed.

coreSecurity Pro includes Robots.txt feature that can add new directives to this virtual Robots file. Right now, these directives are for disallowing access to the website to many AI (Artificial Intelligence) and LLM (Large Language Model) bots.

If you have a real ‘robots.txt’ in your website root directory, you will need to manually add new directives to this file, and plugin will show directives to add on the Robots.TXT panel.

AI and LLM User Agent to Disallow

To learn more about the AI and LLM bots, what they do, additional information about each one, check out the Block AI Crawlers and Scrappers user guide, with information about blocking them.

Robots.TXT feature currently has settings to Disallow access to selected bots user agents. Bots are split into several groups: Scrappers, Crawlers, and Assistants. Every bot you enable here, will be added to the Robots.txt file, and it will be disallowed from accessing your website content.

Settings for Disallowing user agents bots
Settings for Disallowing user agents bots

Important Notices

  • There is no guarantee that bots accessing website will obey rules in the Robots file. Reputable companies run bots that obey the robots file. If you notice that disallowed bots continue to access your content, you can ban it in more forceful way using firewall.
  • User Agent value can be faked and spoofed, and any software access website can set any user agent they want, even an empty user agent.
  • For now, based on user experience, AI and LLM bots are behaving correctly, and are obeying the disallow rules.
Rate this article
0
0
2486

You are not allowed to rate this post.

Leave a Comment

0
0
0
0
0
0
0
0
0