dizzysoft

web development for search engine optimization

Crawler Record Plugin for WordPress

Before your customers can find your website in search results, the search engine needs to be able to see your website. This is true whether your customers are searching in an LLM (such as ChatGPT, Claude, or AI Mode) or a traditional search engine. When was the last time one of these systems viewed your website? That’s what this WordPress plugin can tell you.

  • What is a user-agent and why should I care?
  • What the Crawler Record tells you about your site
  • What this doesn’t tell you about your site
  • How you might use the Crawler Record plugin
  • Plugin Features
  • What can you do with the Crawler Record data?
    • It’s been a few weeks and a crawler has not visited my site at all.
    • How do I get crawlers to come to my website more frequently?
    • I have prevented a certain crawler from visiting my website using the robots.txt file, but it’s still coming.
    • Commonly encountered 3xx and 4xx status codes.

What is a user-agent and why should I care?

A user-agent (also called a “crawler”) is a computer program designed to access a website. A user-agent might be an individual’s web browser or a search engine spider. Each user-agent has a unique identifier that tells the web host (or CDN) who it is.

When it comes to website marketing (especially for search and LLM strategies), how a platform’s user-agent interacts with your website matters. Once a user-agent has viewed your website, it can add your site to its index. Once you’re in the index for this tool, it can draw on what it’s discovered from the pages on your site and potentially show your site to people looking for what you offer.

Unfortunately, visits from crawlers can be expensive for your web host, so they might be preventing them from coming to your website. Some CDNs (I’m looking at you, Cloudflare) might prevent them from reaching your site as well (presented as protecting intellectual property, but really it’s an expense to them, too). It can be hard to tell whether these systems are preventing platforms from accessing your site (and preventing them from serving your site to people who might be interested). That’s the main problem Crawler Record is trying to solve– but it offers several other advantages along the way.

Of course, the crawl (or viewing of pages) doesn’t mean these systems will or want to show your site- just that it can access it. If you want to show up on these sites, that’s called “search engine optimization” (SEO) and involves much more than this plugin can do. This data demonstrates that the first step to SEO is indeed possible. The same is true if you want your website to be shown in LLMs- the first step is to ensure their user-agent (or “crawler”) has viewed your website.

Since the goal of this plugin is to identify search and LLM crawlers, it is checking for these particular user-agents to show up:

What the Crawler Record tells you about your site

  1. The last time a particular user-agent visited your website. If it has accessed your site, you might be in this tool’s index. This plugin only tracks bots after you’ve installed the plugin, so you might have to wait several days to see any information.
  2. If a particular user-agent has not visited your website, you can check whether your site allows it to view your website by checking your robots.txt or robots meta tag. This plugin will inform you if something on your site is preventing the agent from accessing it, but it cannot determine whether your host or CDN is blocking access. This matters because some web hosts and CDNs prevent bots from accessing your site. If they do, these platforms cannot serve you to people searching for what you have to offer. If, after waiting several weeks, certain bots still aren’t accessing your site, check with your web host, CDN, or developer.
  3. Learn which pages each user-agent visited on your site and when (this is throttled to prevent overwhelming your system or database).
  4. Are these user-agents encountering problems when viewing a page on your website?
    1. If a user-agent encounters a 4xx header code (such as 404), it will not be able to index that page- because it does not exist! If that’s what you intend, great! If you want it to view that page, then you will need to fix it.
    2. If a user agent tries to view a page and gets a 3xx header code (such as a 301 or 302 redirect), make sure that is the expected result.
    3. You want a user-agent to encounter a 2xx header code (like 200). That tells the agent that it successfully accessed that page. Just make sure it’s accessing the pages you want it to- and not those you don’t (such as those containing personal information or behind paywalls).
  5. View a particular page of your site, and this plugin will tell you the last time one of these agents viewed that specific page.

How is this different from Google Search Console or Bing Webmaster Tools?

If I’m concerned about how Google and Bing view my site, I’d check there before using this plugin. Those two tools (Search Console and Bing Webmaster Tools) are excellent, free tools you should use.

For now, however, LLMs do not have this type of data available to website owners. This plugin hopes to fill that gap.

Can’t you get this information (and more) from your web server’s logs?

Yes, you can. Unfortunately, most web hosts that support WordPress don’t allow log access. This plugin replaces that missing information. However, if you can get into your server’s logs, you’d learn a lot more than this plugin can tell you!

What this doesn’t tell you about your site

  1. That your site is “optimized” for SEO, LLMs, GEO, AIO, SAO, AEO (or whatever you want to call it). Just because a platform’s user-agent has visited a page on your site, it does not mean that page is in their index. If it is not in the index, then clearly you need to improve it.
  2. That people are visiting your website. This information reflects search engines or chatbots that are looking for opportunities to learn about the information on your site so they can send people to it. Just because a bot came, however, doesn’t mean people come. That’s what Google Analytics is for.
  3. That your host or CDN is preventing this user-agent from accessing your site only shows whether or not an agent has accessed it. If it’s been a while (a month or so) and one of these agents hasn’t accessed your site (and there’s nothing in your robots.txt/meta preventing it), you should investigate further.

How you might use the Crawler Record plugin:

  1. Can search systems view your WordPress website? If so, great: you’re capable of getting served by a search engine or LLM- if the pages on your website are deemed good enough quality. If not, you need to fix the problem.
  2. When was the last time a search engine visited your website? This information is valuable after launching a new website to learn how quickly it will be re-indexed.
  3. Are search systems consistently viewing your website? How often should they visit? It depends- more popular sites are visited more frequently. If a specific user-agent hasn’t visited in over a month, you might look into it.
  4. Are there search systems you don’t want to view your website, but are viewing it anyway? For example, most websites that serve ads (and generate revenue from page views) may not want an LLM user-agent to index their site. Unfortunately, some LLM agents still come through.

Plugin Features:

  1. From the admin, in the left column, select “Crawler Record” to learn:
    1. When crawlers last visited your website. Click those buttons to limit the following data to recent, not recent, or never-visited crawlers.
    2. Crawler activity such as
      1. Visits by crawlers in the last 28 days
      2. Unique URLs in recently recorded visits
      3. How many times crawlers landed on the correct page (determined by a 2xx header code). Click that button to see which pages by which crawlers (also accessible from the sidebar.
      4. How many times crawlers encountered a 3xx header code because a page was redirected. Click that button or the link in the sidebar to see this list.
      5. How many times a crawler hit a 4xx header code. You can also access this information by clicking on that button or from the admin sidebar.
    3. If you’re viewing all crawler or recently active crawler data (based on the top filter), you’ll see:
      1. The platforms whose crawlers visited your site the most in the last 28 days.
      2. The crawler activity on your website in the last 28 days.
      3. A chart representing response codes for crawlers’ visits.
      4. The most frequent URLs in the latest recorded visits.
    4. If you’re viewing all crawlers or not recently seen crawler data (based on the top filter), you’ll see:
      1. Crawlers not seen recently and when they stopped coming.
      2. A chart representing response codes for crawlers’ visits.
    5. A list of every user-agent this plugin tracks, organized by the platform (Google, Bing, ChatGPT, etc). If you filter your data (at the top of this admin page) by recent, not recent, or never-visited crawlers, this list will be filtered too. With this list, it will tell you:
      1. The last time and last page each crawer/user-agent has accessed.
      2. The HTTP status code with that visit. Learn more about these status codes from the sidebar or the buttons at the top of this dashboard.
      3. If a robot directive is preventing one (or all) of these agents from accessing your site.
      4. Click the manager for each user agent to learn more about these bots.
      5. Click on “View Recent Pages” to see which pages the crawler has recently visited.
  2. Visit any page and learn:
    1. If you’re viewing a page or post from the front-end, look at the admin bar. Hover over “Crawler Record” to see all managers of relevant user-agents. Hover over a manager to learn the last time each of these specific user-agents accessed this particular page.
    2. Edit the page to access a list of all monitored user-agents, sorted by the last time they accessed this page or post. From this, you will also learn if any robot directives are preventing one of these bots from accessing this specific page.

Please note:

  • The information about each user-agent (both “Last Seen” and “Last Page”) is only valid from when you installed this plugin. It cannot determine visits from any user-agent before that time. To avoid overwhelming your WordPress system, this data is throttled.
  • This plugin detects whether your WordPress install is blocking one of these agents, but not whether your host or CDN is blocking this user-agent.

What can you do with the Crawler Record data?

It’s been a few weeks and a crawler has not visited my site at all. What should I do?

First of all, it’s okay that not every crawler has visited your site. Focus on the platforms (not necessarily individual user agents) that matter most. I’d suggest those include Google, ChatGPT, and Claude. Whether or not any crawler can access your site comes down to three factors: permission, ability, and desire.

Are you permitting crawlers? The first thing to check is your robots.txt status in the list of platforms and user-agents at the bottom of the plugin’s dashboard- are you allowing them to access your site in the first place? If not, allow them!

Do crawlers have the ability to access your site? To ensure Google can access your site, verify it in Google Search Console. If Google can’t access your site, you will learn about it there.

A second check you can do in Google is to use the site: operator. That means searching Google for site:yourdomain.com, which shows all the pages Google knows are on your website.

Unfortunately, ChatGPT and Claude don’t have equivalent systems, so you have to work from a different angle- ask them. If they cannot, ask them to tell you why they cannot,

Do crawlers desire to access your site? Just because you permit them and they can access it doesn’t mean they want to come to your website! Why would they want to look at your website? Does anyone else talk about your website? Does anyone else mention your website on theirs? If not, these platforms might not care enough to visit. This will involve some digital PR or link building- but be careful.

How do I get crawlers to come to my website more frequently?

There’s no real advantage to crawlers coming more frequently to your website. What you want to know is that they come to see new pages on your website so they can consider adding them to their respective index.

One thing you can do to help crawlers know something’s new on your website is by using the IndexNow WordPress plugin. This tells platforms that there’s something new to look at on your website. It doesn’t guarantee they will come, but it does give them more information.

The biggest factor in whether crawlers want to visit your website more often is their desire (as I mentioned earlier).

  • Is there something new and worth seeing, or is it the same old content?
  • Are your pages unique and helpful? Are you saying anything new, or are you just rephrasing what everyone else has already said? If you’re using AI to write your website’s content, it might seem helpful, but it won’t necessarily be unique enough for crawlers to care.
  • Does anyone else think you’re an expert in your field? This is where digital PR can help. However, never pursue a link building campaign without understanding Google’s web spam guidelines first (even if you’re focusing on LLMs).

Will llms.txt help? No platform has accepted this standard. It won’t hurt you, but it probably won’t help either. Will creating a markdown version of each page help? It’s not necessary. In fact, most LLMs will read your page and create their own markdown anyway. If you’re running WordPress, your site is already easily read by most user agents- so even if crawlers wanted to see this, it wouldn’t be necessary.

Will Schema on my web pages help? It couldn’t hurt, but it’s not the panacea everyone thinks it might be. Be reasonable but don’t break your back on it.

I have prevented a certain crawler from visiting my website using the robots.txt file, but it’s still coming. What should I do?

There are certain pages on any website you don’t want crawlers to visit. For instance, admin pages, private information, or paywalled content. If a crawler visits anyway (based on the 2xx report in Crawler Record)…

  • Remember: the robots directive (whether robots.txt or a meta tag) is a suggestion, not a command. Some user-agents (as described in the admin Dashboard) ignore the robots directives because they are sent because of a user-generated (human) initiative.
  • The only sure-fire way to prevent a user agent from accessing anything on a website is to put it behind a server-side password. That might be too draconian for your use, but it’s the only guaranteed solution.

Commonly encountered 3xx and 4xx HTTP status codes.

Common 3xx HTTP status codes you might encounter- and what to do about them:

  • facicon.ico is the standard request for your website’s favicon. If you have a favicon installed on your site (and you should), it should redirect to the proper location of your favicon file using a 302 redirect. This is one instance where you do not want to use a 301 redirect.
  • Sitemap redirects. There are several valid filenames for sitemap files (which I’d recommend you have on your website): sitemap.xml, sitemap.txt, sitemap_index.xml, sitemap.xml.gz, etc. Redirects are fine as long as they end up at the correct filename for your particular sitemap. You’re seeing these because a platform is looking for your sitemap in different locations, which is good. You can help a user-agent find your proper sitemap filename by adding the reference in your robots.txt file or submitting a sitemap through Google Search Console or Bing Webmaster Tools.

Common 4xx HTTP status codes you might encounter- and what to do about them:

  • ads.txt: if you’re not running ads on your website and this comes up as a 404 error, that’s okay. You want the crawlers to see this 404 error because you’re not running ads. However, if you’re running ads, check with your ad platform to see whether you need this file and what it should contain.
  • sitemap.txt is a valid format for a website’s sitemap, but you might be using a different format (and I hope you are). If you find a user-agent is looking for your sitemap.txt, you might set up a 301 redirect to help it see it’s not there, but point it to the correct filename for your sitemap.

Need Help?

If you need support for technical issues with the plugin, please reach out via the plugin page on WordPress.org (accessible from the plugin).

If you need help understanding why bots aren’t coming to your site- or optimizing your site for search engines or LLM chats, please reach out to my agency: Reliable Acorn provides internet marketing consulting and has expertise in search engine optimization.

Tags:

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Examples

Recent Comments

Topics