Technical basics

Crawling and indexing: know the difference

A page can be easy for a person to open and still have instructions that affect search engines. Before changing those instructions, decide whether the page should be public, eligible for search, or genuinely private. Those are different goals.

SEO Fieldwork · Prepared October 8, 2026 · AI-assisted educational draft

Keep this in mind

robots.txt controls crawling, not reliable removal from search. A crawler must be able to access a page to see a noindex rule. Neither is a privacy barrier.

A URL is not the same as an indexed page

A crawler may discover a URL through a link or sitemap, request its content, and then analyze it for indexing. These steps can happen at different times. A page returning successfully in your browser proves that you can access it; it does not prove that Google has indexed it.

If you have access to the site’s Search Console property, URL Inspection can help investigate how Google sees a page. This publication cannot perform that inspection and the workbench does not connect to your account. Keep the difference clear when you record a problem: “I can open the page” and “Google reports the page indexed” are different observations.

Do not assume a technical setting will create a ranking result. Eligibility and discovery are foundations. They are not guarantees of inclusion or of placement for a particular query.

What robots.txt does

A robots.txt file gives crawlers instructions about which URLs they may request. Respectful crawlers follow supported rules, but the file does not enforce access control. Google describes it primarily as a way to manage crawling traffic, not as a reliable mechanism for keeping a web page out of search.

If a disallowed URL is linked from elsewhere, Google may still discover and show that URL without crawling its content. That is why “Disallow” is not a synonym for “remove from the index.” A robots.txt file is also public. Do not list confidential information in it and expect the file to hide that information.

Avoid blocking resources that a crawler needs to understand the page, such as important styles or scripts. Before changing a site-wide rule, review the paths it matches. A small typo or an overly broad rule can affect more pages than the one you intended to change.

What noindex does

A noindex rule can be supplied in an HTML meta tag or an X-Robots-Tag HTTP response header. Google supports it as an instruction not to index the resource. The crawler needs to access the page and see the rule before it can act on it.

If you also block that page in robots.txt, the crawler may never see the noindex instruction. Do not combine these settings without understanding their interaction. Google does not support placing a noindex instruction in robots.txt as a replacement for the meta tag or header.

Changes may take time to be processed because the page needs to be revisited. A noindex rule is not immediate removal, and it is not password protection. Anyone who can access the URL may still read the page. Sensitive content needs real access controls, not a search visibility setting.

Choose the setting from the page’s purpose

A public class guide intended for search needs to be accessible and should not carry an unintended noindex rule. A temporary public draft may need noindex, but it is still a public draft. A private member record needs appropriate authentication and access controls. These examples illustrate different goals, not a complete security design.

Many website editors expose a setting such as “discourage search engines” or “hide from search.” Check what it actually emits. A checkbox label is not evidence that the published HTML, HTTP headers and robots.txt all have the intended behavior. Hosting headers can also add a noindex rule that your editor does not show.

For a small site, keep a note of what should be searchable and why. If you are unsure how a setting affects the whole website, avoid experimenting on a live site without a tested way to undo the change.

Try it: check one public page

Open the important page, verify that its content and links work, and confirm that its intended visibility is public. Review your editor’s search-visibility settings. If you know how to inspect HTML and headers, check the robots meta value and any X-Robots-Tag header. Also review the actual robots.txt at the site root.

Record observations rather than guesses: the URL, the response you saw, the rule you found and the page’s intended purpose. If you use URL Inspection on your own property, keep that observation separate from the browser check. You may need help from your hosting administrator to change an HTTP header.

The SEO Fieldwork checklist is a manual planning aid. Checking a box does not inspect these rules or change them for you. When the technical basics are sound, return to the reader’s question and improve the page itself. That is the center of the work.

Documentation behind this guide

Original explanations informed by these primary documents. Examples and worksheets are our own, not Google scoring methods.

  • Google Search Central: SEO Starter Guide
  • Google Search Central: Introduction to robots.txt
  • Google Search Central: Block Search indexing with noindex

Primary documents reviewed October 8, 2026. Source destinations are not linked in this edition. Read our sourcing approach.