How Does Indexing Work on the Internet?
The term indexing generally stands for the recording of information. On the internet, it refers to the inclusion of documents in the index of search engines.
In the process, similar to libraries, content is tagged with descriptors that allow it to be located specifically. While typical descriptors in libraries are author names or ISBN numbers, with search engines they are keywords.
In the field of search engine optimization indexing plays a central role. Webmasters or site operators can deliberately ensure that pages are included in the index more quickly. Sometimes, on the other hand, it makes sense to prevent the indexing of a website or parts of it.
In any case: only if a web page is in the index of a search engine will it also be found by it.
How Does Indexing Work on the Internet?
The indexing of web pages by large search engines such as Google or Bing is a complex process that is made up of several processes:
- Crawlers search the internet, read the source code of pages, and send it to the index.
- The retrieved content is then processed in a separate step and – provided Google classifies it as suitable – stored in the index (indexing). A crawled page is therefore not automatically indexed.
- The indexed pages then appear for relevant search queries.
In order to make the user experience during a search as positive as possible, search engines regularly optimize indexing. For this reason, and because both ranking factors and websites are constantly changing, the Google index is dynamic. New pages are added and the hierarchy – and thus the order of search results – changes. In addition, Google removes pages that massively violate its own guidelines from the index. As a result, these pages no longer appear in search results either.
The Indexing of a Web Page
When a new web page goes online, it is not automatically in a search engine’s index. Although Google’s crawlers constantly search the internet for new content, it can still take time until they index a page and it becomes findable.
The process can be accelerated by webmasters providing Google with a sitemap. This can be done in two ways: by referencing the sitemap URL via the Sitemap: directive in the robots.txt, or by actively submitting the sitemap to the Google Search Console.
A sitemap can be created easily with third-party tools. To have a single page crawled and indexed again, you now use the “Request indexing” function in the URL inspection tool of the Google Search Console. The formerly common sitemap “ping” no longer exists: this endpoint was switched off at the end of 2023 and has since returned a 404 error.
A single URL can be submitted via the URL inspection tool of the Google Search Console using “Request indexing”. This specifically triggers the (re)crawling and indexing of a particular page (instructions).
Preventing Indexing
Sometimes webmasters want to prevent pages from being indexed by the search engine or at least postpone it. There are several possible reasons for this:
- The page is under construction or in a relaunch. In that case, it is better if it does not appear in the search results until it is finished.
- These are admin access areas.
- The page is low-quality. In an online shop, this can be a category page with few products, on which visitors are unlikely to find what they are looking for.
- There are data protection or copyright reasons for “hiding” the page.
- The web page is intended exclusively for private use.
- There is a risk of duplicate content, which has a negative effect on SEO.
To prevent indexing, site operators or webmasters can proceed in different ways:
- They integrate the noindex meta tag into the HTML code of the relevant page or return a “noindex” header in the HTTP request.
- They inform crawlers via a robots.txt file about pages or files that they are not allowed to request.
- They store confidential content in a password-protected server directory. The Google crawler cannot access it.
To correctly mark up web pages with duplicate content, you insert canonical tags into the header of the pages. This is relevant, for example, in blogs when an article is displayed under several categories. A canonical tag prevents the inclusion of duplicate content in the index and a possible devaluation by the search engine.
Noindex or robots.txt: When Is Which Instrument the Right One?
Google advises against using a robots.txt file to keep web pages from appearing in search results. Because if other pages link to this page with descriptive text, it can still be indexed – then without a description in the search results.
robots.txt files, on the other hand, are useful for managing crawling traffic and avoiding overload, for preventing the display of video, image, and audio files in search results, or for blocking unimportant resource files. To avoid indexing, it is safer to fall back on noindex or password-protected directories.
robots.txt does NOT prevent indexing:
A page blocked via robots.txt can still end up in the index (and appear without a snippet) if other pages link to it – because robots.txt only controls crawling, not inclusion in the index. To reliably keep a page out of the index, the noindex tag (or a password-protected directory) is the right means. Crawl control (robots.txt) and index control (noindex) must therefore be clearly separated.
The central diagnostic tool for checking whether pages are actually indexed is the Search Console report “Page indexing”. It lists indexed and non-indexed URLs together with reasons (e.g. “Crawled – currently not indexed”, “Excluded by ‘noindex’ tag”). As a rough quick check, the “site:domain.de” query is also available, which, however, only provides an imprecise estimate.
Quellen
- Google, Build and submit a sitemap: https://support.google.com/webmasters/answer/183668#addsitemap
- Google, Ask Google to recrawl your URLs: https://developers.google.com/search/docs/crawling-indexing/ask-google-to-recrawl?hl=de
- Google, Introduction to indexing: https://developers.google.com/search/docs/guides/intro-indexing?hl=de
SEO & GEO check
Want more visibility – in search engines and AI answers?
We analyse your potential and show concrete next steps. Free and non-binding.