SEO Basics & Strategy

How does a search engine work?

How does a search engine work?

To optimize a website for search engines, it is essential to know how a search engine works.

And that is exactly what I would like to explain to you in this blog article.

Search engines explained

Search engines (also called: general search engines, web search engines, universal search engines, algorithmic search engines) aim to cover the web as completely as possible. Due to the size of the internet, however, this is not possible.

Web search engines focus on specific content: documents that are stored on web servers. The links between the various documents create a network (World Wide Web).

unsorted data and websites

Search engines like Lycos worked from the very beginning according to the same system as Google does today: the documents on the web are captured by following the links between them.

Search engines are those that capture the content of the web by following links, through so-called crawling. Crawling starts with known documents and follows the links. This discovers new documents, whose links are then followed as well. In this way, an (incomplete) representation of the documents on the internet is created, which is then made searchable.

Important: YouTube is not a true search engine according to the definition, because YouTube only allows a search for the videos on the platform itself and does not represent videos across the entire World Wide Web.

Other types of search engines:

  • Specialized search engines: They deliver only documents for specific websites, such as search engines for children (e.g. Frag Finn)
  • Meta search engines: For a search query, the results are gathered from one or more “real” search engines. Example: Metager.de – here a large number of different search engines are queried.
  • Web directories: Here the websites are recorded manually by humans and stored in the index – a major difference from current search engines like Google or Bing!

How search engines work

First of all: when we search with a search engine, we never search the content of the web itself, but always draw only on a prepared copy of the World Wide Web. As already mentioned, the crawlers move from page to page and follow the links there in order to find further documents. The documents found are first stored in the so-called “local store”. This provides the raw material for the index. Through a so-called indexer, the raw material is prepared so that it can be searched efficiently.

A kind of representation is created for each document, and categories are set up so that they can be found quickly. The searcher has the role of intermediary between the stored information and the user.

How does a search engine work
Fig.: How a search engine works, source: Prof. Dr. Dirk Lewandowski, Suchmaschinen verstehen, page 31

Obtaining content: How does the crawler work?

The World Wide Web (WWW) consists above all of documents in HTML format, which have a unique address (URL) and are connected to one another via links.

Through so-called crawlers, search engines capture the documents on the WWW and their links to one another. The crawlers follow these links and try to find as much as possible on the WWW in the process. Unfortunately, this goal can never be achieved, because there are pages that do not want to be captured by search engine crawlers. This is referred to as the deep web – these are pages that are protected by a password or that do not have appropriate links.

Missing links are generally a problem for the crawlers. New websites that have no links at all cannot be found. Alternatively, search engines find new information via “feeds”. Through these, information is communicated to the search engines. For example, online shops provide search engines with their product catalog via a feed in the form of XML documents. Normal websites can also use an XML sitemap to show search engines which content can be found on their website.

  • Digression:
  • Website = a complete domain
  • Web page = a specific page on a domain

Crawling must take place regularly, otherwise the search engine will not find any new content or updated texts, nor any new links. Heavily linked websites and popular pages are visited more often, since the update intervals here are frequently shorter (example: news websites).

A crawler can only read out texts and the structural information in the source code of a website. Documents with many graphic elements, such as images and videos, and only little text can therefore be understood by the crawlers only with difficulty. For searching and capturing specific files, such as videos and images, there are crawlers that obtain content for so-called specialized search engines, such as Google Image Search or Google Shopping. In image search, for example, the crawler can recognize the image in the HTML document and understand or interpret it with the help of additional information, such as the title of the image, the surrounding text and the alt tag.

Raw data prepared: How does the indexer work and how is the index structured?

The indexer has the task of preparing and organizing the raw data from the crawler's local store so that it can be processed efficiently for search. From the multitude of documents, an index is created that constitutes a representation. When searching, it is not the document itself that is searched, but the representation (index) created by the indexer.

The system of syntax analysis, also called parsing module breaks down all the documents found into units, such as words and word stems, marks them and stores them. Analogous to a book index, where terms in the book are listed with a page number, documents are listed in the index according to specific terms. This is referred to as an inverted index – you do not have to read through the entire book to arrive at a particular term, but can turn directly to the relevant place. The search engine index is structured in a similar way. An inverted index and a book index differ, however, in that a book index does not always contain all the words that can be found in the book, but only those that are relevant to the section of text. With a representation, nothing can be found for a search query for a particular term if that term is not present in the representation. In this respect an inverted index differs from a book index.

Due to the huge mass of documents on the WWW, there are distributed indexes in order to provide search results as quickly as possible for the user – there is not one large index, but several distributed systems. In this way, the indexes can be searched efficiently and kept up to date.

The principle was originally established by Google with the MapReduce algorithm established; today the large volumes of data run over successor systems, but distributed crawling and indexing across many computers remains at the core of the architecture.

How the search engine understands search queries: The searcher

In order to be able to present the users of the search engine with suitable results, an intermediary is needed, the searcher. This matches the search query against the prepared data holdings of the search engine. The list of results is sorted and displayed according to the specifications and ranking factors of the search engine. The searcher therefore has the task of understanding the search query (the user's search intent) and showing the user the appropriate correct results.

In order to be able to serve the user relevant results, especially for complex search queries, search engines make use of query interpretation. The search query is linked with contextual information, such as the user's search history, their clicks on certain results, the paths of other users with a similar search history and much more. Through this interpretation, also called query understanding it is possible for the searcher to display relevant results for complex search queries.

On top of the classic structure of crawler, local store, indexer and searcher, search engines like Google today add an additional AI layer: AI Overviews deliver generative answers directly at the top of the SERPs, the Knowledge Graph provides structured knowledge, and ML components such as RankBrain, BERT and MUM help the searcher to better understand even complex or ambiguous search queries.

In addition, LLM-based searches such as Perplexity or ChatGPT Search have become established, using a language model as a core component and drawing on a classic web index only as a supplement. The classic crawler-indexer-searcher model remains the foundation – but it is noticeably extended by this additional layer.

Summary

Even though a search engine aims to find all possible documents on the web and make them available to the user, this is not possible, because on the one hand the mass of documents is extremely large, and on the other hand there are pages that have no links or pages that simply do not want to be found by classic search engines and are protected by a password.

Nevertheless, search engines have made the internet tangible for us users and make it possible to search it for relevant information. This article has only scratched the surface; in one of the next blog articles I will go into more detail on the topic of information retrieval Information retrieval means making stored, unstructured information retrievable and analyzable. This can be described as the science behind search engines.

Sources:

Rate this text

Average rating: 4.9 / 5. | Number of ratings: 7

Need a GEO check?

Get in touch for more visibility – in search engines and AI answers.

  • More visibility and leads via Google and AI chats
  • Sustainable SEO & GEO strategies
  • Personal support
Free initial analysis →