WDF*IDF explained simply and clearly
WDF*IDF is a formula that is widely used in search engine optimization, especially in German-speaking regions. Unlike keyword density, the focus here is not on the percentage occurrence of a keyword within a text. Instead, the frequency of a main keyword and related terms is set in relation to other pages.
The abbreviation WDF stands for “Within Document Frequency” and refers to the frequency of a term (with WDF*IDF, one tends to speak of a “term” rather than a “keyword”) within a document. IDF stands for “Inverse Document Frequency”. This means that the number of all known documents is set in relation to the number of texts containing the corresponding term.
The two formulas look as follows:
i: word
j: document
L: the total number of words in the document j
Freq(i,j): the frequency of the word i in the document j
ND: total number of documents;
ft: number of documents in which the term appears.
How WDF*IDF tools work
In order to subject documents to a WDF*IDF analysis, appropriate tools are required. These generally work as follows:
- The user enters a main keyword and selects a country.
- The tool examines the pages from the respective country that rank best for the keyword. On this basis, it determines relevant terms (proof keywords) that should be contained in the text, along with an ideal term weighting.
- When the user then enters their text, the tool shows which terms are missing, underrepresented, or used too often. On this basis, they can revise the text until it corresponds as closely as possible to the ideal ratio of proof keywords.
Common tools with WDF*IDF analysis include TermLabs.io, Seobility and SurferSEO. They differ above all in the number of comparison pages evaluated and in how they present the proof keywords.
The importance of WDF*IDF for SEO
The formula has existed for decades; in on-page SEO, it was made popular above all by Karl Kratz. By now, many SEO tools offer a WDF*IDF analysis.
Effectiveness and impact on Ranking are disputed. Advocates see the formula as a modern successor to keyword density. Critics point out that WDF*IDF tools do not take into account many synonyms, strong term clustering within a paragraph, or keyword combinations. Above all, readability, tone, and formatting play no role – even though they are crucial for texts to resonate with the target audience.
Ultimately, WDF*IDF comes down to mathematical calculations. These alone are not enough to optimize texts successfully, but they can provide important insights into the content – for example, whether an author has missed the topic or whether important thematic areas are missing. More on on-page optimization.
WDF*IDF is distinct from TF*IDF: TF*IDF (Term Frequency * Inverse Document Frequency) is the older, classic approach from information retrieval and uses the simple, linear term frequency – similar to keyword density. WDF*IDF, by contrast, compresses the term frequency within the document logarithmically (WDF = Within Document Frequency) and thereby reacts less sensitively to individual outliers and mere word repetitions: a tenfold occurrence is weighted with heavy damping instead of tenfold. In practice, WDF*IDF therefore rates the balanced use of relevant terms higher than the mere inflation of a single keyword.
Quellen
SEO & GEO check
Want more visibility – in search engines and AI answers?
We analyse your potential and show concrete next steps. Free and non-binding.