Chapter 15 10 min read

Entity Mention Clarity

Entity Mention Clarity is a metric that evaluates the density, diversity, and confidence of named entities recognized within a page's content.

Named Entity Recognition (NER) is the process of identifying and categorizing key information into predefined classes such as Person, Organization, Location, Product, and so on. For an LLM, clear and abundant entities serve as crucial anchors, connecting the text to a real-world knowledge graph. Content that is rich in well-defined entities is more semantically precise and provides clearer context, making it easier for an AI to understand relationships and verify facts, thus increasing its value as a source document.

Calculation Methodology

The score is a composite index based on three factors derived from running the page's content through an NER model (such as spaCy or Flair).

Entity Extraction

Process the clean text of the page's main content through a pre-trained NER model to extract all entities, their labels (e.g., ORG, PERSON), and the model's confidence score for each prediction.

Diversity Score (40% weight)

This measures the breadth of entity types mentioned. Count the number of unique entity labels (e.g., PERSON, ORG, GPE, DATE, PRODUCT) found on the page. The score is normalized based on the total number of possible entity types in the model. A page that discusses people, organizations, and locations is semantically richer than one that only mentions dates.

Density Score (40% weight)

This measures the frequency of entity mentions. Count the total number of identified entities. Normalize this count per 1000 words of content. A log-scaling function can be applied to prevent extremely high frequencies from dominating the score, rewarding a healthy density without over-optimizing for entity stuffing.

Confidence Score (20% weight)

This measures the reliability of the entity predictions. Calculate the average confidence score provided by the NER model across all identified entities on the page. A higher average confidence indicates that the entities are well-defined and unambiguous.

Calculating The Entity Mention Clarity Score

The calculation of the Entity Mention Clarity Score is a multi-step process that involves analyzing the HTML structure of the page. Here is a simplified pseudo-code representation of how this score is calculated.

Pseudo-code for Entity Mention Clarity Score Calculation

BEGIN
 FETCH and EXTRACT clean text from the webpage.
 PROCESS text through a Named Entity Recognition (NER) model to get a list of entities with labels and confidence scores.

 // Calculate Diversity
 COUNT the number of unique entity labels.
 NORMALIZE against the total number of possible labels in the model.
 COMPUTE diversity_score.

 // Calculate Density
 COUNT the total number of entities.
 NORMALIZE against word count.
 COMPUTE density_score.

 // Calculate Confidence
 AVERAGE the confidence scores from the NER model for all entities.
 COMPUTE confidence_score.

 CALCULATE final_score as the weighted average of the three components.
 RETURN final_score and a summary of entities.
END

Conclusion

Entity Mention Clarity is a crucial metric for GEO, as it directly impacts how well AI systems can understand and connect your content to the real world. By focusing on this metric, you can improve the semantic precision of your content and increase its value as a source for generative engines.

Key Takeaways

  • Entity Mention Clarity evaluates the density, diversity, and confidence of named entities in your content.
  • The score is a composite of diversity, density, and confidence scores from an NER model.
  • A high score indicates that your content is semantically rich and easy for AI to understand.
  • To improve your score, focus on including a variety of well-defined entities in your content.