Chapter 14 9 min read

Pre-formatted Info Blocks

Pre-formatted Info Blocks refers to the identification and quality assessment of content enclosed within HTML <pre> and <code> tags.

These tags are semantically significant as they instruct browsers—and by extension, AI crawlers—to preserve whitespace, line breaks, and tabs, displaying the content in a fixed-width font. This is essential for accurately representing structured information like source code, log files, configuration data, or ASCII art. For an LLM, correctly identifying these blocks is crucial to prevent misinterpretation; for example, it must understand that the indentation in a Python code snippet is syntactically vital and not just stylistic. This metric counts the number of such blocks and evaluates their appropriate use.

Calculation Methodology

The evaluation is a two-part process: quantitative identification and qualitative assessment.

Identification and Counting

Scrape the page's full HTML content. Use a parsing library (like BeautifulSoup) to find all instances of <pre> and <code> tags. Count the total number of each tag type. A higher count can indicate content rich in technical detail.

Quality and Appropriateness Scoring

  • Content Type Analysis: For each block, extract the text content. Use an LLM to classify the type of content within the block (e.g., "Python code," "JSON data," "Command-line output," "Plain text," "ASCII art").
  • Semantic Correctness: Score the appropriateness of the tag used. Content identified as source code should ideally be wrapped in <pre><code>...</code></pre> for block-level code, or just <code>...</code> for inline code. Award a high score for this correct nesting. Penalize the use of <pre> for content that is simply trying to achieve a monospace font style without being genuinely pre-formatted (e.g., a normal paragraph).
  • Code Validity (for code blocks): For content identified as a specific programming language, a basic syntax validity check can be performed using a relevant linter or library. This provides a signal of quality and attention to detail.

Calculating The Pre-formatted Info Blocks Score

The calculation of the Pre-formatted Info Blocks Score is a multi-step process that involves analyzing the HTML structure of the page. Here is a simplified pseudo-code representation of how this score is calculated.

Pseudo-code for Pre-formatted Info Blocks Score Calculation

BEGIN
 FETCH and PARSE the webpage HTML.
 FIND and COUNT all <pre> and <code> tags.
 INITIALIZE total_quality_score to 0.

 FOR EACH <pre> or <code> block:
 EXTRACT the text content.
 USE LLM to CLASSIFY the content type (e.g., "Python code", "JSON", "Plain text").
 ASSIGN a quality score based on semantic correctness (e.g., code is in <pre><code>, not just <pre>).
 (Optional) VALIDATE syntax if content is identified as code.
 ADD quality score to total_quality_score.
 
 CALCULATE average_quality_score.
 RETURN tag counts and average_quality_score.
END

Conclusion

Pre-formatted Info Blocks are a powerful tool for communicating structured information to AI systems. By using them correctly, you can ensure that your technical content is accurately interpreted and used by generative engines, which is essential for establishing your site as a credible source for technical information.

Key Takeaways

  • Pre-formatted Info Blocks are essential for representing structured information like code, logs, and data.
  • The metric evaluates the quantity and quality of <pre> and <code> tags on a page.
  • Correctly nesting <code> inside <pre> for block-level code is a key quality signal.
  • Using these tags appropriately helps prevent misinterpretation by AI models.