Domain-Aware Snippet Extraction via Template Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search engine algorithms lack the ability to identify the most important text from web pages, often extracting irrelevant content for snippets due to their 'one-size-fits-all' approach, which fails to account for domain-specific structures and layouts.

Innovation Solution

A domain-aware snippet extraction method that identifies templates and tag patterns across multiple web pages within a domain, ranking sections based on importance and relevance to provide targeted and relevant snippets for search results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a conventional one-size-fits-all algorithm is used to extract snippets from all web pages, then the extraction process is simple and fast, but the accuracy and relevance of the extracted snippets deteriorate

Engineering Contradiction:
Improvesnippet extraction speedVSAvoidsnippet relevance accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating domain-specific snippet extraction algorithms tailored to different web page types. Instead of using a uniform approach for all pages, the system analyzes the domain and applies specialized extraction rules appropriate for each domain's structure and content patterns, thereby improving snippet relevance without sacrificing overall system efficiency

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes extraction parameters based on domain characteristics. By detecting the domain type and analyzing web page structure patterns, the algorithm adjusts extraction parameters such as section weighting, keyword importance, and text selection criteria to optimize snippet quality for each specific domain while maintaining fast processing through parameter-based differentiation rather than complete reprocessing

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If domain-specific template identification is implemented to improve snippet accuracy, then the relevance of extracted snippets improves, but the system complexity increases

Engineering Contradiction:
Improvesnippet extraction accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary domain analysis and template identification in advance. By pre-processing web pages to identify domain-specific templates, section patterns, and structural characteristics before snippet extraction, the system builds a knowledge base that guides accurate extraction without adding complexity during the actual snippet generation process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses template copying by identifying recurring structural patterns across multiple web pages within a domain. Once a template is identified from representative pages, it is copied and applied to extract snippets from similar pages, reducing the need for complex individual page analysis while maintaining high accuracy through pattern-based extraction

Inventive Principle:
Principle #26Copying

3Use of energy by moving object

If conventional heuristics are used to extract snippets, then the extraction process is computationally efficient, but the ability to identify important text deteriorates

Engineering Contradiction:
Improvecomputational energy consumptionVSAvoidimportance of extracted text
Core Design Contradiction:
Use of energy by moving objectVSLoss of information

Solution Approach 1:

The system implements feedback by analyzing the results of snippet extraction and using this information to refine domain-specific templates and extraction parameters. By continuously learning from extraction outcomes and adjusting templates based on identified patterns, the system improves its ability to identify important text while maintaining computational efficiency through refined, experience-based rules

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8195634B2Domain-aware snippets for search results
Publication Date: 2012.06.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8195634B2 patent drawing
  • US8195634B2 patent drawing
  • US8195634B2 patent drawing

AI summary

Techniques are disclosed for providing a domain-aware snippet for a search result. With such techniques, a domain classification component is provided for identifying a template used to generate a plurality of web pages of a domain, associating the template and content of the web pages related to the template with a Uniform Resource Locator pattern of the plurality of web pages, and storing the associated template, the related content, and the Uniform Resource Locator pattern in a database. A snippet extraction component is also provided for extracting text from a section of a web page of the plurality of web pages for a snippet of a search result corresponding to a search query, wherein the extracted text is based on a ranking value of the section and the relevance of the extracted text to the search query.