Search Engine Ranking Using Semi-Structured Web Page Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Search engines fail to accurately rank semi-structured web pages based on user intent due to not considering features independent of query content, leading to suboptimal search results and user dissatisfaction.
Innovation Solution
A search engine system that identifies and utilizes features common across semi-structured web pages, such as review counts and social media engagement, through programmatic analysis and machine learning, to rank web pages independently of query content, employing semi-automated wrapper induction and scoring functions to position relevant pages higher in search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If search engines use complex algorithms to rank documents based on query content, then search result relevance is improved, but features independent of query content are not considered
Solution Approach 1:
The patent segments the ranking process into two independent components: query-based relevance scoring and feature-based quality scoring. The ranking system separately extracts features independent of query content (such as review counts, social media engagement, trust signals) and combines them with query-based relevance. This segmentation allows both query matching and independent feature consideration to contribute to final ranking without interfering with each other.
Solution Approach 2:
The patent merges query-based document retrieval with feature-based quality assessment into a unified ranking framework. Documents are first retrieved based on query relevance, then enriched with independent features extracted from semi-structured content. The combined scoring system integrates both query matching results and feature-based quality metrics to produce the final ranked list, ensuring both aspects work together synergistically.
2Device complexity
If search engines consider only query content when ranking, then processing simplicity is maintained, but user engagement metrics are ignored
Solution Approach 1:
The patent applies preliminary action by pre-extracting and storing feature information from semi-structured web pages before the search query is processed. Features such as review counts, social media engagement metrics, and trust signals are identified and cached in advance. When a query is received, these pre-computed features are quickly retrieved and applied to ranking without adding significant processing complexity during the actual search operation.
3Speed
If search engines ignore semi-structured content features, then ranking speed is maintained, but user satisfaction decreases
Solution Approach 1:
The patent applies local quality by selectively extracting only the most relevant features from semi-structured content rather than processing all possible information. Specific high-value features such as review counts, star ratings, social media engagement metrics, and trust signals are identified and extracted from their local positions in the document structure. This targeted extraction maintains processing speed while capturing the most impactful quality indicators for user satisfaction.
Data Source
AI summary
Features automatically extracted from semi-structured web pages are utilized by a search engine to rank documents that include semi-structured web pages. These features include, but are not limited to, a number of reviews, a number of positive reviews, and/or a number of negative reviews from a web page that includes user reviews. These features also include a number of views of a video that is viewable by way of a semi-structured web page. The features also include a number of subscribers to broadcasts of an individual from a social networking web page and a number of contacts of an individual listed on a social networking web page.


