Document Ranking via Context Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing search engine ranking processes are degraded by techniques such as link-based spamming, anchor text spamming, and bombing, which artificially inflate document ranks, leading to lower quality search results.
Innovation Solution
A method and system that analyze the context of links by identifying rare words in text windows surrounding links, creating context identifiers, and ranking documents based on these identifiers to reduce the impact of spamming techniques, improving the relevance of search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If link-based spamming techniques are used to increase document ranks, then the number of links to a document increases, but the quality of search results degrades
Solution Approach 1:
The patent applies local quality by analyzing the specific context (local text window) surrounding each link rather than treating all links uniformly. It examines the semantic content of text windows adjacent to links to determine their genuine relevance, thereby distinguishing spam links from legitimate ones and maintaining accurate document ranking.
Solution Approach 2:
The patent introduces context identifiers as an intermediary element between links and ranking evaluation. These identifiers capture the semantic context of link placements and serve as mediators to assess whether links are genuine or spam, preventing direct manipulation of rank metrics while preserving search quality.
2Measurement precision
If anchor text spamming is used to associate documents with search terms, then the number of documents with matching anchor text increases, but the relevance of search results decreases
Solution Approach 1:
The patent examines the local text context surrounding anchor text to determine its genuine relevance. By analyzing the semantic content of text windows adjacent to anchor text occurrences, it distinguishes between meaningful anchor text and spammy repetitions, preserving search result relevance while allowing legitimate term matching.
Solution Approach 2:
The patent uses context identifiers as intermediaries to evaluate anchor text quality. These identifiers capture the surrounding semantic context and mediate between raw anchor text matching and relevance assessment, preventing spammy anchor text from artificially inflating search result relevance.
3Measurement precision
If bombing techniques are used to manipulate document ranks, then the number of documents with specific anchor text increases, but the accuracy of search result ranking decreases
Solution Approach 1:
The patent applies local quality by analyzing the specific semantic context of each link placement rather than counting links uniformly. It examines text windows surrounding links to assess their genuine relevance, thereby detecting and neutralizing bombing attempts that rely on volume rather than quality.
Solution Approach 2:
The patent introduces context identifiers as intermediaries between link data and ranking calculations. These identifiers capture semantic context and mediate the ranking process, preventing direct manipulation through volume-based bombing while preserving accurate relevance-based ranking.
4Measurement precision
If standard frames with duplicated links are used across multiple documents, then the number of links to certain documents increases, but the quality of ranking information decreases
Solution Approach 1:
The patent applies local quality by analyzing the unique contextual environment of each link instance. It examines the specific text window content surrounding each link to determine its genuine relevance, thereby distinguishing between meaningful repeated links in standard frames and spammy link duplication, preventing artificial rank inflation.
Data Source
AI summary
A system ranks documents based on contexts associated with the documents. The system identifies a reference in a first document, where the reference is associated with a second document. The system analyzes a portion of the first document associated with the reference, identifies a rare word (or words) from the portion, creates a context identifier based on the rare word(s), and ranks the second document based on the context identifier.


