Link Context Language Relevance Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately determining the language of resources, especially when resources have insufficient textual content or when multiple languages are present, leading to difficulties in identifying relevant languages for search engines.
Innovation Solution
A method that generates language relevance scores for incoming and outgoing resource links, allowing for the identification of relevant languages by analyzing the context and features of these links, including anchor text, source and target resource languages, and common source features, to contextualize language relevance without being disproportionately influenced by large groups of links.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional content-based language identification methods are used, then language detection can be performed on resources with sufficient text, but it fails when resources have insufficient textual content or no text at all
Solution Approach 1:
The patent uses link context as an intermediary to infer language information. Instead of directly analyzing the target resource's content (which may be insufficient or nonexistent), the system examines the languages of source resources that link to the target resource. This intermediary approach allows language identification to work even when the target resource has little or no text.
Solution Approach 2:
The patent segments the language identification task into analyzing individual incoming links and their source resource languages. By processing each link separately and aggregating the results, the system can handle resources with insufficient overall text content while maintaining reliable language detection through cumulative evidence from multiple links.
2Adaptability or versatility
If link-based language association is used, then language relevance can be determined for resources with limited text, but the method is disproportionately influenced by large groups of links from the same source
Solution Approach 1:
The patent applies local quality by differentiating between links from the same source resource. Instead of treating all incoming links equally, the system adjusts the weight or treatment of links based on their source. Links from the same source resource are processed together as a group, allowing the system to account for the fact that multiple links from one source may represent a coordinated effort rather than independent language signals.
Solution Approach 2:
The patent uses partial action by selectively counting or weighting links based on their source characteristics. When multiple links come from the same source resource, the system applies a modification to the counting process, effectively using partial information (the group of links from a single source) rather than treating each link as an independent complete signal. This prevents excessive influence from large groups of links while still utilizing the available link data.
Data Source
AI summary
Methods, systems, and apparatus, including computer program products, for identifying languages that are relevant to resource. In an aspect, language features are identified for incoming resource links to a resource and outgoing resource links from the resource. The language features or use by a language classification model to generate language relevance scores. The language relevance scores for each of the incoming resource links and outgoing resource links are used to generate a corresponding relevance measure for each of a plurality of languages. Each relevance measure is a measure of the relevance of the language to the resource.


