Link Context Language Relevance Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in accurately determining the language of resources, especially when resources have insufficient textual content or when multiple languages are present, leading to difficulties in identifying relevant languages for search engines.

Innovation Solution

A method that generates language relevance scores for incoming and outgoing resource links, allowing for the identification of relevant languages by analyzing the context and features of these links, including anchor text, source and target resource languages, and common source features, to contextualize language relevance without being disproportionately influenced by large groups of links.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional content-based language identification methods are used, then language detection can be performed on resources with sufficient text, but it fails when resources have insufficient textual content or no text at all

Engineering Contradiction:
Improvelanguage identification capabilityVSAvoidlanguage detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent uses link context as an intermediary to infer language information. Instead of directly analyzing the target resource's content (which may be insufficient or nonexistent), the system examines the languages of source resources that link to the target resource. This intermediary approach allows language identification to work even when the target resource has little or no text.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the language identification task into analyzing individual incoming links and their source resource languages. By processing each link separately and aggregating the results, the system can handle resources with insufficient overall text content while maintaining reliable language detection through cumulative evidence from multiple links.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If link-based language association is used, then language relevance can be determined for resources with limited text, but the method is disproportionately influenced by large groups of links from the same source

Engineering Contradiction:
Improvelanguage relevance detectionVSAvoidlanguage relevance accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by differentiating between links from the same source resource. Instead of treating all incoming links equally, the system adjusts the weight or treatment of links based on their source. Links from the same source resource are processed together as a group, allowing the system to account for the fact that multiple links from one source may represent a coordinated effort rather than independent language signals.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses partial action by selectively counting or weighting links based on their source characteristics. When multiple links come from the same source resource, the system applies a modification to the counting process, effectively using partial information (the group of links from a single source) rather than treating each link as an independent complete signal. This prevents excessive influence from large groups of links while still utilizing the available link data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9098582B1Identifying relevant document languages through link context
Publication Date: 2015.08.04 GOOGLE LLC
  • US9098582B1 patent drawing
  • US9098582B1 patent drawing
  • US9098582B1 patent drawing

AI summary

Methods, systems, and apparatus, including computer program products, for identifying languages that are relevant to resource. In an aspect, language features are identified for incoming resource links to a resource and outgoing resource links from the resource. The language features or use by a language classification model to generate language relevance scores. The language relevance scores for each of the incoming resource links and outgoing resource links are used to generate a corresponding relevance measure for each of a plurality of languages. Each relevance measure is a measure of the relevance of the language to the resource.