Social Graph Language Detection for Short Text Fragments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods struggle to accurately determine the language of short text fragments in social networking posts, especially when they contain abbreviations and incorrect grammar, which complicates language identification.

Innovation Solution

The method involves analyzing social graphs associated with the author, including statistics on posts and comments, to determine the language of the text, and uses a combination of social graph data to weight language statistics, thereby improving language detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional language detection methods are used on short text fragments, then the processing speed is fast, but the language identification accuracy deteriorates due to abbreviations and incorrect grammar

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoiddetection method complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces social graph data as an intermediary element to bridge the gap between short text fragments and accurate language identification. Instead of directly analyzing the limited text content, the system uses the author's social connections, contact languages, and interaction patterns as mediating information to infer the language, thereby resolving the contradiction between text brevity and detection accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from analyzing only the text content dimension to incorporating the social relationship dimension. By adding social graph data (contacts, interactions, language preferences of connected users) as an additional dimension of analysis, the system compensates for the insufficiency of short text fragments and achieves accurate language identification despite the text's limited length

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If social graph data is used to determine language, then language detection accuracy is improved, but the data processing complexity increases

Engineering Contradiction:
Improvelanguage detection accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-collecting and storing social graph data, including user contacts, interaction histories, and language preferences, before language detection is needed. This pre-processing allows the actual language detection to use readily available aggregated data, reducing the complexity of real-time processing while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges multiple data sources (author's own posts, contact languages, interaction patterns, and social graph relationships) into a unified language detection framework. By combining these diverse elements into a single probabilistic model, the system achieves accurate language identification without managing the complexity of separate processing pipelines for each data source

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If only the author's own posts are analyzed, then the processing is simple, but the language determination accuracy deteriorates for short fragments

Engineering Contradiction:
Improvelanguage determination accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent adds the social relationship dimension to the data analysis, incorporating information from contacts and interactions beyond the author's own posts. This dimensional expansion provides additional contextual clues for language identification, compensating for the limited data volume in short text fragments while maintaining manageable data quantities through focused social graph sampling

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS8832188B1Determining language of text fragments
Publication Date: 2014.09.09 GOOGLE LLC
  • US8832188B1 patent drawing
  • US8832188B1 patent drawing
  • US8832188B1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving content data, the content data comprising text, identifying an author of the text, the author being a user of a social networking service, retrieving data corresponding to one or more social graphs associated with the user, the data being stored in a computer-readable storage device, and determining the language of the text based on the data corresponding to the one or more social graphs associated with the author.