Multilingual Cognate Mining for Translation Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine translation techniques are ineffective at producing accurate and reliable translations due to limited training corpora and variations in language expressions, resulting in low accuracy rates, especially for English-to-Chinese translations.

Innovation Solution

The system identifies and utilizes multilingual cognates from user profiles to translate text by matching phrases across languages, updating a language model, and employing cognates to assist in search queries and content alignment, thereby enhancing translation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine translation techniques are used with limited training corpora, then the system complexity remains low, but translation accuracy deteriorates significantly

Engineering Contradiction:
Improvetranslation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by collecting and storing multilingual cognate data from user profiles before translation tasks are performed. The system proactively builds a comprehensive training corpus by mining cognates across multiple languages from social media profiles, job descriptions, and user-generated content, preparing this data in advance to improve translation accuracy when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transitions from traditional single-language training corpora to a multidimensional approach by incorporating cognate relationships across multiple languages simultaneously. The system creates a multidimensional translation model that considers cognates in source language, target language, and intermediate languages, adding dimensional depth to the translation process beyond conventional single-pair translation methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If simple word substitution methods are used for machine translation, then the processing speed remains high, but translation reliability deteriorates due to inability to recognize whole phrases

Engineering Contradiction:
Improvetranslation reliabilityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies segmentation by dividing the translation process into distinct stages: cognate identification, phrase matching, context analysis, and translation generation. The system segments text into meaningful units (cognates, phrases, sentences) and processes each segment through appropriate analysis methods, improving reliability by ensuring whole phrases are recognized and translated accurately rather than relying on simple word-by-word substitution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism in the form of a cognate-based translation model that mediates between source and target languages. This intermediary layer uses identified cognates from multilingual profiles to bridge language gaps, enabling more reliable phrase-level translation while maintaining efficient processing through pre-identified cognate relationships.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If extensive multilingual training data is collected from user profiles, then translation accuracy improves, but data processing complexity and time increase

Engineering Contradiction:
Improvetranslation accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing and storing cognate relationships from user profiles during off-peak times or in the background. The system proactively mines multilingual data from social media profiles, job descriptions, and user content, organizing this data into structured cognate databases in advance, so that when translation tasks are performed, the processing time is minimized by querying pre-organized data rather than processing raw profiles in real-time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential cognate information from extensive user profile data, separating relevant translation data from unnecessary profile details. The system selectively extracts cognate pairs, phrase equivalencies, and contextual relationships from multilingual profiles, filtering out redundant information and focusing processing on high-value translation data, thereby reducing processing time while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10114817B2Data mining multilingual and contextual cognates from user profiles
Publication Date: 2018.10.30 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10114817B2 patent drawing
  • US10114817B2 patent drawing
  • US10114817B2 patent drawing

AI summary

Techniques for identifying multilingual cognates and using the multilingual cognates are provided. In one technique, multilingual cognates identified from multiple user profiles are used to train one or more translation models. In another technique, multilingual cognates identified from a single user's profile are used to translate text provided by that user. In another technique, multilingual cognates from a single user are used to align sentences in one language to sentences in another language and the aligned sentences are used to train a language model. In another technique, multilingual cognates identified from multiple user profiles are used to expand search queries. In another technique, multilingual cognates identified from multiple user profiles are used to translate other users' profiles into a target language so that users associated with a source language are viewing the other users' profiles.