Bilingual Corpus Expansion via Pivot Language Phrase Mining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Statistics-based machine translation systems face challenges due to data sparseness in bilingual corpora, which affects translation quality.

Innovation Solution

The method involves searching for semantically matching phrases in a source language-pivot language corpus and a pivot language-target language corpus, forming phrase pairs, and storing them in a source language-target language corpus to expand the bilingual data, using modules for pivot language phrase search, source and target language phrase set establishment, phrase pair combination, and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manually creating rules for machine translation is used, then translation quality can be maintained with limited data, but the system cannot be applied to all languages and requires extensive manual effort

Engineering Contradiction:
Improveapplicability to all languagesVSAvoidmanual rule creation effort
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a pivot language as an intermediary to bridge source and target languages. By translating source language to pivot language and then to target language, the system achieves universal applicability without requiring manual rules for every language pair, resolving the contradiction between versatility and complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses parallel corpora to copy translation patterns and phrases from existing high-quality translations. By extracting and reusing proven translation units from parallel corpora, the system achieves high translation quality automatically without manual rule creation for each language pair

Inventive Principle:
Principle #26Copying

2Quantity of substance

If a larger amount of data is used in the bilingual corpus, then translation quality improves, but data sparseness problem persists in the initial stage of corpus establishment

Engineering Contradiction:
Improveamount of data in corpusVSAvoidtranslation quality due to data sparseness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent merges data from multiple sources including parallel corpora, pivot language corpora, and extracted phrase pairs to build a comprehensive bilingual corpus. By combining these diverse data sources, the system overcomes initial data sparseness and achieves both large quantity and high reliability of training data

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary data collection and processing by extracting phrase pairs from parallel corpora before actual translation tasks. This preliminary action of pre-processing and pre-collecting high-quality phrase pairs ensures that sufficient training data is available from the outset, eliminating the data sparseness problem

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If statistics-based machine translation is used, then manual rule creation is eliminated and applicability to all languages is achieved, but translation quality depends heavily on corpus quality which suffers from data sparseness

Engineering Contradiction:
Improveautomatic system establishmentVSAvoidtranslation quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent automatically copies high-quality translation patterns from parallel corpora through phrase extraction. By systematically copying proven translation units from existing high-quality parallel data, the system maintains high translation quality automatically without manual intervention, resolving the contradiction between ease of manufacture and reliability

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9953024B2Method and device for expanding data of bilingual corpus, and storage medium
Publication Date: 2018.04.24 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US9953024B2 patent drawing
  • US9953024B2 patent drawing
  • US9953024B2 patent drawing

AI summary

Disclosed are a method and a device for expanding data of a bilingual corpus. The method for expanding data of a bilingual corpus includes: searching, in a source language-pivot language corpus, for at least one first pivot language phrase semantically matching a first source language phrase; searching, in the source language-pivot language corpus, for at least one second source language phrase semantically matching each of the first pivot language phrases to form a source language phrase set by the second source language phrases; searching, in a pivot language-target language corpus, for at least one first target language phrase semantically matching each of the first pivot language phrases to form a target language phrase set by the first target language phrases; combining the second source language phrases in the source language phrase set with the first target language phrases in the target language phrase set, so as to form at least one phrase pair in which a source language phrase and a target language phrase semantically match; and storing the formed at least one phrase pair in which the source language phrase and the target language phrase semantically match into a source language-target language corpus. Data in a bilingual corpus is expanded, so that the problem of data sparseness in the bilingual corpus is solved.