Transliteration Detection via Consonant Correspondence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating transliteration dictionary data is inefficient due to the time-consuming manual process and the difficulty in preparing learning data for machine learning, as it is challenging to determine the necessary rules for creating appropriate transliteration rules.

Innovation Solution

A transliteration processing device and method that acquire alphabetic character strings representing words in different languages and determine transliteration relationships based on the correspondence between consonant elements, using a predetermined correspondence rule to efficiently detect transliteration pairs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If transliteration dictionary data is created by hand, then accuracy of transliteration pairs is improved, but time and effort required increases

Engineering Contradiction:
Improveaccuracy of transliteration pairsVSAvoidtime and effort required
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an automatic detection system that acts as an intermediary between manual creation and machine learning. The system uses consonant element correspondence rules as intermediaries to automatically identify transliteration pairs, reducing the need for both manual verification and extensive machine learning training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical process of manual transliteration pair creation with an automated computational system. The determination unit automatically analyzes consonant element correspondence between character strings to identify transliteration pairs, eliminating the need for manual word-by-word analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If machine learning is used to create transliteration dictionary data, then productivity is improved, but difficulty in preparing learning data increases

Engineering Contradiction:
Improveefficiency of creating transliteration dictionary dataVSAvoiddifficulty in preparing learning data
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts the essential feature for transliteration detection - consonant element correspondence - from the complex process of creating transliteration dictionary data. By focusing only on consonant elements rather than requiring complete learning datasets, the system simplifies the data preparation process while maintaining productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter used for transliteration detection from comprehensive machine learning models to specific consonant element correspondence rules. This parameter change allows for more efficient data preparation by focusing on specific linguistic features rather than requiring extensive training corpora.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive learning rules are prepared for machine learning, then accuracy of transliteration detection is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of transliteration detectionVSAvoidcomplexity of learning rules
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex task of transliteration detection into a simpler sub-task: analyzing consonant element correspondence. By dividing the problem into focused components (consonant elements rather than complete words or phrases), the system achieves accurate detection with simpler rules.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10185710B2Transliteration apparatus, transliteration method, transliteration program, and information processing apparatus
Publication Date: 2019.01.22 RAKUTEN GROUP INC
  • US10185710B2 patent drawing
  • US10185710B2 patent drawing
  • US10185710B2 patent drawing

AI summary

A transliteration processing device according to one embodiment includes a character string acquisition unit that acquires a first alphabetic character string representing by alphabet a first word written in a first language having a specified script and a second alphabetic character string representing by alphabet a second word written in a second language having a different script from the first language, a determination unit that makes a determination whether a first consonant element included in the first alphabetic character string and a second consonant element included in the second alphabetic character string have a predetermined correspondence, and determines whether the first word and the second word have a transliteration relationship based on a result of the determination, and an output unit that outputs, as a transliteration pair, the first word and the second word determined to have a transliteration relationship by the determination unit.