Cross-Language Data Consolidation via Phonetic Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to effectively consolidate data records from various databases stored in different languages, leading to data fragmentation and inaccessibility, as they cannot accurately identify and link records with discrepancies in attributes, especially when records are in multiple languages.

Innovation Solution

A method and system that converts user identifications from text records into speech records, then into data records, grouping and consolidating identical or similar records across different languages to create a unified record for a single user, using text-to-speech conversion engines and data coding, while identifying the language of each record for accurate association.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional systems locate and delete duplicate data records within a database, then data redundancy is reduced, but the system cannot identify records in different languages as duplicates, leading to data fragmentation

Engineering Contradiction:
Improvedata consolidation accuracyVSAvoidlanguage compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces phonetic transcription as an intermediary representation layer between different language scripts. By converting user identifications from various languages into phonetic symbols (such as IPA), the system creates a language-independent intermediate form that enables accurate matching of records regardless of their original language, thus resolving the contradiction between maintaining language versatility and achieving reliable duplicate detection

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transforms the parameter of user identification from visual text representation (which varies by language) to phonetic representation (which is language-universal). This parameter change allows the same user identified in different languages to be recognized as identical through phonetic equivalence, thereby improving data consolidation accuracy while preserving multi-language capability

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If speech-to-text conversion is used for translation, then non-written language input can be processed, but pronunciation variations cause inaccurate word conversion

Engineering Contradiction:
Improvespeech input capabilityVSAvoidword conversion accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Instead of relying on speech-to-text conversion which produces text copies subject to pronunciation errors, the patent directly transcribes speech into phonetic symbol representations. This phonetic copying method preserves the actual pronunciation characteristics without the intermediate text conversion step, thereby maintaining ease of speech input while significantly improving conversion accuracy by avoiding the pronunciation-interpretation bottleneck

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10324949B2Method and system for consolidating data retrieved from different sources
Publication Date: 2019.06.18 PHONEMIX LTD
  • US10324949B2 patent drawing

AI summary

A method is provided for consolidating data retrieved from different text records stored in different languages and associated with a single user. According to an embodiment, the method comprises the steps of: extracting a plurality of users' identifications from a plurality of text records and converting them into a corresponding plurality of speech records, each being essentially identical to the pronunciation of a corresponding user identification in a language which its respective text record has been stored; converting each speech record to a respective data record; extracting from the data records obtained, at least one group of data records comprising two or more data records essentially identical to each other; for each of the groups, retrieving information comprised in two or more text records which are stored in different languages from each other; and storing the information retrieved in a consolidated text record.