Phoneme-Based Call Word Learning Data for Changed-Word Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing call word recognition models require large amounts of recorded data for changed call words and rely on grapheme-to-phoneme technology, leading to reduced accuracy.

Innovation Solution

A call word learning data generation device that decomposes utterance data into phoneme units, compares user input with phoneme data, and generates learning data without recording the changed call word, using techniques like G2P conversion, silence removal, normalization, and boundary correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a call word recognition model uses grapheme-to-phoneme technology to recognize changed call words, then the model can handle variable call words, but the recognition accuracy is reduced

Engineering Contradiction:
Improveability to recognize changed call wordsVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the call word into phoneme units and creates separate phoneme-wise learning data for each phoneme. This allows the model to handle variable call words by combining phoneme-level information rather than relying on grapheme-to-phoneme conversion, thereby maintaining high recognition accuracy while adapting to changed call words.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation from grapheme-based to phoneme-based. By converting call words into phoneme units and creating learning data based on phoneme characteristics rather than grapheme patterns, the system achieves both adaptability to changed call words and high recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If the model records actual utterances of changed call words to improve recognition accuracy, then accuracy improves, but the time required for data collection and model update increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidtime for data collection and model update
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-processing existing utterance data into phoneme units and creating phoneme-wise learning data before the model needs to recognize changed call words. This allows the model to quickly adapt to new call words by combining pre-existing phoneme data rather than collecting new recorded data, significantly reducing update time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

3Speed

If the model uses general phonetic symbols from grapheme-to-phoneme conversion, then the processing speed is fast, but the recognition accuracy is reduced

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent creates copying of phoneme data from existing utterance data, generating phoneme-wise learning data that replicates the acoustic characteristics of actual utterances. This copied phoneme data maintains the processing speed benefits of phoneme-based representation while achieving high recognition accuracy by preserving the temporal and acoustic properties of real speech.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12380878B2Call word learning data generation device and method
Publication Date: 2025.08.05 HYUNDAI MOTOR CO LTD
  • US12380878B2 patent drawing
  • US12380878B2 patent drawing
  • US12380878B2 patent drawing

AI summary

The present disclosure relates to a device and method for generating call word learning data, and the call word learning data generation device includes a processor and storage. The storage stores utterance data and an utterance phrase corresponding to the utterance data. The processor is configured to decompose the utterance data into phoneme units based on the utterance data and the utterance phrase, to receive a call word through a user input, to decompose the received call word into phoneme units, to compare phoneme data of the call word with phoneme data of the utterance data, and to generate call word learning data by combining phoneme data matched as the comparison result.