Context-Aware Voice Data Creation for Additional Word Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to accurately recognize additional words without context, leading to insufficient recognition accuracy and high costs due to the need for generating voice data for these words.

Innovation Solution

A voice data creation device that extracts and selects text corpora likely to include additional words in context, using a language model to determine optimal sentence examples, synthesizes voices for these examples, and outputs them as voice data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If voice data of only words is learned without context, then recognition cost is reduced, but recognition accuracy is insufficient

Engineering Contradiction:
Improvevoice data creation costVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent extracts only the necessary context information from complete sentences surrounding the additional word, rather than using full sentences. This extraction approach provides sufficient contextual information for accurate recognition while reducing the data volume and synthesis cost, resolving the contradiction between cost reduction and accuracy improvement.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary extraction and selection of optimal sentence examples containing the additional word before voice synthesis. By pre-selecting sentences with appropriate context and calculating occurrence likelihood, the system prepares optimized training data in advance, reducing subsequent processing costs while ensuring recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a person generates voice data by generating additional words, then recognition accuracy is improved, but cost and labor are greatly increased

Engineering Contradiction:
Improverecognition accuracyVSAvoidvoice data creation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

Instead of manually generating voice data for each additional word, the system copies and synthesizes voice from pre-selected sentence examples containing the additional word. This automated copying approach from existing text corpora eliminates manual recording costs while maintaining contextual information necessary for accurate recognition.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses automated voice synthesis technology to generate voice data from selected text corpora without requiring human recording. The synthesis process automatically creates voice data with proper context, eliminating manual labor while ensuring consistent quality and contextual accuracy.

Inventive Principle:
Principle #25Self-service

3Reliability

If text corpora with highest likelihood of occurrence are selected as optimal sentence examples, then context information quality is improved, but processing complexity increases

Engineering Contradiction:
Improvecontext information qualityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex manual evaluation of context quality with automated language model-based likelihood calculation. The system uses computational algorithms to automatically determine which sentence examples have the highest probability of containing the additional word in natural contexts, substituting mechanical processing complexity with efficient computational methods that maintain high reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12531050B2Voice data creation device
Publication Date: 2026.01.20 NTT DOCOMO INC
  • US12531050B2 patent drawing
  • US12531050B2 patent drawing
  • US12531050B2 patent drawing

AI summary

A voice data creation device is a device configured to create voice data including an additional word which is a word to be added to a recognition target in a speech recognition system, and includes: a sentence example extraction unit configured to extract one or more text corpora including the additional word from a text corpus group including a plurality of text corpora consisting of sentence examples including a plurality of words; a sentence example selection unit configured to select a text corpus having a highest measure indicating a likelihood of occurrence as a sentence among the text corpora extracted by the sentence example extraction unit 11 as an optimal sentence example for the additional word; and a voice creation unit configured to output a synthesized voice of the optimal sentence example generated by a predetermined voice synthesis system as voice data corresponding to the additional word.