Speech Recognition Training With Converted Text Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in accurately learning from speech data without corresponding text data, and existing methods for generating pseudo data are inadequate.

Innovation Solution

An information processing system that includes a first text data acquisition unit, a text data conversion unit, and a learning unit to generate converted speech data, which are used to enhance the learning process of a speech recognition unit by augmenting the data with converted text data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition learning is performed using speech data without corresponding text data, then the system can process more diverse data sources, but the learning accuracy deteriorates due to lack of precise alignment between speech and text

Engineering Contradiction:
Improvedata source versatilityVSAvoidlearning accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces converted text data as an intermediary element. This converted text data serves as a mediator between the speech data and the learning process, providing the necessary textual information alignment without requiring original paired speech-text data. The converted text data acts as a bridge that enables accurate learning from diverse data sources while maintaining precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If existing pseudo learning data generation methods are used, then the system can generate training data without speech recognition, but the generated data are inadequate for improving speech recognition accuracy

Engineering Contradiction:
Improvetraining data generation easeVSAvoidlearning data quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent transforms text data into converted text data by applying specific conversion rules that simulate speech characteristics. This parameter transformation process converts static text into text that reflects speech patterns, errors, and variations. The converted text data then serves as high-quality training data that maintains both ease of generation and high learning quality, as the conversion process preserves essential speech-related characteristics.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If text data is converted to generate training data, then more learning data can be generated from limited original data, but the conversion process increases system complexity

Engineering Contradiction:
Improvetraining data quantityVSAvoidconversion process complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent creates converted text data by copying and transforming original text data according to predefined conversion rules. This copying process generates additional training data that replicate the characteristics of real speech data. The conversion rules are designed to be straightforward text transformation operations, which increases data quantity while keeping the added complexity manageable through systematic rule-based conversion rather than complex generative models.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250246179A1Information processing system, information processing method, and non-transitory recording medium
Publication Date: 2025.07.31 NEC CORP
  • US20250246179A1 patent drawing
  • US20250246179A1 patent drawing
  • US20250246179A1 patent drawing

AI summary

An information processing system includes: a first text data acquisition unit that acquires first text data: a text data conversion unit that converts the first text data, thereby to generate converted text data; a converted speech data generation unit that generates converted speech data corresponding to the converted text data; and a learning unit that performs learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first text data and the converted speech data as inputs.