Context-Symbol Training for Speech Recognition and Word-Usage Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle with misrecognition due to insufficient consideration of how words are used in context, leading to inefficiencies in learning and accuracy.

Innovation Solution

An information processing system that includes a first text data acquisition unit, a speech data generation unit, a context symbol acquisition unit, a text data generation unit, and a learning unit, which generates and utilizes context symbols to enhance speech recognition by considering how words are used in context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition learning is performed without context symbols, then the learning process is simpler and faster, but the speech recognition accuracy deteriorates due to insufficient consideration of word usage in context

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Context symbols are introduced as intermediary elements that mediate between the speech input and the recognition output. These symbols represent contextual information (such as grammatical role, semantic category, or discourse function) and are inserted into the text data to provide additional guidance during the recognition process, thereby improving accuracy without fundamentally changing the recognition architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Context symbols are prepared and inserted into the text data in advance of the speech recognition process. This preliminary action enriches the training data with contextual information before the actual recognition task, allowing the system to learn contextual patterns without adding complexity to the real-time recognition operation

Inventive Principle:
Principle #10Preliminary action

2Reliability

If context symbols are inserted into text data to improve learning effectiveness, then speech recognition accuracy is improved, but the data processing time and computational load increase

Engineering Contradiction:
Improverecognition reliabilityVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Context symbol insertion is performed as a preliminary data preparation step before the speech recognition training or operation. By pre-processing the text data to include context symbols, the system avoids the need for complex real-time contextual analysis during speech recognition, thereby maintaining high reliability while reducing actual processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of modifying the original text data permanently or creating complex data structures, the system creates enhanced copies of the text data with context symbols inserted. These copied datasets are used for training purposes, allowing the original data to remain unchanged and easily reusable

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250246180A1Information processing system, information processing apparatus, information processing method, and recording medium
Publication Date: 2025.07.31 NEC CORP
  • US20250246180A1 patent drawing
  • US20250246180A1 patent drawing
  • US20250246180A1 patent drawing

AI summary

An information processing system includes: a first text data acquisition unit that acquires first text data; a speech data generation unit that generates first speech data corresponding to the first text data; a context symbol acquisition unit that acquires a context symbol corresponding to a word included in the first text data; a text data generation unit that generates second text data by inserting the context symbol into the first text data; and a learning unit that performs learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first speech data and the second text data as inputs.