User-Adaptive Speech Recognition for Repeated Speech Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition accuracy varies due to individual differences in repeated speech, leading to inconsistent recognition results.

Innovation Solution

A speech recognition apparatus that divides target speech into sections, reproduces each section, recognizes repeated speech by a user, generates text information, and stores user-specific learning data for improved recognition using a user-tailored recognition engine.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition is performed on repeated speech without user-specific adaptation, then the system complexity is low, but speech recognition accuracy varies due to individual differences

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary speech collection and learning before actual recognition tasks. User-specific speech data is collected and stored in advance, creating a personalized learning database that the recognition engine can utilize during transcription tasks to improve accuracy without adding complexity during the actual recognition process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The recognition engine automatically learns from user-specific speech data and adapts to individual speech patterns without requiring manual configuration or intervention. The system self-improves by utilizing the collected learning data to enhance its recognition capabilities for each user

Inventive Principle:
Principle #25Self-service

2Measurement precision

If user-specific learning data is collected and stored, then speech recognition accuracy is improved, but data storage requirements and processing complexity increase

Engineering Contradiction:
Improverecognition accuracyVSAvoiddata storage volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential and relevant features from user speech data for storage and processing. Rather than storing complete raw speech recordings, the system identifies and stores key acoustic characteristics and speech patterns that are most valuable for improving recognition accuracy, reducing the overall data volume required

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The learning data storage is organized in a distributed manner across multiple storage units, with each unit holding specific types of speech data or features. This allows for efficient data management, selective access to relevant data portions, and reduced processing overhead compared to centralized storage of all raw data

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20260011333A1Speech recognition device, speech recognition method, and program
Publication Date: 2026.01.08 NEC CORP
  • US20260011333A1 patent drawing
  • US20260011333A1 patent drawing
  • US20260011333A1 patent drawing

AI summary

A speech recognition apparatus (100) includes: a speech reproduction unit (102) that reproduces, for each predetermined section, target speech for speech recognition being divided for each predetermined section; a speech recognition unit (104) that recognizes, for each target speech, spoken speech acquired by repeating the target speech by a user; a text information generation unit (106) that generates text information about the spoken speech, based on a recognition result of the speech recognition unit (104); and a storage processing unit (108) that stores, as learning data, identification information by the user, the spoken speech, and the recognition result corresponding to the spoken speech in association with one another, in which the speech recognition unit (104) performs recognition by using a recognition engine that learns the learning data by the user.