User-Adaptive Speech Recognition for Repeated Speech Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition accuracy varies due to individual differences in repeated speech, leading to inconsistent recognition results.
Innovation Solution
A speech recognition apparatus that divides target speech into sections, reproduces each section, recognizes repeated speech by a user, generates text information, and stores user-specific learning data for improved recognition using a user-tailored recognition engine.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition is performed on repeated speech without user-specific adaptation, then the system complexity is low, but speech recognition accuracy varies due to individual differences
Solution Approach 1:
The system performs preliminary speech collection and learning before actual recognition tasks. User-specific speech data is collected and stored in advance, creating a personalized learning database that the recognition engine can utilize during transcription tasks to improve accuracy without adding complexity during the actual recognition process
Solution Approach 2:
The recognition engine automatically learns from user-specific speech data and adapts to individual speech patterns without requiring manual configuration or intervention. The system self-improves by utilizing the collected learning data to enhance its recognition capabilities for each user
2Measurement precision
If user-specific learning data is collected and stored, then speech recognition accuracy is improved, but data storage requirements and processing complexity increase
Solution Approach 1:
The system extracts only the essential and relevant features from user speech data for storage and processing. Rather than storing complete raw speech recordings, the system identifies and stores key acoustic characteristics and speech patterns that are most valuable for improving recognition accuracy, reducing the overall data volume required
Solution Approach 2:
The learning data storage is organized in a distributed manner across multiple storage units, with each unit holding specific types of speech data or features. This allows for efficient data management, selective access to relevant data portions, and reduced processing overhead compared to centralized storage of all raw data
Data Source
AI summary
A speech recognition apparatus (100) includes: a speech reproduction unit (102) that reproduces, for each predetermined section, target speech for speech recognition being divided for each predetermined section; a speech recognition unit (104) that recognizes, for each target speech, spoken speech acquired by repeating the target speech by a user; a text information generation unit (106) that generates text information about the spoken speech, based on a recognition result of the speech recognition unit (104); and a storage processing unit (108) that stores, as learning data, identification information by the user, the spoken speech, and the recognition result corresponding to the spoken speech in association with one another, in which the speech recognition unit (104) performs recognition by using a recognition engine that learns the learning data by the user.


