Speech Recognition Correction Database for Diverse User Voices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face difficulties in recognizing speech commands from young children, individuals with strong dialects, and those with poor pronunciation, as well as generating learning data suitable for all speakers, leading to suboptimal performance in diverse user scenarios.
Innovation Solution
An artificial intelligence device equipped with a database to store correction data for speech commands, a microphone to receive and store unclear commands, and a processor to acquire and map correction data from authorized users, enabling improved speech recognition by retrieving and applying similar patterned commands and text data for intention analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cloud server processes speech data in real time with vast database, then speech recognition capability is improved, but difficulty in recognizing speeches from children, strong dialects, and poor pronunciation remains
Solution Approach 1:
The system performs preliminary actions by collecting speech samples from multiple users beforehand and generating correction data in advance. The database stores pre-processed correction data that can be quickly retrieved during speech recognition, avoiding the need to process diverse speech patterns in real-time while maintaining high accuracy for various speakers including children and dialect speakers.
Solution Approach 2:
The system creates copies of speech data by generating correction data from multiple users' speech samples. These correction data copies are stored in the database and can be used to improve recognition accuracy for similar speech patterns, allowing the system to handle diverse speech characteristics without processing each unique pattern individually.
2Productivity
If speech data is converted to text and analyzed in cloud server, then real-time search result is provided, but speech recognition fails for speeches hard to recognize
Solution Approach 1:
The system implements feedback by using correction data from multiple users to improve recognition of difficult speech patterns. When speech recognition is performed, the system retrieves relevant correction data from the database and uses it to adjust and improve recognition accuracy, creating a feedback loop that enhances reliability without sacrificing real-time processing capability.
Solution Approach 2:
The system performs preliminary processing by pre-collecting and pre-processing speech samples from multiple users to generate correction data before actual speech recognition is needed. This preliminary action ensures that reliable correction data is already available in the database, enabling both high-speed real-time processing and improved reliability for difficult speech patterns.
3Adaptability or versatility
If learning data is generated for all speakers, then speech recognition coverage is improved, but difficulty in generating and applying suitable learning data remains
Solution Approach 1:
The system achieves universality by collecting speech samples from multiple different users and generating correction data that can serve all speakers. The database stores correction data that is universally applicable to various speech patterns, eliminating the need to create separate learning data for each individual speaker while maintaining broad coverage and adaptability.
Solution Approach 2:
The system creates copies of correction data from multiple users and stores them in the database. These copied correction data can be retrieved and applied to improve recognition for any speaker, providing broad coverage without the complexity of generating custom learning data for each individual. The copying approach simplifies the process while maintaining versatility.
Data Source
AI summary
An artificial intelligence device for performing speech recognition includes a database configured to store correction data replacing a predetermined speech command, a microphone configured to receive a first speech command from a first user, and a processor configured to store the first speech command in the database when operation to be performed with respect to the first speech command is not determined, acquire correction data replacing the first speech command from a second user, and map and store the first speech command and the correction data in the database.


