Robot Learning Data Collection for Environment-Matched Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition models struggle to accurately adapt to diverse user environments due to variations in user speech tendencies, gender, and surrounding noise, leading to performance issues.
Innovation Solution
A robot equipped with a speaker, microphone, driver, communication interface, and processor actively collects learning data by controlling external devices to output noise and user speech, using pre-stored environment information to learn a speech recognition model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If virtual learning data is generated using noise and reverberation from large-capacity databases, then the quantity of learning data is increased, but the accuracy of speech recognition in actual user environments deteriorates due to environmental mismatches
Solution Approach 1:
The system performs preliminary actions by having external devices output noise and user speech in advance before actual speech recognition tasks. This creates pre-collected learning data that reflects the specific user environment, allowing the speech recognition model to be trained with environment-matched data rather than generic database noise.
Solution Approach 2:
The system changes the parameters of learning data by collecting actual noise characteristics from the user's specific environment (residence type, use space, surrounding noise) rather than using standardized noise from databases. This parameter change ensures the learning data matches the actual operating conditions, improving recognition accuracy.
2Adaptability or versatility
If speech recognition models are trained on generic noise data, then the model can be applied broadly, but it fails to adapt to specific user environments with unique noise characteristics and speech tendencies
Solution Approach 1:
The system applies local quality by collecting noise and speech data specific to each user's local environment (their home, surrounding noise, residence type) rather than using uniform generic data. This localized approach allows the speech recognition model to adapt to specific environmental characteristics while maintaining reliability in that particular context.
Solution Approach 2:
The system uses feedback by having the robot collect actual noise and speech data from the user environment, use this data to train the speech recognition model, and then improve performance based on the results. This closed-loop feedback mechanism enables continuous adaptation to the specific user environment.
3Measurement precision
If the robot collects learning data by controlling external devices to output noise and user speech, then the quality of learning data improves, but the device complexity and control requirements increase
Solution Approach 1:
The system applies universality by using a standardized control protocol that can work with multiple types of external devices (smartphones, computers, smart speakers) regardless of their specific functions. The robot sends generic commands to output noise and speech, and these commands can be executed by various devices, reducing the complexity of device-specific control while maintaining high learning data quality.
Data Source
AI summary
A robot transmits a command to control an external device around the robot based on pre-stored environment information while the robot is operating in a learning mode. The external device makes a noise as part of its operation. Also, the robot outputs user speech for learning while the external device is operating. The robot learns a speech recognition model based on the noise and speech of a user acquired through a microphone of the robot. The speech recognition model is then used by the robot or by another device to better understand the user when the user talks. The robot is then able to more accurately understand and properly execute speech commands from the user.


