Voice Recognition Model Training via Cross-Device Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition models trained in optimal environments often fail to provide optimal performance in real-world scenarios due to noise, distance, and obstacles, necessitating the need for data that accounts for actual utterance environments to effectively update these models.
Innovation Solution
A method where one electronic device captures voice data and recognition results, transmitting them to another device for use in training its speech recognition model, allowing for the sharing and updating of voice recognition models across multiple devices to improve accuracy in diverse environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition models are trained in optimal environments with good soundproofing and no obstacles, then training data quality is improved, but the models fail to perform optimally in real-world noisy and distant environments
Solution Approach 1:
The system performs preliminary voice recognition in optimal environments (first electronic device with good soundproofing) to obtain accurate recognition results before the actual utterance occurs. These preliminary results are then transmitted to devices in suboptimal environments for model training, allowing the models to learn from high-quality reference data while adapting to various environmental conditions.
Solution Approach 2:
The patent introduces an intermediary mechanism (communication between electronic devices) that transfers voice data and recognition results from devices in optimal environments to devices in suboptimal environments. This intermediary system enables the sharing of training data across different environmental conditions, allowing models to be trained on high-quality data while improving performance in diverse real-world settings.
2Adaptability or versatility
If speech recognition models are trained using data from multiple electronic devices in diverse environments, then environmental adaptability is improved, but data collection complexity increases
Solution Approach 1:
The patent creates a universal data sharing system where electronic devices can serve multiple functions: acting as both data collectors in their own environments and as recipients of reference data from other devices. This multi-functionality allows the system to gather diverse environmental data without requiring each device to independently collect all types of training data, simplifying the overall data collection architecture.
Solution Approach 2:
The system uses copying by transmitting voice data and recognition results from one electronic device to another. Instead of requiring each device to independently generate training data in all possible environments, devices copy relevant training data from peers who have captured it in optimal conditions, significantly reducing the complexity of data collection while maintaining environmental diversity.
3Reliability
If each electronic device collects its own voice data for training, then data representativeness is improved, but training data quantity is insufficient for devices in suboptimal environments
Solution Approach 1:
The patent merges training data resources across multiple electronic devices by transmitting voice data and recognition results between them. Devices in suboptimal environments combine their own locally collected data with reference data received from devices in optimal environments, creating a larger and more representative training dataset than any single device could accumulate independently.
Solution Approach 2:
The system adds a new dimension to data collection by utilizing spatial distribution across multiple devices. Instead of relying solely on temporal accumulation of data from a single device, the patent leverages the spatial dimension by collecting data simultaneously from multiple devices in different environments and combining them, thereby increasing both data quantity and representativeness efficiently.
Data Source
AI summary
Methods for providing and obtaining data for training and electronic devices thereof are provided. The method for providing data for training includes obtaining first voice data for a voice uttered by a user at a specific time through a microphone of the electronic device and transmitting the voice recognition result to a second electronic device which obtained second voice data for the voice uttered by the user at the specific time, for use as data for training a voice recognition model. In this case, the voice recognition model may be trained using the data for training and an artificial intelligence algorithm such as deep learning.


