Voice Recognition Model Training via Representative Device Data Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition models trained in optimal environments often fail to perform optimally in real-world environments due to noise and obstacles, necessitating updates that account for actual utterance environments, and require efficient data for training across multiple devices.

Innovation Solution

An electronic device can efficiently update its voice recognition model using voice data and recognition results from a representative device, which can then share these updates with other devices to improve recognition accuracy, even in noisy or distant environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a speech recognition model is trained using data collected in an optimal environment (soundproof, no obstacles), then the training data quality is high and the model achieves good baseline performance, but the model fails to perform optimally in real-world environments with noise and obstacles

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by collecting voice data from multiple electronic devices in their respective real-world environments before model training. This advance data collection from diverse environments (noisy, distant, obstructed) allows the model to be pre-exposed to various conditions, improving its adaptability while maintaining training quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses a representative electronic device that can serve multiple functions: it acts as both a regular consumer device and as a data collection platform for training other devices' models. The collected data from this representative device can be universally applied to improve multiple speech recognition models across different devices and environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If speech recognition models are trained independently for each electronic device to account for device-specific environments, then each device achieves optimal performance in its environment, but the data collection and model training process becomes complex and time-consuming

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates a copy of the speech recognition model trained on representative device data and applies it to other electronic devices. Instead of training each device independently, the model learned from the representative device's real-world environment is copied and adapted to improve other devices, significantly reducing training complexity while maintaining environmental adaptability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system merges the training process by using data collected from multiple electronic devices to train a single representative model. This consolidated approach combines environmental characteristics from various devices into one model that can then be applied universally, reducing the overall complexity of training multiple separate models.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If data is collected from multiple electronic devices to improve model adaptability, then the model becomes more versatile across different environments, but the amount of data to be processed and the coordination required increases significantly

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoiddata processing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system introduces an intermediary approach by using a representative electronic device as a mediator between multiple source devices and the final model application. Data from multiple devices is aggregated and processed through this representative device's model, which then serves as an intermediary solution for improving other devices, reducing direct processing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The representative electronic device and its collected data serve a universal purpose: the data collected from this single device can be used to improve speech recognition models across multiple other devices. This multi-functional use of the representative device's data maximizes productivity by avoiding redundant data collection and processing for each individual device.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3785258B1Electronic device and method for providing or obtaining data for training thereof
Publication Date: 2024.11.27 SAMSUNG ELECTRONICS CO LTD
  • EP3785258B1 patent drawingFigure 1a
  • EP3785258B1 patent drawingFigure 1b
  • EP3785258B1 patent drawingFigure 2

AI summary

Methods for providing and obtaining data for training and electronic devices thereof are provided. The method for providing data for training includes obtaining first voice data for a voice uttered by a user at a specific time through a microphone of the electronic device and transmitting the voice recognition result to a second electronic device which obtained second voice data for the voice uttered by the user at the specific time, for use as data for training a voice recognition model. In this case, the voice recognition model may be trained using the data for training and an artificial intelligence algorithm such as deep learning.