Voice Recognition Model Tag-Based Data Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Creating a voice recognition model with high accuracy is challenging due to the difficulty in preparing a data set with a sufficient amount of voice data collected under specific conditions, such as environment and speaker characteristics.
Innovation Solution
A voice recognition device that acquires and processes voice data using a first recognition model, extracts significant tags based on recognition results, and creates a second model by combining data sets associated with these tags to enhance recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a recognition model is created based on a data set with sufficient data amount, then recognition accuracy is improved, but data preparation difficulty increases
Solution Approach 1:
The patent creates synthetic voice data by copying and transforming existing voice data characteristics. A learning model generates artificial voice data that mimics real voice patterns, allowing the system to expand the training data set without requiring additional real-world data collection. This copying approach enables sufficient data quantity for high recognition accuracy while avoiding the difficulty of collecting enough real voice samples.
Solution Approach 2:
The system performs self-service by automatically generating its own training data. The learning model uses existing voice data to create synthetic samples, eliminating the need for external data collection efforts. This self-service mechanism allows the system to independently expand its data set to achieve high recognition accuracy without relying on difficult external data preparation.
2Measurement precision
If voice data is collected under specific conditions (environment, speaker characteristics), then recognition accuracy under those conditions is improved, but data collection complexity increases
Solution Approach 1:
The patent applies local quality by generating voice data with specific local characteristics (environmental conditions, speaker features) as needed. Rather than collecting comprehensive data under all possible conditions, the learning model synthesizes voice data with particular local qualities required for specific recognition scenarios, reducing data collection complexity while maintaining high accuracy for target conditions.
Solution Approach 2:
The system uses parameter changes to transform existing voice data into synthetic samples with different environmental and speaker characteristics. By adjusting parameters such as background noise levels, speaker identity features, and recording conditions during synthesis, the system can generate data for specific conditions without physically collecting data under those conditions, thereby reducing collection complexity.
Data Source
AI summary
According to one embodiment, a recognition device includes storage and a processor. The storage is configured to store a first recognition model, a first data set, and tags, for each first recognition model. The processor is configured to acquire a second data set, execute recognition processing of the second recognition target data in the second data set by using the first recognition model, extract a significant tag of the tags stored in the storage in association with the first recognition model, based on the recognition processing result and the second correct data in the second data set, and create a second recognition model based on the acquired second data set and the first data set stored in the storage in association with the extracted tag.


