Feature Information Mining Using Simulated Audio Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face inefficiencies and high costs in feature information mining due to limited audio data, leading to low accuracy and long training times, especially in diverse usage scenarios.
Innovation Solution
A method that simulates usage scenarios to generate target audio data by combining real, synthesized, and recorded audio data, along with environmental noise, to iteratively extract feature information, enhancing coverage and reducing false alarms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large amount of audio data is collected for speech model training, then the accuracy of speech recognition is improved, but the time-consuming and cost of data collection and processing increases significantly
Solution Approach 1:
The patent applies preliminary action by pre-collecting and organizing audio data from multiple sources (real scenario data, speech synthesis data, recorded audio data) before actual model training. This pre-prepared data pool enables rapid model iteration without repeated data collection, significantly reducing time consumption while maintaining data quality for accurate speech recognition.
Solution Approach 2:
The patent uses speech synthesis data to generate synthetic audio samples that copy and augment real speech patterns. This copying approach creates additional training data without requiring proportional increases in real data collection time, thereby improving model accuracy while controlling data processing time.
2Reliability
If diverse usage scenario data is collected to improve model adaptability, then the reliability of speech recognition is improved, but the complexity of data management and processing increases
Solution Approach 1:
The patent segments diverse usage scenario data into distinct categories (real scenario data, speech synthesis data, recorded audio data, other media data). This segmentation allows organized management of different data types with specific processing pipelines for each, reducing overall management complexity while maintaining comprehensive scenario coverage for reliable speech recognition.
Solution Approach 2:
The patent creates a universal data processing framework that handles multiple data types (audio, text, media) through integrated pipelines. This multi-functional system manages diverse usage scenario data uniformly, reducing complexity by providing a single management interface while maintaining the ability to process various data formats for improved model reliability.
3Reliability
If extensive audio data is accumulated to reduce false alarms, then the robustness of speech recognition is improved, but the cost of data storage and processing increases
Solution Approach 1:
The patent creates a composite data structure that integrates multiple data types (real scenario audio, synthesized speech, recorded data, and other media) into a unified training dataset. This composite approach achieves robust speech recognition with reduced false alarms by leveraging the complementary strengths of different data sources, rather than relying solely on large volumes of single-type audio data.
Solution Approach 2:
The patent transforms data parameters by converting text to synthesized speech, transforming one type of media into another. This parameter transformation creates diverse training samples from limited source material, reducing the total amount of raw audio data needed while maintaining robustness against false alarms through varied data representations.
Data Source
AI summary
A method for mining feature information, an apparatus for mining feature information and an electronic device are disclosed. The method includes: determining a usage scenario of a target device; obtaining raw audio data including real scenario data, speech synthesis data, recorded audio data and other media data; generating target audio data of the usage scenario by simulating the usage scenario based on the raw audio data; and obtaining feature information of the usage scenario by performing feature extraction on the target audio data.


