Speech Recognition Model Pretraining via Near-Field Data and Chat Text
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately recognizing speech data collected in Far Field environments due to low signal-to-noise ratios and require large amounts of learning data for effective training, which is difficult and costly to collect.
Innovation Solution
The system generates first and second learning data by processing speech data from Near Field and Far Field environments, respectively, and uses these data sets for pretraining and fine-tuning the speech recognition model, enabling improved accuracy in Far Field speech recognition without relying on manually annotated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech data is collected in Far Field environment, then the system can capture more speech scenarios, but the signal-to-noise ratio decreases and recognition accuracy deteriorates
Solution Approach 1:
The system performs preliminary action by collecting Near Field speech data in advance to create pretraining learning data. This pretraining phase prepares the speech recognition model before actual Far Field recognition tasks, enabling the model to adapt to Far Field conditions without requiring manual annotation of Far Field data. The preliminary collection of high-quality Near Field data resolves the contradiction by providing a foundation that improves Far Field recognition accuracy.
2Reliability
If large amounts of learning data are collected for model training, then speech recognition accuracy improves, but data collection cost and difficulty increase
Solution Approach 1:
The system applies copying by using text data from chat services as a substitute for manually annotated speech recognition data. Instead of collecting and annotating actual speech data, the system copies existing text data and uses it as learning data for the speech recognition model. This copying approach maintains high recognition accuracy while dramatically reducing data collection costs and effort.
Solution Approach 2:
The system implements self-service by automatically generating pretraining learning data from readily available chat service text data without requiring manual intervention for data collection or annotation. The process autonomously converts text data into suitable learning formats, eliminating the need for expensive manual labor in data preparation while still providing sufficient training data for accurate speech recognition.
Data Source
AI summary
An information processing method includes obtaining speech data based on a distance between a sound collection device and a speaker, obtaining text data input in a service for exchanging messages, and outputting first learning data that is based on the speech data and second learning data that includes the text data.


