Speech Recognition Training Set Generation via Audio-Video Text Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep learning-based automatic speech recognition (ASR) models require extensive and manually labeled speech data for improved generalization, which is time-consuming and inefficient.
Innovation Solution
A method and apparatus for generating a speech recognition training set by acquiring audio and video data where the video includes text information, recognizing the audio to obtain audio text, recognizing video text using OCR, and using the consistency between audio and video texts to create the training set, thereby automating the process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual labeling is used to construct speech recognition training sets, then labeling accuracy can be ensured, but the construction process becomes time-consuming and inefficient
Solution Approach 1:
The system uses its own video recognition capabilities to automatically generate text information from video content, which then serves as the ground truth for labeling speech audio. This self-service mechanism eliminates the need for external manual labeling while maintaining high accuracy, directly resolving the contradiction between construction efficiency and time consumption.
2Reliability
If extensive speech data is collected to improve generalization performance, then model robustness improves, but data acquisition and processing complexity increases
Solution Approach 1:
The system processes both video and audio data through a unified framework where video recognition automatically generates text labels that are then used to supervise speech recognition training. This multi-functional approach allows the same system to handle diverse data types (video, audio, text) and improves model generalization across different modalities while managing complexity through integration.
Data Source
AI summary
Disclosed in the present disclosure are a method and apparatus for generating a speech recognition training set. The method may include: acquiring a to-be-processed audio and a to-be-processed video, where the to-be-processed video comprises text information corresponding to the to-be-processed audio; recognizing the to-be-processed audio to obtain an audio text; recognizing text information in the to-be-processed video to obtain a video text; and using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain the speech recognition training set.


