Speech Recognition Training Set Generation via Audio-Video Text Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning-based automatic speech recognition (ASR) models require extensive and manually labeled speech data for improved generalization, which is time-consuming and inefficient.

Innovation Solution

A method and apparatus for generating a speech recognition training set by acquiring audio and video data where the video includes text information, recognizing the audio to obtain audio text, recognizing video text using OCR, and using the consistency between audio and video texts to create the training set, thereby automating the process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual labeling is used to construct speech recognition training sets, then labeling accuracy can be ensured, but the construction process becomes time-consuming and inefficient

Engineering Contradiction:
Improvetraining set construction efficiencyVSAvoidtime for manual labeling
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system uses its own video recognition capabilities to automatically generate text information from video content, which then serves as the ground truth for labeling speech audio. This self-service mechanism eliminates the need for external manual labeling while maintaining high accuracy, directly resolving the contradiction between construction efficiency and time consumption.

Inventive Principle:
Principle #25Self-service

2Reliability

If extensive speech data is collected to improve generalization performance, then model robustness improves, but data acquisition and processing complexity increases

Engineering Contradiction:
Improvemodel generalization performanceVSAvoiddata acquisition and processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system processes both video and audio data through a unified framework where video recognition automatically generates text labels that are then used to supervise speech recognition training. This multi-functional approach allows the same system to handle diverse data types (video, audio, text) and improves model generalization across different modalities while managing complexity through integration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240233708A1Method and device for generating speech recognition training set
Publication Date: 2024.07.11 JINGDONG TECH HLDG CO LTD
  • US20240233708A1 patent drawing
  • US20240233708A1 patent drawing
  • US20240233708A1 patent drawing

AI summary

Disclosed in the present disclosure are a method and apparatus for generating a speech recognition training set. The method may include: acquiring a to-be-processed audio and a to-be-processed video, where the to-be-processed video comprises text information corresponding to the to-be-processed audio; recognizing the to-be-processed audio to obtain an audio text; recognizing text information in the to-be-processed video to obtain a video text; and using, based on consistency of the audio text with the video text, the to-be-processed audio as a speech sample and the video text as a label to obtain the speech recognition training set.