Audio Content Recognition via Fragment Segmentation and Model Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic subtitle adding technologies for videos face challenges in accuracy, especially when dealing with diverse audio contents like singing, speaking voices, coughs, laughs, and multiple languages, making it difficult to improve user experience efficiently and cost-effectively.

Innovation Solution

An audio content recognition method and apparatus that segments audio into voice and non-voice fragments, determines the type and language of each voice fragment, and performs voice recognition using specific models based on these characteristics to improve recognition accuracy for both speaking and music fragments, and handles multiple languages effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single voice recognition model is used for all audio contents, then the device complexity is reduced, but the recognition accuracy deteriorates when dealing with diverse audio contents like singing, speaking voices, and multiple languages

Engineering Contradiction:
Improverecognition accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments audio contents into different types (speaking voice, singing voice, etc.) and selects different recognition models accordingly. This segmentation allows the system to achieve high recognition accuracy for each audio type while managing model complexity through selective model application rather than using a single complex model for all cases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different recognition models to different local segments of audio based on their characteristics. Speaking voice segments use one model while singing voice segments use another, ensuring each segment receives the most appropriate processing quality for its specific type, thereby improving overall recognition accuracy without uniformly increasing complexity.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If artificial subtitle adding is used, then the recognition accuracy is improved, but the productivity and cost deteriorate

Engineering Contradiction:
Improvesubtitle accuracyVSAvoidsubtitle generation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements an automatic recognition system that processes audio contents without human intervention. The system automatically segments audio, identifies audio types, selects appropriate models, and generates subtitles, enabling self-service operation that maintains high accuracy while dramatically improving productivity and reducing costs compared to manual subtitle adding.

Inventive Principle:
Principle #25Self-service

3Productivity

If existing automatic subtitle adding technology is used, then the productivity is improved, but the recognition accuracy deteriorates when dealing with various audio contents

Engineering Contradiction:
Improvesubtitle generation efficiencyVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces dynamic model selection based on real-time audio content analysis. Instead of using a fixed single model, the system dynamically adjusts the recognition model according to the detected audio type (speaking, singing, etc.), maintaining high productivity through automation while improving accuracy through adaptive model selection matched to content characteristics.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If audio with fragments in multiple languages is processed using a single language model, then the device complexity is reduced, but the recognition accuracy deteriorates

Engineering Contradiction:
Improvemulti-language recognition accuracyVSAvoidlanguage model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments multi-language audio into language-specific fragments and applies appropriate language models to each fragment. This segmentation approach enables accurate recognition of multiple languages while managing complexity by only loading and applying the specific language model needed for each audio segment rather than maintaining all language models simultaneously.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11783808B2Audio content recognition method and apparatus, and device and computer-readable medium
Publication Date: 2023.10.10 DOUYIN VISION CO LTD
  • US11783808B2 patent drawing
  • US11783808B2 patent drawing
  • US11783808B2 patent drawing

AI summary

Embodiments of the present disclosure disclose an audio content recognition method and apparatus, an electronic device and a non-transitory computer-readable medium. A specific implementation of the method includes: obtaining a voice fragment collection and a non-voice fragment collection by segmenting audio; determining a type and language information of each voice fragment in the voice fragment collection; obtaining, for each voice fragment in the voice fragment collection, a first recognition result by performing voice recognition on the voice fragment based on the type and the language information of the voice fragment. In the implementation, speaking and music fragments in the audio are recognized by different models, so that two audio contents may both have better recognition effects. Moreover, audio of different language contents is recognized by using different models, thereby further improving a voice recognition effect.