Audio Content Recognition via Fragment Segmentation and Model Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic subtitle adding technologies for videos face challenges in accuracy, especially when dealing with diverse audio contents like singing, speaking voices, coughs, laughs, and multiple languages, making it difficult to improve user experience efficiently and cost-effectively.
Innovation Solution
An audio content recognition method and apparatus that segments audio into voice and non-voice fragments, determines the type and language of each voice fragment, and performs voice recognition using specific models based on these characteristics to improve recognition accuracy for both speaking and music fragments, and handles multiple languages effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single voice recognition model is used for all audio contents, then the device complexity is reduced, but the recognition accuracy deteriorates when dealing with diverse audio contents like singing, speaking voices, and multiple languages
Solution Approach 1:
The patent segments audio contents into different types (speaking voice, singing voice, etc.) and selects different recognition models accordingly. This segmentation allows the system to achieve high recognition accuracy for each audio type while managing model complexity through selective model application rather than using a single complex model for all cases.
Solution Approach 2:
The patent applies different recognition models to different local segments of audio based on their characteristics. Speaking voice segments use one model while singing voice segments use another, ensuring each segment receives the most appropriate processing quality for its specific type, thereby improving overall recognition accuracy without uniformly increasing complexity.
2Measurement precision
If artificial subtitle adding is used, then the recognition accuracy is improved, but the productivity and cost deteriorate
Solution Approach 1:
The patent implements an automatic recognition system that processes audio contents without human intervention. The system automatically segments audio, identifies audio types, selects appropriate models, and generates subtitles, enabling self-service operation that maintains high accuracy while dramatically improving productivity and reducing costs compared to manual subtitle adding.
3Productivity
If existing automatic subtitle adding technology is used, then the productivity is improved, but the recognition accuracy deteriorates when dealing with various audio contents
Solution Approach 1:
The patent introduces dynamic model selection based on real-time audio content analysis. Instead of using a fixed single model, the system dynamically adjusts the recognition model according to the detected audio type (speaking, singing, etc.), maintaining high productivity through automation while improving accuracy through adaptive model selection matched to content characteristics.
4Measurement precision
If audio with fragments in multiple languages is processed using a single language model, then the device complexity is reduced, but the recognition accuracy deteriorates
Solution Approach 1:
The patent segments multi-language audio into language-specific fragments and applies appropriate language models to each fragment. This segmentation approach enables accurate recognition of multiple languages while managing complexity by only loading and applying the specific language model needed for each audio segment rather than maintaining all language models simultaneously.
Data Source
AI summary
Embodiments of the present disclosure disclose an audio content recognition method and apparatus, an electronic device and a non-transitory computer-readable medium. A specific implementation of the method includes: obtaining a voice fragment collection and a non-voice fragment collection by segmenting audio; determining a type and language information of each voice fragment in the voice fragment collection; obtaining, for each voice fragment in the voice fragment collection, a first recognition result by performing voice recognition on the voice fragment based on the type and the language information of the voice fragment. In the implementation, speaking and music fragments in the audio are recognized by different models, so that two audio contents may both have better recognition effects. Moreover, audio of different language contents is recognized by using different models, thereby further improving a voice recognition effect.


