Multi-Modal Key Segment Detection in Audio Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting key segments in audio or video require manual review and selection, which is inefficient and costly, and most users lack the capability to accurately identify key segments.
Innovation Solution
A method that extracts multi-modal features (visual, acoustic, and natural language) from audio or video, determines candidate key segments, and uses automatic speech recognition (ASR) text to filter and identify key segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review and selection methods are used to detect key segments, then users can identify key segments with human judgment, but the process is inefficient and costly
Solution Approach 1:
The system performs automatic key segment detection without requiring manual user intervention. The multi-modal analysis system independently identifies key segments by processing visual, acoustic, and natural language features, eliminating the need for users to manually review and select key segments while maintaining high accuracy through automated multi-modal feature fusion
Solution Approach 2:
The patent replaces manual mechanical review processes with an automated computational system. Instead of users manually watching and selecting key segments, the system uses multi-modal feature extraction and fusion algorithms to automatically detect key segments, substituting human labor with machine-based automated analysis
2Reliability
If manual review methods are used, then users can make subjective judgments about key segments, but the process requires significant time and cost investment
Solution Approach 1:
The system performs preliminary feature extraction and analysis on all video segments before final key segment identification. By pre-processing visual, acoustic, and natural language features and fusing them in advance, the system prepares candidate key segments for rapid identification, reducing the time required for final detection while maintaining reliable quality through comprehensive multi-modal analysis
Solution Approach 2:
The patent extracts and separates different modalities of features (visual, acoustic, natural language) into distinct processing streams. Each modality is independently analyzed and then fused to identify key segments, allowing the system to efficiently process and evaluate multiple feature types without requiring complete manual review of the entire video content
3Productivity
If automated multi-modal analysis is used to detect key segments, then efficiency and accuracy are improved, but the system complexity increases
Solution Approach 1:
The system divides the complex task of key segment detection into separate processing modules for different modalities: visual feature extraction, acoustic feature extraction, and natural language feature extraction. Each modality is processed independently through its own feature extraction pipeline, and the results are subsequently fused to identify key segments, making the overall complex system manageable through modular segmentation
Data Source
AI summary
The present disclosure provides a method, a device, a computer-readable storage medium, and a computer program product for detecting key segments in an audio or video. The method includes: obtaining multi-modal features of an audio or video, where the multi-modal features include a visual feature, an acoustic feature, and a natural language feature; determining candidate key segments in the audio or video based on the multi-modal features; obtaining a keyword list based on automatic speech recognition (ASR) text of the candidate key segments; and determining a key segment in the audio or video based on the keyword list.


