Multi-Modal Key Segment Detection in Audio Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting key segments in audio or video require manual review and selection, which is inefficient and costly, and most users lack the capability to accurately identify key segments.

Innovation Solution

A method that extracts multi-modal features (visual, acoustic, and natural language) from audio or video, determines candidate key segments, and uses automatic speech recognition (ASR) text to filter and identify key segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review and selection methods are used to detect key segments, then users can identify key segments with human judgment, but the process is inefficient and costly

Engineering Contradiction:
Improveaccuracy of key segment identificationVSAvoidefficiency of key segment detection
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs automatic key segment detection without requiring manual user intervention. The multi-modal analysis system independently identifies key segments by processing visual, acoustic, and natural language features, eliminating the need for users to manually review and select key segments while maintaining high accuracy through automated multi-modal feature fusion

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical review processes with an automated computational system. Instead of users manually watching and selecting key segments, the system uses multi-modal feature extraction and fusion algorithms to automatically detect key segments, substituting human labor with machine-based automated analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual review methods are used, then users can make subjective judgments about key segments, but the process requires significant time and cost investment

Engineering Contradiction:
Improvequality of key segment selectionVSAvoidtime required for key segment detection
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary feature extraction and analysis on all video segments before final key segment identification. By pre-processing visual, acoustic, and natural language features and fusing them in advance, the system prepares candidate key segments for rapid identification, reducing the time required for final detection while maintaining reliable quality through comprehensive multi-modal analysis

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and separates different modalities of features (visual, acoustic, natural language) into distinct processing streams. Each modality is independently analyzed and then fused to identify key segments, allowing the system to efficiently process and evaluate multiple feature types without requiring complete manual review of the entire video content

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If automated multi-modal analysis is used to detect key segments, then efficiency and accuracy are improved, but the system complexity increases

Engineering Contradiction:
Improveautomation level of key segment detectionVSAvoidsystem complexity for multi-modal feature processing
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the complex task of key segment detection into separate processing modules for different modalities: visual feature extraction, acoustic feature extraction, and natural language feature extraction. Each modality is processed independently through its own feature extraction pipeline, and the results are subsequently fused to identify key segments, making the overall complex system manageable through modular segmentation

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250182746A1Method, device and medium for detecting key segments in audio or video
Publication Date: 2025.06.05 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250182746A1 patent drawing
  • US20250182746A1 patent drawing
  • US20250182746A1 patent drawing

AI summary

The present disclosure provides a method, a device, a computer-readable storage medium, and a computer program product for detecting key segments in an audio or video. The method includes: obtaining multi-modal features of an audio or video, where the multi-modal features include a visual feature, an acoustic feature, and a natural language feature; determining candidate key segments in the audio or video based on the multi-modal features; obtaining a keyword list based on automatic speech recognition (ASR) text of the candidate key segments; and determining a key segment in the audio or video based on the keyword list.