Multi-modal video feature extraction method and device, equipment and medium
Through the methods of hierarchical multimodal decomposition and adaptive sparse coding, the problem of insufficient feature expression in multimodal video feature extraction is solved, and efficient compression and improved discriminability of multimodal video features are achieved, and the ability to adapt to dynamic scenes is enhanced.
Patent Information
- Application Number
- CN202510828110.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-16
AI Technical Summary
The multimodal video feature extraction methods in existing technologies are difficult to adapt to changes in video content in dynamic scenes, resulting in insufficient feature expression and unable to meet the needs of application scenarios such as urban management, medical prediction and financial prediction.
The original multimodal data is obtained through the acquisition equipment. After preprocessing, the hierarchical multimodal decomposition technology is used to extract multi-scale features. Adaptive sparse coding is used to perform joint sparse representation optimization, and the sparse threshold and dictionary are dynamically adjusted to generate a compact multimodal feature vector.
It significantly improves the expressive power of multimodal video features and their ability to adapt to dynamic scenes, solves the problem of insufficient feature expression, and enhances the execution effect of downstream tasks.
Smart Images

Figure CN120656107A_ABST
Abstract
Claims
1. A multimodal video feature extraction method, characterized in that: include: Obtaining original multimodal data through acquisition equipment; Preprocessing the original multimodal data; Performing multimodal decomposition on the preprocessed original multimodal data to generate hierarchical features of each modality; Performing joint sparse representation optimization on the hierarchical features through adaptive sparse coding to obtain optimized sparse coding; The optimized sparse coding is integrated to generate a compact multimodal feature vector for output to downstream tasks.
2. The multimodal video feature extraction method according to claim 1, wherein: The step of performing joint sparse representation optimization on the hierarchical features by adaptive sparse coding to obtain optimized sparse coding includes: Dynamically adjust the sparse threshold based on scene complexity, and generate sparse codes for each modality according to the hierarchical features and the sparse threshold through online dictionary learning; The optimized sparse coding is obtained by optimizing the sparse coding by transferring saliency information between the sparse codings of each modality through a cross-modal attention mechanism.
3. The multimodal video feature extraction method according to claim 2, wherein: The step of dynamically adjusting the sparse threshold based on scene complexity and generating sparse codes of each modality according to the hierarchical features and the sparse threshold through online dictionary learning includes: By analyzing the number of moving objects, audio energy change rate and text keyword density in the hierarchical features, a scene complexity value is comprehensively calculated; adjusting the sparse threshold according to the scene complexity value; The hierarchical features are used as input, and a sparse code is generated by iteratively optimizing an objective function according to the sparse threshold.
4. The multimodal video feature extraction method according to claim 3, wherein: The steps of optimizing the sparse coding by transferring saliency information between the sparse codings of each modality through a cross-modal attention mechanism to obtain the optimized sparse coding include: Projecting the sparse codes of each modality into a query vector, a key vector, and a value vector respectively; Calculating an inter-modality attention weight based on the query vector, the key vector, and the value vector; Based on the attention weight weighted fusion value vector, the optimized sparse code is generated.
5. The multimodal video feature extraction method according to claim 1, wherein: The original multimodal data includes a video stream, an audio stream, and a text stream, and the step of preprocessing the original multimodal data includes: Performing frame alignment and color normalization processing on the video stream to obtain a standardized video stream; Performing noise reduction and spectrum normalization processing on the audio stream to obtain a standardized audio stream; The text stream is subjected to word segmentation and embedding vectorization processing to obtain a standardized text stream.
6. The multimodal video feature extraction method according to claim 5, characterized in that: The hierarchical features include multi-resolution spatial features, time-frequency domain energy distribution features, and phrase-level semantic unit features. The step of performing multimodal decomposition on the pre-processed original multimodal data includes: Performing multi-scale spatial decomposition on the standardized video stream, and generating the multi-resolution spatial features through a Laplacian pyramid; Performing frequency band energy decomposition on the standardized audio stream, and generating the time-frequency domain energy distribution feature through a learnable Mel filter bank; The standardized text stream is semantically segmented, and the phrase-level semantic unit features are generated through phrase-level division of a pre-trained language model.
7. The multimodal video feature extraction method according to claim 1, wherein: The step of fusing the optimized sparse coding to generate a compact multimodal feature vector for output to a downstream task includes: Splicing the optimized sparse codes of each modality along the feature dimension to form an initial fusion feature; Performing a linear transformation on the initial fusion features through a lightweight fully connected layer to generate a compact multimodal feature vector; The multimodal feature vector is input into the downstream task module.
8. A multimodal video feature extraction device, characterized in that: Used to perform the multimodal video feature extraction method as described in any one of claims 1 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 can be implemented.
Citation Information
Cited By
Lightweight storage method and system for time series data of Internet of Things
CN121255109A