Compressed video action recognition method based on cross-modal progressive CLIP
Through the cross-modal progressive CLIP feature extraction network combined with the motion vector and residual features and text description of compressed video, the problem of degradation of action recognition accuracy caused by the loss of background information in compressed video is solved, and a more efficient behavior recognition effect is achieved.
Patent Information
- Application Number
- CN202510234814.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing compressed video behavior recognition method causes background information to be lost due to the reduction of I-frames, making it difficult to accurately capture and understand complex action scenes, affecting the recognition accuracy of actions and their contexts.
The compressed video action recognition method based on cross-modal progressive CLIP is adopted, and the cross-modal progressive CLIP feature extraction network is used, combining the motion vector and residual features of the compressed video, and text descriptions related to the action are fused. Multi-level feature extraction and progressive feature fusion modules are used to enhance the understanding of the action and its context.
It significantly improves the accuracy and robustness of compressed video behavior recognition, maintains efficient processing while enhancing the ability to understand the actions and their context.
Smart Images

Figure CN120236325A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a compressed video action recognition method based on cross-modal progressive CLIP. Background Art
[0002] Action recognition is an important branch in the field of computer vision, focusing on automatically analyzing and understanding the actions of humans or other subjects through video data. Research in this field aims to accurately identify various actions and activities from videos, such as daily behaviors such as walking, running, jumping, as well as more complex interactive behaviors or actions related to specific tasks. This task is widely used in many scenarios such as intelligent monitoring, human-computer interaction, sports analysis, and healthcare. RGB-based video action recognition methods rely on high-resolution raw videos, but their huge storage requirements and computational costs limit their practical applications.
[0003] Compressed video behavior recognition utilizes information in the compressed domain, which not only significantly reduces storage requirements but also improves computational efficiency, because modal information such as motion vectors and residuals can be directly extracted from compressed videos, avoiding complex decoding and feature extraction processes. Combined with multimodal information in the compressed domain, the accuracy of behavior recognition can be further improved through multimodal fusion. The complementarity between different modalities enables the model to better understand and classify complex behaviors. In compressed videos, I frames (Intra-coded pictures) save static information of the background similar to RGB frames, while motion vectors and residuals record the dynamic characteristics of the action, which helps to capture the moving direction and speed of objects in the video and enhance the recognition of different actions. Therefore, while maintaining recognition accuracy as much as possible, compressed video behavior recognition greatly reduces computational and storage overheads, which is suitable for application scenarios that require efficient processing of large amounts of video data, such as smart city monitoring systems and large-scale video database retrieval. Compressed video behavior recognition methods provide new possibilities for the practical deployment and widespread application of behavior recognition technology.
[0004] Although the method of compressed video behavior recognition has many advantages, due to the reduction of I-frames in compressed videos, there is often insufficient background information, which in turn affects the accuracy of recognition. For example, in the compressed domain, for actions such as applying lipstick and brushing teeth, it is difficult to accurately distinguish them without sufficient background information. To solve this problem, we propose to use a learning framework based on the CLIP model, which supplements the semantic relevance of actions by introducing text descriptions, effectively making up for the semantic deficiency caused by the lack of background information. Specifically, this method combines compressed domain features such as motion vectors and residuals in the compressed video, and integrates text descriptions related to actions, enhancing the ability to understand actions and their contexts. This multi-modal fusion strategy not only improves the recognition accuracy of the model for behaviors, but also strengthens the understanding of the scene background, thus improving the accuracy and robustness of behavior recognition while maintaining efficient processing.
[0005] On this basis, in order to more effectively extract action information from compressed videos, we introduce a progressive feature fusion module. Specifically, this module ensures that key action information is fully captured by gradually enhancing and refining low-level features such as motion vectors and residuals extracted from compressed videos. The progressive feature fusion module can dynamically adjust the feature extraction process according to the importance and complexity of actions, enabling the model to focus more on action-related details and filter out irrelevant noise. This method provides a more comprehensive and efficient solution for compressed video behavior recognition, especially suitable for application scenarios that require efficient processing of a large amount of video data.
[0006] Most existing methods for compressed video behavior recognition have not fully addressed the problem that the accuracy of action recognition significantly decreases due to the loss or simplification of background information during the video compression process. Due to the reduction of I-frames, the loss of background information makes it difficult for the model to accurately capture and understand complex action scenes, thus affecting the ability to understand actions and their contexts. Summary of the Invention
[0007] The present invention provides a method for recognizing compressed video actions based on cross-modal progressive CLIP, which solves the problem in the prior art that due to the reduction of I-frames and the loss of background information, it is difficult for the model to accurately capture and understand complex action scenes, thus affecting the ability to understand actions and their contexts, and realizes better utilization of multi-modal information to improve the overall performance of behavior recognition.
[0008] In a first aspect, the present invention provides a method for recognizing compressed video actions based on cross-modal progressive CLIP, the method comprising:
[0009] Obtain a video sequence to be recognized, and convert the video sequence to be recognized into a re-encoded video sequence;
[0010] Obtain multiple text descriptions corresponding to the video sequence to be recognized;
[0011] Input the re-encoded video sequence into the trained cross-modal progressive CLIP feature extraction network to obtain the compressed video action recognition result; wherein, the cross-modal progressive CLIP feature extraction network includes:
[0012] A data processing module, a visual encoder branch, a motion encoder branch, a detail encoder branch, a text processing branch, a progressive feature fusion module, and a contrast module;
[0013] The data processing module is used to preprocess the re-encoded video sequence to obtain a sampled GOP group;
[0014] The visual encoder branch is used to perform I-frame feature extraction on the I-frame data of the sampled GOP group based on the M-layer Transformer encoder to obtain an I-frame extraction feature set; where M > 0;
[0015] The motion encoder branch is used to perform motion feature extraction on the P-frame data of the sampled GOP group based on the F-layer Transformer encoder to obtain a motion extraction feature set; where N > 0 and N < M;
[0016] The detail encoder branch is used to perform residual feature extraction on the P-frame data of the sampled GOP group based on the F-layer Transformer encoder to obtain a residual extraction feature set;
[0017] The text processing branch is used to perform text feature extraction on multiple text descriptions respectively based on a single-layer Transformer encoder to obtain a text feature vector set;
[0018] The progressive feature fusion module fuses the I-frame extraction feature set, the motion extraction feature set, and the residual extraction feature set to obtain a fused feature vector set;
[0019] The contrast module is used to calculate the similarity between the fused feature vector set and the text feature vector set, and determine the video action in the video sequence to be recognized according to the similarity.
[0020] In a possible implementation, the data processing module is used to preprocess the re-encoded video sequence to obtain a sampled GOP group, including:
[0021] Sample the re-encoded video sequence to obtain an initial sampled GOP group;
[0022] Perform data augmentation on the initial sampled GOP group to obtain a sampled GOP group.
[0023] In a possible implementation, the visual encoder branch is used to extract I-frame features from the I-frame data of the sampled GOP group based on an M-layer Transformer encoder, obtaining an I-frame feature set, including:
[0024] Obtain the I-frame images in each GOP of the sampled GOP group to obtain a first I-frame data set;
[0025] Adjust the sizes of the I-frame images in the first I-frame data set to obtain an updated first I-frame data set;
[0026] Normalize each I-frame image in the updated first I-frame data set to obtain a normalized first I-frame data set;
[0027] Map each I-frame image in the normalized first I-frame data set to a feature space to obtain a mapped I-frame feature set;
[0028] Perform position embedding on each I-frame image in the mapped I-frame feature set to obtain a position-embedded I-frame feature set;
[0029] Extract features from each I-frame image in the position-embedded I-frame feature set based on an M-layer Transformer encoder to obtain an I-frame extraction feature set.
[0030] In a possible implementation, the motion encoder branch is used to extract motion features from the P-frame data of the sampled GOP group based on an F-layer Transformer encoder, obtaining a motion extraction feature set, including:
[0031] Obtain multiple P-frames in each GOP of the sampled GOP group to obtain a P-frame data set;
[0032] Randomly sample each P-frame data in the P-frame data set to obtain a sampled P-frame data set;
[0033] Use the coviar method to extract motion modalities from the sampled P-frame data set to obtain a motion vector set;
[0034] Adjust the sizes of the motion vectors in the motion vector set to obtain a preprocessed motion vector set;
[0035] Perform position embedding on each motion vector in the preprocessed motion vector set respectively to obtain a position-embedded motion vector set;
[0036] Use an F-layer Transformer encoder to extract motion features from each motion vector in the position-embedded motion vector set respectively to obtain a motion extraction feature set.
[0037] In a possible implementation, the detail encoder branch is used to perform residual feature extraction on the P-frame data of the sampled GOP group based on an F-layer Transformer encoder, obtaining a residual extraction feature set, including:
[0038] Obtain multiple P-frames in each GOP of the sampled GOP group to obtain a P-frame data set;
[0039] Randomly sample each P-frame data in the P-frame data set to obtain a sampled P-frame data set;
[0040] Use the coviar method to perform residual mode extraction on the sampled P-frame data set to obtain a residual set;
[0041] Adjust the magnitude of each residual in the residual set to obtain a preprocessed residual set;
[0042] Perform position embedding on each residual in the preprocessed residual set respectively to obtain a position-embedded residual set;
[0043] Use the F-layer Transformer encoder to perform residual feature extraction on each residual in the position-embedded residual set respectively to obtain a residual extraction feature set.
[0044] In a possible implementation, the text processing branch is used to perform text feature extraction on multiple text descriptions respectively based on a single-layer Transformer encoder, obtaining a text feature vector set, including:
[0045] Tokenize each text description and convert the tokenized text description into a text embedding vector;
[0046] Perform position embedding on the text embedding vector to obtain a position-embedded text embedding vector;
[0047] Input the position-embedded text embedding vector into the single-layer Transformer encoder for text feature extraction to obtain a text feature vector set.
[0048] In a possible implementation, the progressive feature fusion module fuses the I-frame extraction feature set, the motion extraction feature set, and the residual extraction feature set to obtain a fused feature vector set, including:
[0049] Add the motion vectors and residuals belonging to the same GOP in the motion extraction feature set and the residual extraction feature set value by value to obtain a first feature set;
[0050] Use the self-attention module in the progressive feature fusion module to process the first feature set to obtain a second feature set;
[0051] Process the second feature set and the feature extraction set of the I-frame by using the cross-attention module in the progressive feature fusion module to obtain a third feature set;
[0052] Add the third feature set to the feature extraction set of the I-frame, and input the added features into the temporal feature fusion module in the progressive feature fusion module to obtain a set of fused feature vectors.
[0053] In a possible implementation, the loss function of the cross-modal progressive CLIP feature extraction network is expressed as:
[0054]
[0055] where L video represents the video loss; L text represents the text loss.
[0056] In a possible implementation, the contrast module is used to calculate the similarity between the set of fused feature vectors and the set of text feature vectors, and determine the video actions in the video sequence to be recognized according to the similarity, including:
[0057] Calculate the similarity matrix of the corresponding vectors in the set of fused feature vectors and the set of text feature vectors;
[0058] Pass the similarity matrix through the softmax layer for the video actions in the video sequence to be recognized.
[0059] In a second aspect, the present invention provides a compressed video action recognition device based on cross-modal progressive CLIP, including: an input unit, a processing unit, and an output unit;
[0060] The input unit is used to obtain the video sequence to be recognized, convert the video sequence to be recognized into a re-encoded video sequence; and obtain a plurality of text descriptions corresponding to the video sequence to be recognized;
[0061] The processing unit is used to obtain the compressed video action recognition result according to the trained cross-modal progressive CLIP feature extraction network;
[0062] The output unit is used to output the compressed video action recognition result.
[0063] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0064] By adopting a cross-modal progressive CLIP feature extraction network, the present invention can determine the I-frame features and P-frame features of the video sequence to be recognized, namely spatial features and motion features. The self-attention module, cross-attention module, and temporal feature fusion module in the progressive feature fusion module gradually enhance and refine low-level features through multi-level feature extraction, enabling the cross-modal progressive CLIP feature extraction network to focus more on action-related details, ensuring that key action information is fully captured, filtering out irrelevant noise, making better use of multi-modal information, and improving the overall performance of action recognition; by combining the motion vectors and residual features in the compressed video in the cross-modal progressive CLIP feature extraction network, using text descriptions to supplement the semantic relevance of actions, and fusing compressed domain features such as motion vectors and residuals in the compressed video, the ability to understand actions and their contexts is enhanced; effectively solving the problem in the prior art that the hierarchical interactivity between modalities is ignored, resulting in the model being difficult to fully mine the effective information in the compressed video, thereby realizing the enhancement of the network's understanding of the scene, maintaining high efficiency in processing, and significantly improving the accuracy and robustness of action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 FIG. is a flowchart of the steps of a method for recognizing actions in a compressed video based on cross-modal progressive CLIP provided by an embodiment of the present invention;
[0066] Figure 2 FIG. is a processing flowchart of a cross-modal progressive CLIP feature extraction network provided by an embodiment of the present invention;
[0067] Figure 3 FIG. is a processing flowchart of a progressive feature fusion module provided by an embodiment of the present invention;
[0068] Figure 4 FIG. is a schematic diagram of a device for recognizing actions in a compressed video based on cross-modal progressive CLIP provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] The present invention provides a method for recognizing actions in a compressed video based on cross-modal progressive CLIP. Refer to Figure 1 , the method includes the following steps S101 to S103.
[0071] S101. Obtain the video sequence to be recognized and convert the video sequence to be recognized into a re-encoded video sequence;
[0072] Exemplarily, convert the video sequence V to be recognized original into a re-encoded video sequence V using the H.263 video encoding method reencoded . In the re-encoded video sequence V reencoded , the size of a group of consecutive pictures (Group of Pictures, GOP) is 12. Among them, the first frame in a GOP is an I frame, and the subsequent 11 frames are P frames, denoted as G i = [I i , P i,1 , P i,2 , …, P i,11 . The P frame contains information in two modalities of motion vectors and residuals, denoted as P i,j = (M j , R j ). Here, it should be noted that when using different video encoding methods, the number of frames in each GOP obtained is different.
[0073] S102. Obtain multiple text descriptions corresponding to the video sequence to be recognized;
[0074] S103. Input the re-encoded video sequence into the trained cross-modal progressive CLIP feature extraction network to obtain the compressed video action recognition result.
[0075] For the cross-modal progressive CLIP feature extraction network, see Figure 2 , which includes: a data processing module, a visual encoder branch, a motion encoder branch, a detail encoder branch, a text processing branch, a progressive feature fusion module, and a contrast module.
[0076] The data processing module is used to preprocess the re-encoded video sequence to obtain a sampled GOP group.
[0077] Specifically, when the data processing module is used to preprocess the re-encoded video sequence to obtain a sampled GOP group, it includes:
[0078] (1) Sample the re-encoded video sequence to obtain an initial sampled GOP group;
[0079] (2) Perform data augmentation on the initial sampled GOP group to obtain a sampled GOP group.
[0080] Exemplarily, for the I-frames in the initial sampled GOP group, data augmentation is performed through image multi-scale cropping, random horizontal flipping, random color jittering, random grayscaling, Gaussian blurring, random exposure, etc. For the motion vectors and residuals in the P-frames of the initial sampled GOP group, data augmentation is performed through central cropping, random horizontal flipping, and data normalization.
[0081] The visual encoder branch is used to extract I-frame features from the I-frame data of the sampled GOP group based on the M-layer Transformer encoder to obtain an I-frame extraction feature set; where M is greater than 0;
[0082] Specifically, the visual encoder branch is used to extract I-frame features from the I-frame data of the sampled GOP group based on the M-layer Transformer encoder to obtain an I-frame feature set, including:
[0083] (1) Obtain the I-frame images in each GOP in the sampled GOP group to obtain a first I-frame data set;
[0084] (2) Adjust the sizes of the I-frame images in the first I-frame data set to obtain an updated first I-frame data set;
[0085] (3) Normalize each I-frame image in the updated first I-frame data set to obtain a normalized first I-frame data set;
[0086] (4) Map each I-frame image in the normalized first I-frame data set to the feature space to obtain a mapped I-frame feature set;
[0087] (5) Perform position embedding on each I-frame image in the mapped I-frame feature set to obtain a position-embedded I-frame feature set;
[0088] (6) Extract features from each I-frame image in the position-embedded I-frame feature set based on the M-layer Transformer encoder to obtain an I-frame extraction feature set. Here, M in the M-layer Transformer encoder is 12.
[0089] Exemplarily, for the visual encoder branch, given the input compressed video V reencoded , sample T key frames {I1, I2,..., I T} from it, and the size of each frame image I t is H×W×C, where H is the height, W is the width, and C is the number of channels.
[0090] Preprocess each I-frame image in the first I-frame data set and adjust the size of each I-frame image in the first I-frame data set to 224×224:
[0091] I tI' = Resize(I t , (H', W'));
[0092] Where H' represents the target height; W' represents the target width; Resize(·) represents the image resizing operation. Normalize each I-frame image in the updated first I-frame dataset:
[0093] I t ″ = Normalize(I t ′, μ, σ);
[0094] Where μ is the mean value; σ is the standard deviation; Normalize(·) represents the image normalization operation.
[0095] Based on each normalized I-frame image, obtain the first I-frame dataset V', that is, generate the batch representation as V', V' = [I1″, I2″,..., I T ″] ∈ R T×H′×W′×C .
[0096] The visual encoder branch consists of M = 12 Transformer encoder layers, and each layer contains a multi-head self-attention mechanism and a feed-forward neural network.
[0097] The visual encoder branch maps the first I-frame dataset V' to the feature space to obtain the mapped I-frame feature set P I , specifically expressed as:
[0098]
[0099] Where N represents the number of small blocks into which each I-frame is divided, the size of each small block is 16×16, d model represents the embedding dimension, which is taken as 768 here, and PacthEmbed(·) represents the mapping function.
[0100] Add positional embeddings to each I-frame image in the mapped I-frame feature set P I to generate the I-frame feature set P I ' with positional embeddings:
[0101] P I ' = P I + PositionEncoding(P I );
[0102] Where PositionEncoding(·) is the positional embedding function.
[0103] Input the I-frame feature set P I ' with positional embeddings into the transformer network to obtain the I-frame extraction feature set ZI , specifically expressed as:
[0104]
[0105] Among them, TransformerEncoder(·) represents the encoder in the transformer network.
[0106] The motion encoder branch is used to extract motion features from the P-frame data of the sampled GOP group based on the F-layer Transformer encoder, obtaining a motion extraction feature set; where F is greater than or equal to 0 and less than M;
[0107] Specifically, the motion encoder branch is used to extract motion features from the P-frame data of the sampled GOP group based on the F-layer Transformer encoder, obtaining a motion extraction feature set, including:
[0108] (1) Obtain multiple P-frames in each GOP of the sampled GOP group to obtain a P-frame data set;
[0109] (2) Randomly sample each P-frame data in the P-frame data set to obtain a sampled P-frame data set;
[0110] (3) Use the coviar method to extract the motion modality from the sampled P-frame data set to obtain a motion vector set;
[0111] (4) Adjust the magnitude of each motion vector in the motion vector set to obtain a preprocessed motion vector set;
[0112] (5) Perform position embedding on each motion vector in the preprocessed motion vector set respectively to obtain a position-embedded motion vector set;
[0113] (6) Use the F-layer Transformer encoder to extract motion features from each motion vector in the position-embedded motion vector set respectively to obtain a motion extraction feature set; here, N = 2.
[0114] Exemplarily, for the motion encoder branch, obtain the GOP group corresponding to the sampled P-frame data set, randomly sample three P-frames in each GOP, and the P-frame includes a motion vector modality.
[0115] Resize the magnitude of each motion vector MV in the motion vector set (i,j) After resizing, the motion vector MV (i ′ ,j) has a dimension of 56×56×2, and each motion vector MV in the preprocessed motion vector set (i ′ ,j) is expressed as:
[0116] MV (i ′ ,j) = Resize(MV (i,j) , (H′ MV , W M ′ V ));
[0117] Among them, H′ MV represents the target height of the motion vector; W M ′ V represents the target width of the motion vector.
[0118] Represent the pre - processed motion vector set as MVT; that is, generate a batch represented as MVT, where MVT includes: I represents the number of P - frame samplings in the GOP.
[0119] The motion encoder branch consists of 2 layers of Transformer encoder layers, and the detail encoder also consists of 2 layers of Transformer encoder layers.
[0120] Add position embeddings to each pre - processed motion vector in the pre - processed motion vector set MVT to generate each motion vector P after position embedding corresponding to each pre - processed motion vector mv , and according to each motion vector P after position embedding mv , obtain the motion vector set P after position embedding MV ; among them, the calculation method of each motion vector P after position embedding mv is:
[0121] P mv = MVT + PositionEncoding(MVT);
[0122] Among them, PositionEncoding(·) represents the position embedding function;
[0123] Input the motion vector set P after position embedding into the motion encoder branch to obtain the motion extraction feature set Z MV , specifically expressed as: res
[0124]
[0125] Among them, TransformerEncoderMV(·) represents the motion encoder branch; N represents the number of small pieces into which the running vector is divided; d model represents the embedding dimension.
[0126] The detail encoder branch is used to extract residual features from the P-frame data of the sampled GOP group based on the F-layer Transformer encoder, obtaining a residual extraction feature set; here, N = 2.
[0127] Specifically, the detail encoder branch is used to extract residual features from the P-frame data of the sampled GOP group based on the F-layer Transformer encoder, obtaining a residual extraction feature set, including:
[0128] (1) Obtain multiple P-frames in each GOP of the sampled GOP group to get a P-frame data set;
[0129] (2) Randomly sample each P-frame data in the P-frame data set to get a sampled P-frame data set;
[0130] (3) Use the coviar method to extract the residual mode from the sampled P-frame data set to get a residual set;
[0131] (4) Adjust the size of each residual in the residual set to get a preprocessed residual set;
[0132] Perform position embedding on each residual in the preprocessed residual set respectively to get a position-embedded residual set;
[0133] (5) Use the F-layer Transformer encoder to extract residual features from each residual in the position-embedded residual set respectively, obtaining a residual extraction feature set.
[0134] Exemplarily, for the detail encoder branch, obtain the GOP group corresponding to the sampled P-frame data set, randomly sample three P-frames in each GOP, and the P-frame contains the residual mode.
[0135] Resize the size of the residual RES (i,j) in the residual set to get a preprocessed residual set, and the residual RES ( ′ i,j) in the processed residual set has a dimension of 224×224×3;
[0136] RES ( ′ i,j) = Resize(RES (i,j) ,(H R ′ ES ,W R ′ ES ));
[0137] where, H R ′ ES is the target height of the residual; W R ′ ES is the target width of the residual.
[0138] Represent the preprocessed residual set as REST, that is, generate batch REST, where REST includes I represents the number of P-frame samplings in the GOP.
[0139] The detail encoder branch consists of 2 Transformer encoder layers.
[0140] Add positional embeddings to the preprocessed residual set REST to generate the residual set P after positional embedding RES , where each residual P after positional embedding res is calculated as:
[0141] P res = REST + PositionEncoding(REST);
[0142] Input the residual set P after positional embedding RES into the detail encoder branch respectively, and the residual extraction feature set Z mv ;
[0143]
[0144] where TransformerEncoderRES(·) is the detail encoder branch.
[0145] The text processing branch is used to extract text features from multiple text descriptions respectively based on a single-layer Transformer encoder to obtain a text feature vector set;
[0146] Specifically, the text processing branch is used to extract text features from multiple text descriptions respectively based on a single-layer Transformer encoder to obtain a text feature vector set, including:
[0147] (1) Tokenize each text description and convert the tokenized text description into a text embedding vector;
[0148] (2) Perform positional embedding on the text embedding vector to obtain the text embedding vector after positional embedding;
[0149] (3) Input the text embedding vector after positional embedding into a single-layer Transformer encoder for text feature extraction to obtain a text feature vector set.
[0150] Exemplarily, for a text description, first, text data preprocessing is required. For text description D, first tokenize and convert it into a text embedding vector E D :
[0151] D = {w1, w2,..., wL};
[0152]
[0153] Among them, L represents the sentence length, w represents the word, represents the dimension of the model, which is the same as the d model dimension.
[0154] The text processing branch embeds the input text into the vector E D adds position embedding to generate the text embedding vector E' after position embedding D :
[0155] E' D = E D + PositionEncoding(E D );
[0156] Among them, PositionEncoding(·) represents the position embedding function;
[0157] The text embedding vector E' after position embedding D is input into the transformer network to obtain the text feature vector Z corresponding to a single text description d ;
[0158]
[0159] Among them, TransformerEncoder(·) represents the encoder in the transformer network.
[0160] According to the text feature vector Z corresponding to a single text description d , the text feature vector set Z D is obtained.
[0161] See Figure 3 , the progressive feature fusion module fuses the I-frame extraction feature set, the motion extraction feature set, and the residual extraction feature set to obtain the fused feature vector set;
[0162] Specifically, the progressive feature fusion module fuses the I-frame extraction feature set, the motion extraction feature set, and the residual extraction feature set to obtain the fused feature vector set, including:
[0163] (1) Add the motion vectors and residuals belonging to the same GOP in the motion extraction feature set and the residual extraction feature set value by value to obtain the first feature set;
[0164] (2) Use the self-attention module in the progressive feature fusion module to process the first feature set to obtain the second feature set;
[0165] (3) Process the second feature set and the I-frame extracted feature set using the cross-attention module in the progressive feature fusion module to obtain a third feature set;
[0166] (4) Add the third feature set to the I-frame feature set, and input the added features into the temporal feature fusion module in the progressive feature fusion module to obtain a fused feature vector set.
[0167] Exemplarily, take the feature Z I , Z res , Z mv and fuse them through the progressive feature fusion module to obtain the fused feature vector Z F , and this progressive feature fusion module consists of a self-attention module, a cross-attention module, and a temporal feature fusion module.
[0168] First, add the motion extraction feature set Z res and the residual extraction feature set Z mv value by value to obtain the first feature set Z compress :
[0169] Z compress = Z res + Z mv ;
[0170] Pass the first feature set Z compress through the self-attention module to obtain the second feature set Z c ' ompress , and then input the second feature set Z c ' ompress and the I-frame extracted feature set Z I into the cross-attention module to obtain the third feature set Z c '' ompress :
[0171] Z c ' ompress = SelfAttention(Z compress );
[0172] Z c '' ompress = CrossAttention(Z c ' ompress , Z I );
[0173] Among them, CrossAttention(·) and SelfAttention(·) are the cross-attention mechanism and self-attention mechanism functions respectively.
[0174] Input the third feature set Z c ''ompress is added to the I-frame feature set Z I and then sent to the temporal feature fusion module to obtain the fused feature vector set Z F :
[0175] Z F = TemporalFusion(Z c ″ ompress + Z I );
[0176] where TemporalFusion(·) represents the temporal feature fusion module.
[0177] The comparison module is used to calculate the similarity between the fused feature vector set and the text feature vector set, and determine the video actions in the video sequence to be recognized according to the similarity.
[0178] Specifically, the comparison module is used to calculate the similarity between the fused feature vector set and the text feature vector set, and determine the video actions in the video sequence to be recognized, including:
[0179] (1) Calculate the similarity matrix of the corresponding vectors in the fused feature vector set and the text feature vector set;
[0180] (2) Pass the similarity matrix through the softmax layer for the video actions in the video sequence to be recognized.
[0181] Exemplarily, the similarity calculation formula is expressed as:
[0182]
[0183] where v i represents the i-th fused feature vector in the fused feature vector set, and t j represents the text feature vector corresponding to the j-th text description; v i T represents the transpose of the i-th fused feature vector in the fused feature vector set.
[0184] When training the cross-modal progressive CLIP feature extraction network, the loss function of the cross-modal progressive CLIP feature extraction network is expressed as:
[0185]
[0186] where L video represents the video loss; L text represents the text loss.
[0187]
[0188] It should be noted that although the Transformer is used as the feature extraction network in this embodiment, other feature extraction networks, such as CNN, Resnet, and deep learning models of convolutional neural networks such as U-Net, can also be used as alternatives. In this embodiment, the MLP is selected as the projection head, but other models, such as support vector machines, can also be used for this purpose.
[0189] Most of the existing methods for compressed video action recognition fail to fully address the significant decline in action recognition accuracy caused by the loss or simplification of background information during the video compression process. In contrast, the present invention combines compression domain features such as motion vectors and residuals in compressed videos and fuses text descriptions related to actions, thereby enhancing the ability to understand actions and their contexts. By leveraging the information contained in the text to enhance semantic relevance, this method not only strengthens the understanding of the scene background but also significantly improves the accuracy and robustness of action recognition while maintaining efficient processing.
[0190] The existing methods for compressed video action recognition usually have a relatively simple way of fusing compressed video modalities, mainly by directly concatenating or weighted summing modality features such as I-frames, motion vectors, and residuals. This approach ignores the hierarchical interactions and semantic alignments between modalities, making it difficult for the model to fully exploit the effective information in compressed videos. To address this issue, a progressive feature fusion module is introduced. Specifically, this module can design the feature extraction process according to the unique characteristics of compressed videos, enabling the model to focus more on action-related details and filter out irrelevant noise. Therefore, compared with existing methods, the method provided by the present invention is more refined and efficient in feature fusion, can better utilize multi-modal information, and improve the overall performance of action recognition.
[0191] To improve the effect of compressed video action recognition based on cross-modal progressive CLIP, this embodiment provides a compressed video action recognition device based on cross-modal progressive CLIP, see Figure 4 , including: an input unit, a processing unit, and an output unit;
[0192] The input unit is used to obtain the video sequence to be recognized, convert the video sequence to be recognized into a re-encoded video sequence; and obtain multiple text descriptions corresponding to the video sequence to be recognized
[0193] The processing unit is used to obtain the compressed video action recognition result according to the trained cross-modal progressive CLIP feature extraction network;
[0194] The output unit is used to output the compressed video action recognition result.
[0195] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. All or part of the present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0196] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the present invention; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present invention.
Claims
1. A method for compressed video action recognition based on cross-modal progressive CLIP, characterized in that: include: Acquire a video sequence to be identified, and convert the video sequence to be identified into a re-encoded video sequence; Acquire multiple text descriptions corresponding to the video sequence to be identified; The re-encoded video sequence is input into a trained cross-modal progressive CLIP feature extraction network to obtain a compressed video action recognition result; wherein the cross-modal progressive CLIP feature extraction network comprises: Data processing module, visual encoder branch, motion encoder branch, detail encoder branch, text processing branch, progressive feature fusion module and contrast module; The data processing module is used to pre-process the re-encoded video sequence to obtain a sampled GOP group; The visual encoder branch is used to extract I frame features from the I frame data of the sampled GOP group based on an M-layer Transformer encoder to obtain an I frame extraction feature set; wherein M is greater than 0; The motion encoder branch is used to extract motion features from the P frame data of the sampled GOP group based on an F-layer Transformer encoder to obtain a motion extraction feature set; wherein F is greater than or equal to 0 and less than M; The detail encoder branch is used to perform residual feature extraction on the P frame data of the sampled GOP group based on the F-layer Transformer encoder to obtain a residual extraction feature set; The text processing branch is used to extract text features from multiple text descriptions based on a single-layer Transformer encoder to obtain a text feature vector set; The progressive feature fusion module fuses the I frame extraction feature set, the motion extraction feature set and the residual extraction feature set to obtain a fused feature vector set; The comparison module is used to calculate the similarity between the fused feature vector set and the text feature vector set, and determine the video action in the to-be-identified video sequence according to the similarity.
2. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The data processing module is used to pre-process the re-encoded video sequence to obtain a sampled GOP group, including: Sampling the re-encoded video sequence to obtain an initial sampling GOP group; Data enhancement is performed on the initial sampled GOP group to obtain a sampled GOP group.
3. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The visual encoder branch is used to extract I frame features from the I frame data of the sampled GOP group based on an M-layer Transformer encoder to obtain an I frame extraction feature set, including: Acquire an I frame image in each GOP in the sampled GOP group to obtain a first I frame data set; Adjusting the size of each I-frame image in the first I-frame data set to obtain an updated first I-frame data set; Normalizing each I frame image in the updated first I frame data set to obtain a normalized first I frame data set; Mapping each I-frame image in the normalized first I-frame data set to the feature space to obtain a mapped I-frame feature set; Perform position embedding on each I-frame image in the mapped I-frame feature set to obtain the I-frame feature set after position embedding; Based on the M-layer Transformer encoder, feature extraction is performed on each I-frame image in the I-frame feature set after position embedding to obtain the I-frame extraction feature set.
4. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The motion encoder branch is used to extract motion features from the P frame data of the sampled GOP group based on the F-layer Transformer encoder to obtain a motion extraction feature set, including: Acquire multiple P frames in each GOP in the sampled GOP group to obtain a P frame data set; Randomly sampling each P frame data in the P frame data set to obtain a sampled P frame data set; The coviar method is used to extract motion modes from the sampled P frame data set to obtain a motion vector set; Adjusting the size of each motion vector in the motion vector set to obtain a preprocessed motion vector set; Perform position embedding on each motion vector in the preprocessed motion vector set to obtain a motion vector set after position embedding; The F-layer Transformer encoder is used to extract motion features from each motion vector in the motion vector set after position embedding to obtain a motion extraction feature set.
5. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The detail encoder branch is used to perform residual feature extraction on the P frame data of the sampled GOP group based on the F-layer Transformer encoder to obtain a residual extraction feature set, including: Acquire multiple P frames in each GOP in the sampled GOP group to obtain a P frame data set; Randomly sampling each P frame data in the P frame data set to obtain a sampled P frame data set; The coviar method is used to extract the residual mode of the sampled P frame data set to obtain the residual set; Adjusting the size of each residual in the residual set to obtain a preprocessed residual set; Perform position embedding on each residual in the preprocessed residual set respectively to obtain a residual set after position embedding; The F-layer Transformer encoder is used to extract residual features from each residual in the residual set after position embedding to obtain a residual extraction feature set.
6. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The text processing branch is used to extract text features from multiple text descriptions based on a single-layer Transformer encoder to obtain a text feature vector set, including: Segment each text description and convert the segmented text description into a text embedding vector; Performing position embedding on the text embedding vector to obtain a text embedding vector after position embedding; The text embedding vector after position embedding is input into a single-layer Transformer encoder for text feature extraction to obtain a text feature vector set.
7. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The progressive feature fusion module fuses the I frame extraction feature set, the motion extraction feature set and the residual extraction feature set to obtain a fused feature vector set, including: Adding the motion vectors and residuals belonging to the same GOP in the motion extraction feature set and the residual extraction feature set value by value to obtain a first feature set; Processing the first feature set using a self-attention module in a progressive feature fusion module to obtain a second feature set; Processing the second feature set and the I-frame extraction feature set using a cross attention module in a progressive feature fusion module to obtain a third feature set; The third feature set is added to the I frame extraction feature set, and the added features are input into a time feature fusion module in a progressive feature fusion module to obtain a fused feature vector set.
8. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The loss function of the cross-modal progressive CLIP feature extraction network is expressed as: Among them, L video Indicates video loss; L text Indicates text loss.
9. The method for compressed video action recognition based on cross-modal progressive CLIP according to claim 1, characterized in that: The comparison module is used to calculate the similarity between the fusion feature vector set and the text feature vector set, and determine the video action in the to-be-identified video sequence according to the similarity, including: Calculating a similarity matrix between the fused feature vector set and corresponding vectors in the text feature vector set; The similarity matrix is passed through a softmax layer to identify video actions in the video sequence.
10. A compressed video action recognition device based on cross-modal progressive CLIP, characterized in that: include: Input unit, processing unit and output unit; The input unit is used to obtain a video sequence to be identified, convert the video sequence to be identified into a re-encoded video sequence; and obtain a plurality of text descriptions corresponding to the video sequence to be identified; The processing unit is used to obtain a compressed video action recognition result based on the trained cross-modal progressive CLIP feature extraction network; The output unit is used to output the compressed video action recognition result.
Citation Information
Cited By
Lightweight Transform video action recognition method based on video and text fusion
CN121353996A
A lightweight transform video action recognition method based on video and text fusion
CN121353996B