Video time sequence action detection method based on Mamba2 and bidirectional feature pyramid

By introducing the Mamba2 model and a bidirectional feature pyramid network in video timing action detection, the problem of high computational complexity is solved, and efficient identification of action examples and timely moments in long videos is achieved.

CN119942641APending Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510009896.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has high computational complexity in video timing action detection, making it difficult to effectively identify action examples and moments in long videos.

Method used

The video timing action detection method based on the Mamba2 model and the bidirectional feature pyramid network is adopted. The semi-separable matrix block decomposition algorithm of the Mamba2 model is reduced in computational complexity, and the bidirectional feature pyramid network is fused to improve the action detection ability.

Benefits of technology

This significantly reduces the computational complexity and improves the model's ability to detect actions on different time scales, allowing more efficient identification of category labels, start moments and end moments of action instances in the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942641A_ABST
    Figure CN119942641A_ABST
Patent Text Reader

Abstract

The invention provides a video time sequence action detection method based on Mamba2 and a bidirectional feature pyramid. The method comprises the following steps: (1) extracting features from an input video sequence through a pre-trained feature extractor; (2) constructing a time sequence action detection model based on Mamba2 and a bidirectional feature pyramid, stacking L modules based on Mamba2 to form a bidirectional feature pyramid network, encoding input video features, and extracting key information; and (3) sending the output multi-scale features into a regression head and a classification head, and decoding to obtain a detection result, namely, inputting a category label, a starting moment and an ending moment of an action instance in the video. According to the method, the Mamba2 model and the bidirectional feature pyramid network are used, so that the capability of detecting actions of different time scales is effectively improved, and the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a video temporal action detection method based on Mamba2 and a bidirectional feature pyramid, and belongs to the field of video understanding. Background Art

[0002] Video temporal action detection is an important research direction in the field of video understanding. It involves locating the start and end time of a specific action in an uncut long video and identifying the category label of the action. This technology has a wide range of applications in video content analysis, video editing, intelligent monitoring, virtual reality and other fields.

[0003] Temporal action detection, also known as temporal action localization, is an important area of ​​video understanding. Action recognition can be seen as a pure classification problem, in which the videos to be recognized are basically edited, that is, each video contains a clear action, the video duration is short, and there is a unique action category. In the field of temporal action detection, the video is usually not edited, the video duration is long, the action usually only occurs in a short period of time in the video, and the video may contain multiple actions or no action, that is, the background class. Temporal action detection not only predicts what actions are contained in the video, but also predicts the start and end time of the action. Compared with action recognition, temporal action detection is closer to real scenes. Traditional methods are mainly based on manually designed feature extraction and machine learning algorithms, but these methods have limited accuracy and robustness in complex scenes. In recent years, the rise of deep learning has brought new breakthroughs in temporal action detection.

[0004] With the rise of deep learning, more and more neural network-based methods have been proposed and achieved good results. These methods are divided into multiple modules, including two-stream convolutional neural networks (CNNs), recurrent neural networks (RNNs), 3DCNNs, and Transformer-based methods. In recent years, the success of Transformer in natural language processing tasks has attracted great attention in other fields, including speech, images, and videos, relying entirely on self-attention without using sequence-aligned RNNs or convolutions. The application of Transformer-based action recognition is relatively new, but the number of studies on this topic is increasing and has achieved good results.

[0005] However, the action recognition algorithm based on Transformer has the defect of high computational complexity. Recently, the Mamba2 model has attracted widespread attention in the field of deep learning due to its characteristics of being much faster than Transformer and more accurate. The Mamba2 model uses the "duality" of the SSM state space and the Transformer model to convert the problem of high temporal complexity of Transformer when processing long sequences into the SSM state space model for processing, which significantly reduces the temporal complexity. In the field of video temporal action detection, it is necessary to identify all action instances contained in the video and the start and end times of each action instance. For temporal action detection of long videos, a very large amount of calculation is required, so how to reduce the temporal complexity is an important issue in the field of video temporal action detection. In addition, this paper adopts a special bidirectional feature pyramid structure to fuse features of different scales to improve the image effect. The bidirectional feature map pyramid network mainly solves the multi-scale problem in object detection. Without basically increasing the computational complexity of the original model, the detection performance of small objects is greatly improved by simply changing the network connection. In the field of video temporal action detection, the ability of the model to detect action instances of different time scales can be improved by using a bidirectional feature pyramid network. Summary of the invention

[0006] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a video temporal action detection method based on Mamba2 and a bidirectional feature pyramid. By introducing the Mamba2 model and the bidirectional feature pyramid network, the computational complexity is significantly reduced and the ability of the model to detect action instances of different time scales is improved. In order to achieve the above purpose, the present invention is implemented by the following technical solutions:

[0007] A temporal action detection method based on Mamba2 and bidirectional feature pyramid comprises the following steps:

[0008] S1. Extract features from the input video sequence through a pre-trained feature extractor.

[0009] S2. Build a temporal action detection model based on Mamba2 and bidirectional feature pyramid. Stack L Mamba2 modules to form a bidirectional feature pyramid network, which is used to encode the input video features and extract key information.

[0010] S3. Send the multi-scale features output by the model constructed in step S2 to the regression head and the classification head to obtain the detection results, that is, the category label, start time and end time of the action instance in the input video.

[0011] Furthermore, the feature extractor in step S1 uses an I3D network pre-trained on the Kinetics dataset. The video clip is sampled every 16 frames.

[0012] Furthermore, the step S1 extracts features from the input video sequence through a pre-trained feature extractor, and the specific steps are as follows:

[0013] S11, extracting RGB features of the input video and extracting video optical flow features through the TV-L1 algorithm.

[0014] S12. Input the RGB features and optical flow features into the 3D convolutional neural network to extract action features.

[0015] S13, average the dual-stream action features (RGB features and optical flow features) and use the Softmax formula to fuse the dual-stream features to obtain a video feature sequence. The Softmax formula is as follows:

[0016]

[0017] Among them, z i It means that the video is sampled once every 16 frames to extract the feature vector, where the feature vector extracted for the i-th time, j means that the video is sampled j times in total, and exp(.) represents the power exponent of e.

[0018] Furthermore, the steps of constructing the temporal action detection model in step S2 are as follows:

[0019] S21. Input the extracted features into a shallow convolutional network for projection, and transform them into a D-dimensional embedding space through a convolutional network using the ReLU activation function. D is usually 1024, and Z0 = {E(x1), E(x2), ..., E(x T )},E(x)∈R D Represents the projection function of a shallow convolutional network.

[0020] S22, after a normalization layer, uses the Mamba2 module to encode the projected features. The Mamba2 module uses the SSD algorithm based on semi-separable matrix block decomposition to reduce time complexity. At the beginning of each block, the A, B, and C matrices are generated in parallel, and then the SSD algorithm is used After quantization, the M matrix is ​​obtained, and the final output is F=MZ.

[0021] S23. Adding residual connection after the Mamba2 module can avoid information loss and solve the problem of gradient disappearance.

[0022] S24, after downsampling, the output result is input into the next layer Mamba2 module, and the output of all Mamba2 modules is used to construct a bidirectional feature pyramid network. After multiple bidirectional feature fusions, the feature matrix Z = {Z1, Z2, ..., Z L}.

[0023] Furthermore, the bidirectional feature pyramid network allows information between features of different scales to flow more efficiently by establishing bidirectional connections between the top-down and bottom-up paths. The bidirectional connection allows the integration of information from different levels of the feature pyramid in both directions, effectively capturing multi-scale features.

[0024] Furthermore, in step S3, the multi-scale features are fed into the regression head and the classification head. Each position in the multi-scale features represents a moment and is regarded as an action candidate. The convolutional network is used as a decoder to classify these action candidates and regress the offset between the start frame and the end frame of the action. The classification head predicts the probability of the action category, and the regression head predicts the distance between the action start moment and the end moment, thereby obtaining the result of temporal action detection.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] 1. The present invention introduces the Mamba2 model and utilizes the SSD algorithm based on semi-separable matrix block decomposition of the Mamba2 model to significantly reduce the computational complexity.

[0027] 2. The present invention constructs a bidirectional feature pyramid network based on the Mamba2 model. The bidirectional feature pyramid fuses features through two channels, one top-down and one bottom-up, which significantly improves the model's ability to detect actions at different time scales. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments and technical solutions of the present invention, the following briefly introduces the drawings of the present invention.

[0029] Figure 1 It is a flow chart of the present invention.

[0030] Figure 2 This is a network structure diagram based on the Mamba2 module.

[0031] Figure 3 This is a bidirectional feature pyramid network structure diagram built based on the Mamba2 module. DETAILED DESCRIPTION

[0032] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0033] As shown in the figure, the present invention provides a temporal action detection method based on Mamba2 and a bidirectional feature pyramid, comprising the following steps:

[0034] 1. Extract features from the input video sequence through a pre-trained feature extractor

[0035] First, the system extracts RGB features from the input video and extracts optical flow features through the TV-L1 algorithm. Then, the obtained dual-stream features are input into the 3D convolutional neural network to extract action features. After averaging the dual-stream action features, the Softmax formula is used to fuse the dual-stream features to obtain the video feature sequence X = {x1, x2, ..., x T}.

[0036] 2. The extracted features are input into a shallow convolutional network for projection, and then transformed into a D-dimensional embedding space through a convolutional network using the ReLU activation function. D is usually 1024, and Z0 = {E(x1), E(x2), ..., E(x T )},E(x)∈R D Represents the projection function of a shallow convolutional network.

[0037] 3. After the normalization layer, the projected features are encoded using the Mamba2 module. At the beginning of each block, the A, B, and C matrices are generated in parallel, and then the SSD algorithm is used After quantization, the M matrix is ​​obtained, and a residual connection is added after the Mamba2 module to avoid information loss and solve the problem of gradient disappearance. The final output is F = MZ.

[0038] 4. The output result is input into the next layer Mamba2 module after downsampling, and the output of all Mamba2 modules is used to construct a bidirectional feature pyramid network. After multiple bidirectional feature fusions, the feature matrix Z = {Z1, Z2, ..., Z L The bidirectional feature pyramid network uses downsampling to connect L Mamba2 modules, allowing information between features of different scales to flow more efficiently by establishing bidirectional connections between the top-down and bottom-up paths. The bidirectional connection allows the integration of information from different levels of the feature pyramid in both directions, effectively capturing multi-scale features.

[0039] 5. The multi-scale features output by the bidirectional feature pyramid are fed into the regression head and the classification head. Each position in the multi-scale features represents a moment and is regarded as an action candidate. The convolutional network is used as a decoder to classify these action candidates and regress the offset between the start frame and the end frame of the action to obtain the detection result Y = {y1, y2, ..., y T},in That is, the start time, end time and category label of the action instance in the input video.

[0040] The technical means disclosed in the scheme of the present invention are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical schemes composed of any combination of the above-mentioned technical features. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also regarded as the protection scope of the present invention.

Claims

1. A video temporal action detection method based on Mamba2 and bidirectional feature pyramid, characterized by: The method comprises the following steps: S1, extract features from the input video sequence through a pre-trained feature extractor; S2. Build a temporal action detection model based on Mamba2 and bidirectional feature pyramid. Stack L Mamba2 modules to form a bidirectional feature pyramid network, which is used to encode the input video features and extract key information. S3. Send the multi-scale features output by the temporal action detection model constructed in step S2 to the regression head and the classification head, and decode to obtain the detection results, that is, the category label, start time and end time of the action instance in the input video.

2. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 1 is characterized in that: The step S1 specifically includes the following steps: S11, extracting RGB features of the input video and extracting video optical flow features through the TV-L1 algorithm; S12, inputting RGB features and optical flow features into 3D convolutional neural network to extract action features; S13, average the dual-stream action features and use the Softmax formula to fuse the dual-stream features to obtain a video feature sequence; the Softmax formula is as follows: Among them, z i It means that the video is sampled once every 16 frames to extract the feature vector, where the feature vector extracted for the i-th time, j means that the video is sampled j times in total, and exp(.) represents the power exponent of e.

3. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 1, characterized in that: The feature extractor in step S1 uses an I3D network pre-trained on the Kinetics dataset, and the video clip is sampled every 16 frames.

4. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 1, characterized in that: The step S2 specifically includes the following steps: S21, input the extracted features into a shallow convolutional network for projection and transform to a D-dimensional embedding space, where D is 1024; S22, use the Mamba2 module to encode the projected features; at the beginning of each block, generate the A, B, and C matrices in parallel, then use the SSD algorithm and quantize the output features; S23, the output result of the Mamba2 module is input into the next layer of the Mamba2 module after downsampling, and the output of all the Mamba2 modules is used to construct a bidirectional feature pyramid network. After multiple bidirectional feature fusions, a multi-scale feature matrix is ​​output.

5. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 4 is characterized in that: The step S21 inputs the extracted features into a shallow convolutional network for projection. This step is implemented by a convolutional network using a ReLU activation function, so that the model can better combine the local context of the temporal features and stabilize the training of the Mamba2 model.

6. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 4, characterized in that: The step S22 uses the Mamba2 module to encode the projected features, and adds a residual connection after the Mamba2 module, which can avoid information loss and solve the problem of gradient disappearance.

7. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 4, characterized in that: The bidirectional feature pyramid network formed in step S23 connects L Mamba2 modules to generate an L-layer bidirectional feature pyramid network, and allows information between features of different scales to flow more effectively by establishing bidirectional connections between top-down and bottom-up paths; the bidirectional connection allows information from different levels of the feature pyramid to be integrated in two directions, effectively capturing multi-scale features.

8. The video temporal action detection method based on Mamba2 and bidirectional feature pyramid according to claim 1, characterized in that: The step S3 feeds the multi-scale features into the regression head and the classification head. Each position in the multi-scale features represents a moment and is regarded as an action candidate. The convolutional network is used as a decoder to classify these action candidates and regress the offset between the start frame and the end frame of the action. The classification head predicts the probability of the action category, and the regression head predicts the distance between the action start moment and the end moment, thereby obtaining the result of temporal action detection.

Citation Information

Cited By

  • Electric power operation progress identification method, system, equipment and medium

    CN121121591A

  • Time sequence action detection method based on frequency domain and time domain information interaction

    CN121259702A

  • Signal modulation identification method and system based on feature pyramid and Mama hybrid model, computer equipment and readable storage medium

    CN121547327A

  • A signal modulation recognition method, system, computer device and readable storage medium based on a feature pyramid and Mamba hybrid model

    CN121547327B