Video time sequence action positioning method based on double-helix convolution attention pyramid
By adopting the double helix convolutional attention pyramid method in video timing action positioning technology, the problems of different action durations and low feature distinction are solved, and higher positioning accuracy and applicability are achieved.
Patent Information
- Application Number
- CN202411991666.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The existing video timing action positioning technology is difficult to ensure positioning accuracy when processing videos with varying durations of action, and the Transformer's self-attention mechanism leads to low feature distinction, affecting the accuracy of the action boundary.
Using a method based on the double helix convolution attention pyramid, time-dimensional downsampling and channel dimension upsampling are performed through the spatiotemporal feature optimizer STFO. Combined with the multi-scale double helix attention convolution module Ds-MAC and the double helix feature pyramid network Ds-FPN, the correlation between feature sequences is established and integrated into the output of the attention mechanism.
It improves the accuracy of feature distinction and action boundaries, effectively handles positioning accuracy within different duration ranges, and is suitable for various application scenarios such as intelligent monitoring and video editing.
Smart Images

Figure CN120014225A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a video temporal action localization method based on a double spiral convolution attention pyramid. Background Art
[0002] The goal of the temporal action localization task is to accurately identify the specific moments when an action starts and ends from an unedited video sequence. This technology is the basis of video understanding and is widely used in video editing, video summarization, video recommendation, intelligent monitoring, and abnormal behavior detection. Since the action in the video can last for an arbitrary length of time, and the action instance can appear at any time point in the video, and its duration is also different, it increases the difficulty of locating the start and end moments of the action in the video.
[0003] In recent years, the outstanding performance of Transformer in the field of computer vision has been widely applied to the task of video temporal action localization, and has achieved remarkable results. However, video features are highly similar between different clips, and this similarity is further enhanced by the self-attention mechanism of Transformer, resulting in low differentiation of learned features, which in turn affects the accuracy of action boundaries.
[0004] In addition, the duration of actual action instances in videos usually ranges from a few seconds to several minutes, so it is crucial to ensure that the model can effectively handle these changes in time scale, which is very important for improving the localization accuracy in different duration ranges. A mainstream approach is to construct a feature pyramid network (FPN) to extract video clips of the same size from all pyramid levels through temporal downsampling. However, these methods ignore the problem of detailed feature loss in the process of mapping features from the top layer to the bottom layer, resulting in inaccurate boundary prediction. In summary, there are still many shortcomings in the current technical field. Summary of the invention
[0005] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a video temporal action localization method based on double spiral convolution attention pyramid.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A temporal action localization method based on a double spiral convolution attention pyramid comprises the following steps:
[0008] S1, video features are first extracted by the pre-trained feature extractor, and then the spatiotemporal feature optimizer STFO downsamples the preliminary features in the time dimension and upsamples them in the channel dimension to enhance the feature representation;
[0009] S2, construct a multi-scale double helix attention convolution module Ds-MAC, which first uses the multi-head dynamic mask attention convolution module MDMAC to fine-grainedly encode the features enhanced in step S1, and then uses the dynamic local feature modeling module DLFM to coarse-grainedly encode the features, establish the association between feature sequences, and integrate them into the output of the attention mechanism through adaptive learning;
[0010] S3. Construct a double-helix feature pyramid network Ds-FPN. Each layer of the pyramid performs horizontal and vertical iterative alternating fusion on the features encoded in step S2 to obtain a multi-scale video representation, and sends the representation to the classification head and regression head respectively to obtain the final positioning result.
[0011] Preferably, the process of constructing the spatiotemporal feature optimizer STFO described in step S1 includes:
[0012] S11. Construct a vertical feature mapping layer to reduce the channel dimension of the feature. Its structure is as follows:
[0013] The features are encoded into the vertical feature projection layer before downsampling the feature time dimension and upscaling the channel dimension. This layer consists of a one-dimensional time convolution layer, a layer normalization (Layer Normalization), and a ReLU activation layer. The one-dimensional time convolution is combined with a padding mask. The padding is to expand short sequences to the same length for batch processing. The mask is used to indicate which parts are valid data and which are padding data, so as to avoid the model being disturbed by the padding part when calculating the loss so that it can effectively process variable-length sequences and capture features within sequences of different lengths. In addition, it is necessary to add normalization layers and activation layers to accelerate network convergence and enhance the representability of the model. The calculation formula is as follows:
[0014] Ω(x)=ReLU(LN(Mask1dConv(x)))
[0015] Where x∈R T×C represents the video features extracted by the pre-trained feature extraction network in step S1, T is the length of the time series, and C is the channel dimension corresponding to the feature. ReLU, LN, and Mask 1dConv represent the ReLU activation layer, layer normalization, and one-dimensional temporal convolution with padding mask, respectively. It represents the features after mapping by the vertical feature mapping layer, and C* is the channel dimension corresponding to the features after mapping.
[0016] S12, construct a spatiotemporal feature optimizer STFO that expands the channel dimension and reduces the time series dimension, and its structure is as follows:
[0017] In order to enhance features by exploiting the relationship between different time dimensions and channel dimensions, the present invention designs a spatiotemporal feature optimizer STFO to enhance the channel dimension of temporal features while reducing their temporal dimension at each level of the lateral feature pyramid. These features are first encoded into the spatiotemporal feature optimizer STFO, which consists of a normalization layer and a one-dimensional convolution layer with a padding mask. By downsampling the input feature x in the time dimension to reduce the temporal resolution, the model can focus on a longer time range. At the same time, convolution kernels are used to expand the channel dimension of the features, which not only extracts richer feature representations, but also provides more information for the model to learn features in subsequent steps, improving the model's perception ability. After the spatiotemporal feature optimizer STFO, features with different time and channel dimensions are output. The enhanced features are represented by I(x)∈R T′×C′ The calculation formula is as follows
[0018] I(x)=LN(Mask1dConv(LN(Mask1dConv(Ω(x))))
[0019] Among them, LN and Mask1dConv represent layer normalization and one-dimensional temporal convolution with padding mask, respectively.
[0020] Preferably, the process of constructing the multi-scale double helix attention convolution module Ds-MAC in step S2 includes:
[0021] S21. Construct the multi-head dynamic mask attention convolution module MDMAC in the multi-scale double spiral attention convolution module Ds-MAC, whose structure is as follows:
[0022] The multi-head dynamic masked attention convolution module MDMAC is designed to help features capture long-distance dependencies and encode relative position information before calculating the query (Q), key (K), and value (V) through convolution and layer normalization, replacing the position encoding module in the standard multi-head attention mechanism. The output features of the self-attention layer are then connected to the input features through the residual, allowing the network to learn incremental changes relative to the input features and focus on learning important features rather than just the overall feature map. This makes the network easier to expand in depth, thereby improving during gradient descent, providing a more stable training and optimization process for the model.
[0023] S22. Construct the dynamic local feature modeling module DLFM in the multi-scale double helix attention convolution module Ds-MAC, whose structure is as follows:
[0024] The dynamic local feature modeling module DLFM consists of a linear layer (Linear Layer) with unchanged feature channel dimension, a one-dimensional temporal convolution layer (Conv1d), a GELU activation layer, and a Dropout regularization layer. This module combines the different attention weights assigned by the self-attention mechanism to further integrate feature information and learn feature representations for specific tasks. By adding a linear layer after the self-attention layer, the weighted features can be linearly combined to further integrate feature information. The output features of the dynamic local feature modeler are then residually connected with the output features of the multi-head dynamic mask attention convolution module MDMAC. It can be expressed as the following formula:
[0025]
[0026] in, represents the residual connection, x out is the output feature of the multi-head dynamic mask attention convolution module MDMAC, E out_xi Represents the output result of the multi-scale double spiral attention convolution module corresponding to the i-th layer of the horizontal pyramid.
[0027] Preferably, the process of constructing the double helix feature pyramid network Ds-FPN in step S3 includes:
[0028] S31. Construct a horizontal pyramid feature fusion module PFEM, map the features of each layer of the pyramid according to the linear layer, and ensure the consistency of the feature dimension when iterating to the adjacent layers of the vertical feature pyramid. The specific process is as follows:
[0029] Due to the double-helix nature of horizontal feature fusion and vertical feature iteration in the pyramid structure, the feature dimensions between adjacent layers of the vertical pyramid must remain consistent during temporal downsampling. Therefore, in order to ensure the consistent temporal dimension of each layer of the features on the horizontal pyramid, nearest neighbor upsampling is applied to the temporal dimension. A one-dimensional convolutional layer with a kernel size of 1 is then used to learn the weight relationship between features, and then the features obtained from the shallow layer are fused with the deep features of the horizontal pyramid through linear mapping to better capture the sequence context information. At the same time, the vertical features are iterated on each pyramid level, that is, the feature representation is gradually enhanced through 7 rounds of iterative processing to ensure that the feature information at different scales is fully integrated. The specific formula is as follows:
[0030]
[0031] Among them, Linear and Conv1d represent one-dimensional temporal convolution and linear layers, UpSample represents the nearest neighbor upsampling, represents the sum operation, and Concat represents the connection along the feature channel dimension. F = R T×C′ Represents the fused features.
[0032] S32. The output features from the horizontal pyramid feature fusion module PFFM are residually connected to the input features before entering the multi-scale double spiral attention convolution module Ds-MAC, and these features are fed into the feedforward network FFN. The network consists of a group normalization layer, a one-dimensional temporal convolution layer, a GELU activation layer, and a Dropout regularization layer. The features of the feedforward layer are then passed to the adjacent layers of the vertical feature pyramid. Finally, the output features from each layer of the vertical feature pyramid are used as input to the classification and regression heads of the decoder. Output feature Ψ∈R T×C′ It is expressed as:
[0033]
[0034] Ψ pool =MaxPooling(Ψ)
[0035] Among them, GN represents group normalization, Dropout represents regularization operation, and Ψpool represents the input features of adjacent layers on the vertical pyramid.
[0036] Compared with the prior art, the present invention proposes a temporal action localization method based on a double helical convolutional attention pyramid. To this end, three main modules are designed: a spatiotemporal feature optimizer STFO, a multi-scale double helical attention convolution module Ds-MAC and a double helical feature pyramid network Ds-FPN.
[0037] Specifically, the spatiotemporal feature optimizer STFO: downsamples the time dimension and upsamples the channel dimension to enhance the feature representation, and provides processed features for the multi-scale double helix attention convolution module Ds-MAC and the double helix feature pyramid network Ds-FPN; the multi-scale double helix attention convolution module Ds-MAC: establishes associations between feature sequences and integrates them into the output of the attention mechanism, improving the feature discrimination and the accuracy of action boundaries; the double helix feature pyramid network Ds-FPN: alternates between horizontal feature fusion and vertical feature iteration to obtain the final multi-scale video representation.
[0038] Experimental results and analysis show that the present invention demonstrates its superior performance in the task of video temporal action localization. This method effectively solves the problems existing in the existing model, improves the accuracy of temporal action localization, and is suitable for a variety of application scenarios such as intelligent monitoring and video editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly explain the technical solution of the embodiment of the present invention and make the purpose, characteristics and advantages of the present invention more obvious and easy to understand, the specific implementation methods will be described in detail with reference to the accompanying drawings. It should be noted that the drawings described below are only some exemplary embodiments of the present invention. For ordinary technicians in this field, other drawings and implementation methods can be obtained based on these drawings without creative work.
[0040] Figure 1 is a schematic diagram of the architecture of a video temporal action localization network according to an embodiment of the present invention;
[0041] Figure 2 2 is a schematic diagram of the network structure of the spatiotemporal feature optimizer STFO and the multi-scale double-helix attention convolution module Ds-MAC according to an embodiment of the present invention;
[0042] Figure 3 1 is a schematic diagram of the network structure of a lateral pyramid fusion module PFEM and a feed-forward layer (FFN) in a double helix pyramid Ds-FPN according to an embodiment of the present invention;
[0043] Figure 4 is a schematic diagram of visualization results according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present invention can be implemented in a variety of ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. In addition, the technical features in each embodiment can be combined accordingly without conflicting with each other.
[0045] According to an embodiment of the present invention, a video temporal action localization method based on a double spiral convolution attention pyramid can be used in practical applications as follows: Figure 1 As shown, the network structure of the present invention is deployed, including the following steps S1 to S3:
[0046] S1, video features are first extracted by the pre-trained feature extractor, and then the spatiotemporal feature optimizer STFO downsamples the preliminary features in the time dimension and upsamples them in the channel dimension to enhance the feature representation;
[0047] S2, construct a multi-scale double helix attention convolution module Ds-MAC, which first uses the multi-head dynamic mask attention convolution module MDMAC to fine-grainedly encode the features enhanced in step S1, and then uses the dynamic local feature modeling module DLFM to coarse-grainedly encode the features, establish the association between feature sequences, and integrate them into the output of the attention mechanism through adaptive learning;
[0048] S3. Construct a double-helix feature pyramid network Ds-FPN. Each layer of the pyramid performs horizontal and vertical iterative alternating fusion on the features encoded in step S2 to obtain a multi-scale video representation, and sends the representation to the classification head and regression head respectively to obtain the final positioning result.
[0049] For specific examples, please refer to Figure 2 As shown, a pre-step in step S1 is to encode the features into a vertical feature projection layer. This layer consists of a one-dimensional temporal convolution layer with a padding mask, a layer normalization, and a ReLU activation layer. Appropriate masks are added so that variable-length sequences can be processed effectively. This allows for better capture of features within sequences of varying lengths by padding or ignoring. In addition, normalization layers and activation layers are also needed to accelerate network convergence and enhance the representability of the model. Its expression is:
[0050] Ω(x)=ReLU(LN(Mask1dConv(x)))
[0051] Where ReLU, LN, and Mask 1dConv represent ReLU activation, layer normalization, and one-dimensional temporal convolution with padding mask, respectively. T×C′ represents the result after the vertical feature projection layer. T is the time series length and C′ is the channel dimension of the token.
[0052] Different from the previous spatiotemporal feature optimizer STFO, in order to enhance features by exploiting the relationship between different time dimensions and channel dimensions, the present invention designs a spatiotemporal feature optimizer STFO to increase the channel dimension of temporal features while reducing its temporal dimension at each level of the lateral feature pyramid. These features are first encoded into a spatiotemporal feature optimizer STFO, which consists of a normalization layer and a one-dimensional convolution layer with a padding mask. By downsampling the input feature x in the time dimension to reduce the temporal resolution, the model can focus on a longer time range. At the same time, convolution kernels are used to expand the channel dimension of the features, which not only extracts richer feature representations, but also provides more information for the model to learn features in subsequent steps, improving the model's perception ability. After the spatiotemporal feature optimizer STFO, two types of features with different time and channel dimensions are output. The enhanced features are represented by I(x)∈R T′×C′ The calculation formula is as follows
[0053] I(x)=LN(Mask1dConv(LN(Mask1dConv(Ω(x))))
[0054] Among them, LN and Mask1dConv represent layer normalization and 1D temporal convolution with padding mask, respectively.
[0055] As a specific example, please refer to Figure 2 As shown, in step S2, the multi-scale double spiral attention convolution module Ds-MAC performs fine-grained feature encoding and coarse-grained feature encoding on the features. The multi-head dynamic mask attention convolution module MDMAC is designed to capture the long-term dependencies of the features. Deep convolution and layer normalization are performed before calculating Q, K, and V to encode the relative position information between multiple tokens, replacing the position encoding module in the standard multi-head attention mechanism. The output features of the self-attention layer are then connected to the input features through the residual, enabling the network to learn incremental changes relative to the input features, focusing on learning important features rather than just the overall feature map. This makes it easier for the network to be deeply expanded, thereby improving during the gradient descent process, providing a more stable training and optimization process for the model. Given I(x)∈R T ′×C′ , representing T' time steps with D'-dimensional features, using the matrix and to project I(X) to extract feature representations Q, K, and V, called query, key, and value, where d k =d q The outputs Q, K and V are calculated as follows:
[0056] Q=I(X)W Q ,K=I(X)W K ,V=I(X)W V ,
[0057] The output of self-attention is given by:
[0058]
[0059] The output features are normalized and then input into the dynamic local feature modeling module DLFM, which includes two linear layers with unchanged feature channel dimensions, a deep one-dimensional temporal convolution layer, a GELU activation layer, and a Dropout layer. The self-attention mechanism assigns different attention weights to different parts of the input to emphasize important features. However, the attention mechanism itself does not integrate these weighted features. By adding a linear layer after the self-attention layer, the weighted features can be linearly combined to further integrate feature information. In addition, the linear layer has learnable weights and biases, which can adaptively learn specific feature representations adapted to the task through the gradient descent algorithm in the later stage of the model. Adding a linear layer can improve the model's ability to learn parameters and better fit the training data. The one-dimensional convolution has the characteristic of parameter sharing, which effectively reduces the number of parameters and reduces the risk of overfitting, while enhancing the model's ability to extract local patterns and capture local contextual information of the sequence to a certain extent. Therefore, a deep temporal convolution layer is designed to learn local contextual information between short-distance adjacent tokens. The GELU activation layer is used because it has better performance than ReLU when processing a wider range of inputs and has a faster convergence speed than Sigmoid. Then, before being input to this module, the output features of the dynamic local feature modeler are residually connected with the output features of the multi-head dynamic mask attention convolution module MDMAC. It can be given by the following formula:
[0060]
[0061] in, represents the residual connection, x out is the output feature of the multi-head dynamic mask attention convolution module MDMAC, E out_xi Represents the output result of the multi-scale double spiral attention convolution module corresponding to the i-th layer of the horizontal pyramid.
[0062] As a specific example, please refer to Figure 3As shown, in step S3, since the shallow features contain lower semantic levels but more temporal details, while the deep features have higher semantic levels but lack temporal details, using features at different levels can produce a more comprehensive and richer feature representation. The horizontal pyramid feature fusion module PFFM enables the features to fully utilize the information of each layer of the pyramid. The processed features are transmitted along the vertical feature pyramid in a hierarchical manner. The two types of features enhanced from the spatiotemporal feature optimizer STFO are respectively input into a pair of multi-scale double helix attention convolution modules Ds-MAC for capturing temporal context information, and the output features from each horizontal layer are fused by the feature fusion module of the horizontal pyramid. The pyramid feature fusion module PFEM maps the features of each layer of the pyramid in linear layers, ensuring the consistency of the feature dimension when iterating to the adjacent layers of the vertical feature pyramid. Due to the double helix nature of horizontal feature fusion and vertical feature iteration in the pyramid structure, the feature size between adjacent layers of the vertical pyramid must be consistent during temporal downsampling. Therefore, in order to ensure a consistent temporal dimension across each layer of the features on the horizontal pyramid, the nearest neighbor upsampling is applied to the temporal dimension. Then a one-dimensional convolutional layer with a kernel size of 1 is used to learn the weight relationship between features. After that, the features obtained from the shallow layer through linear mapping are fused with the deep layer of the horizontal pyramid to better capture the sequence context information. Finally, the fused features are fused with the features obtained by nearest neighbor upsampling in the channel dimension. It can be written as:
[0063]
[0064] Among them, Linear and Conv1d represent one-dimensional temporal convolution and linear layers, UpSample represents the nearest neighbor upsampling, represents the sum operation, and Concat represents the connection along the feature channel dimension. F = R T×C′ Represents the fused features.
[0065] The output features from the horizontal pyramid feature fusion module PFEM are residually connected to the input features before entering the multi-scale double spiral attention convolution module Ds-MAC, which are fed into the feed-forward network FFN, which consists of a group normalization layer, a one-dimensional temporal convolution layer, a GELU activation layer, and a Dropout regularization layer. The features of the feed-forward layer are then passed to the adjacent layers of the vertical feature pyramid. Finally, the output features from each layer of the vertical feature pyramid are used as the input of the classification and regression heads of the decoder. Output features Ψ∈R T×C′ It is expressed as:
[0066]
[0067] Ψpool =MaxPooling(Ψ)
[0068] Among them, GN represents group normalization, Dropout represents regularization operation, and Ψpool represents the input features of adjacent layers on the vertical pyramid.
[0069] In order to better illustrate the performance of the video temporal action localization method based on the double spiral convolution attention pyramid provided by the present invention, the following will be compared with the existing method:
[0070] The comparison of the accuracy (%) of this model and existing methods on the THUMOS14 dataset is as follows, comparing the mAP and average mAP (AVG) under different tIoU thresholds [0.3:0.1:0.7]:
[0071]
[0072] In summary, with the help of the above technical solutions, the video temporal action localization method based on the double spiral convolution attention pyramid proposed in the present invention first uses the spatiotemporal feature optimizer STFO to enhance the features by exploring the relationship between different time dimensions and channel dimensions. Secondly, in order to solve the problem of low feature discrimination in temporal feature modeling based on the self-attention mechanism, a multi-scale double spiral attention convolution module (Ds-MAC) is proposed, which can not only accurately capture the long-term dependency of actions, but also can perform fine-grained encoding of local action details. In addition, in order to solve the problem of detailed feature loss in the top-down feature mapping process in the temporal action localization method based on the traditional feature pyramid, a double spiral feature pyramid network Ds-FPN is designed. The network adopts the alternating coding strategy of horizontal fusion and vertical iteration to establish the relationship between shallow and deep features, which significantly improves the performance of temporal action localization.
[0073] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A video temporal action localization method based on double spiral convolution attention pyramid, characterized in that: The following steps are involved: S1, video features are first extracted by the pre-trained feature extractor, and then the spatiotemporal feature optimizer STFO downsamples the preliminary features in the time dimension and upsamples them in the channel dimension to enhance the feature representation; S2, construct a multi-scale double helix attention convolution module Ds-MAC, which first uses the multi-head dynamic mask attention convolution module MDMAC to fine-grainedly encode the features enhanced in step S1, and then uses the dynamic local feature modeling module DLFM to coarse-grainedly encode the features, establish the association between feature sequences, and integrate them into the output of the attention mechanism through adaptive learning; S3. Construct a double-helix feature pyramid network Ds-FPN. Each layer of the pyramid performs horizontal and vertical iterative alternating fusion on the features encoded in step S2 to obtain a multi-scale video representation, and sends the representation to the classification head and regression head respectively to obtain the final positioning result.
2. According to claim 1, a method for localizing video temporal actions based on a double spiral convolution attention pyramid is characterized in that: The specific process of step S1 is as follows: S11. Use the vertical feature mapping layer to reduce the channel dimension of the feature. Its structure is as follows: Before downsampling the time dimension and upsampling the channel dimension of the preliminary features, the features are encoded into the vertical feature projection layer. The vertical feature projection layer consists of a one-dimensional temporal convolution layer with a padding mask, a layer normalization, and a ReLU activation layer. The one-dimensional temporal convolution is combined with a padding mask to effectively process variable-length sequences. The calculation formula is as follows: Ω(x)=ReLU(LN(Mask1dConv(x))) Where x∈R T×C represents the video features extracted by the pre-trained feature extraction network in step S1, T is the time series length, C is the channel dimension corresponding to the feature; ReLU, LN and Mask 1dConv represent the ReLU activation layer, layer normalization and one-dimensional temporal convolution with padding mask, respectively; represents the features after vertical feature mapping layer mapping, C * It is the channel dimension corresponding to the feature after being vertically mapped; S12. Construct a spatiotemporal feature optimizer STFO that expands the channel dimension and reduces the time series dimension. Its structure is as follows: A spatiotemporal feature optimizer STFO is designed to enhance the channel dimension of temporal features while reducing their temporal dimension at each level of the horizontal feature pyramid. The temporal features are first encoded into the spatiotemporal feature optimizer STFO, which consists of a normalization layer and a one-dimensional temporal convolutional layer with a padding mask. After the spatiotemporal feature optimizer STFO, features with different temporal and channel dimensions are output. The enhanced features are represented by I(x)∈R T′×C′ The calculation formula is as follows I(x)=LN(Mask1dConv(LN(Mask1dConv(Ω(x)))) Among them, LN and Mask1dConv represent layer normalization and one-dimensional temporal convolution with padding mask, respectively.
3. The method for localizing video temporal actions based on a double spiral convolution attention pyramid according to claim 1, characterized in that: In step S2, the multi-scale double-helix attention convolution module Ds-MAC performs fine-grained feature encoding and coarse-grained feature encoding on the features; the specific method is: S21. Construct the multi-head dynamic mask attention convolution module MDMAC in the multi-scale double spiral attention convolution module Ds-MAC, whose structure is as follows: The multi-head dynamic masked attention convolution module MDMAC is designed to help features capture long-range dependencies and encode relative position information before calculating the query Q, key K, and value V through deep convolution and layer normalization, replacing the position encoding module in the standard multi-head attention mechanism; then, the output features of the self-attention layer are connected to the input features through the residual, enabling the network to learn incremental changes relative to the input features; S22. Construct the dynamic local feature modeling module DLFM in the multi-scale double helix attention convolution module Ds-MAC, whose structure is as follows: The dynamic local feature modeling module DLFM consists of a linear layer with unchanged feature channel dimension, a one-dimensional temporal convolution layer, a GELU activation layer and a Dropout regularization layer. This module combines the different attention weights assigned by the self-attention mechanism to integrate feature information and learn feature representations for specific tasks. By adding a linear layer after the self-attention layer, the weighted features are linearly combined to further integrate feature information. Subsequently, the output features of the dynamic local feature modeler are residually connected with the output features of the multi-head dynamic mask attention convolution module MDMAC, which can be expressed as the following formula: in, represents the residual connection, x out is the output feature of the multi-head dynamic mask attention convolution module MDMAC, E out_xi Represents the output result of the multi-scale double spiral attention convolution module corresponding to the i-th layer of the horizontal pyramid.
4. The method for localizing video temporal actions based on a double spiral convolution attention pyramid according to claim 1, characterized in that: The double-helix feature pyramid network Ds-FPN used in step S3 combines horizontal feature fusion with vertical feature iteration to optimize the connection between layers; by providing richer detail features from the horizontal pyramid of the base layer and the reference layer to the deep vertical pyramid, it effectively supplements the detail information lost in the vertical pyramid feature mapping process, thereby improving the positioning accuracy.
5. The method for localizing video temporal actions based on a double spiral convolution attention pyramid according to claim 4, characterized in that: The double-helix feature pyramid network Ds-FPN performs horizontal feature fusion on each pyramid level, that is, the features with different temporal dimensions and channel dimensions output from step S2 are fused to retain more detailed features; the pyramid feature fusion module PFFM maps the features of each layer of the pyramid according to the linear layer, ensuring the consistency of the feature dimension when iterating to the adjacent layers of the vertical feature pyramid; the double-helix nature of horizontal feature fusion and vertical feature iteration in the pyramid structure, the feature dimensions between adjacent layers of the vertical pyramid must be kept consistent during the temporal downsampling; to ensure the consistent time dimension of each layer of the features across the horizontal pyramid, the nearest neighbor upsampling is applied to the temporal dimension; then a one-dimensional convolutional layer with a kernel size of 1 is used to learn the weight relationship between the features, and then the features obtained from the shallow layer are fused with the deep features of the horizontal pyramid through linear mapping to better capture the sequence context information; at the same time, the vertical feature iteration is performed on each pyramid level, that is, the feature representation is gradually enhanced through 7 rounds of iterative processing to ensure that the feature information at different scales is fully integrated, and the specific formula is as follows: Among them, Linear and Conv1d represent one-dimensional temporal convolution and linear layers, UpSample represents the nearest neighbor upsampling, represents the addition operation, Concat represents the concatenation operation along the feature channel dimension; F = R T×C′ Represents the fused features.
Citation Information
Cited By
Time sequence action detection method based on frequency domain and time domain information interaction
CN121259702A