Time sequence action positioning method and device based on bidirectional interaction and dynamic feature enhancement
By introducing two-way interaction and dynamic feature enhancement technology into the timing action positioning method, the problem of short-term and long-term dependency balance in timing action positioning is solved, and more efficient timing action positioning performance is achieved.
Patent Information
- Application Number
- CN202510194515.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing timing action positioning methods have the problem of balance between short-term and long-term dependencies when capturing complex timing dependencies, resulting in unstable positioning accuracy in different action scenarios.
The timing action positioning method based on bidirectional interaction and dynamic feature enhancement is adopted. Through the multi-scale dynamic timing modeling module and the global and local adaptive bidirectional interaction module, the receptive field and feature weight are dynamically adjusted to achieve efficient balance and fusion of short-term and long-term dependencies.
It significantly improves the overall performance of timing action positioning tasks, can achieve the current state-of-the-art performance on multiple TAL benchmark data sets, and verifies its powerful ability and effectiveness in capturing different simultaneous sequence modes.
Smart Images

Figure CN120220013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and particularly to a temporal action localization method and device based on two-way interaction and dynamic feature enhancement. Background Art
[0002] Temporal Action Localization (TAL) is an important task in video analysis, aiming to locate and identify the action boundaries and categories in untrimmed videos. With the rapid growth of video data on different platforms, the importance of TAL has increased significantly as it enables efficient video understanding and retrieval, powering applications such as video surveillance, motion analysis, autonomous driving, and human-computer interaction. Existing TAL methods are generally divided into two-stage methods and single-stage methods, each with its own advantages and disadvantages. The two-stage method first generates action proposals and then further processes and classifies these proposals through a classifier, while the single-stage method directly integrates proposal generation and classification into an end-to-end process.
[0003] The two-stage method usually includes two steps: proposal generation and proposal classification. The purpose of proposal generation is to identify potential action segments from the video, usually using methods such as sliding window, boundary prediction, or frame action degree evaluation. The sliding window-based method generates temporal proposals by applying multi-scale sliding windows in the video, such as TURN TAP. The boundary-based method locates the temporal boundaries of actions by predicting the boundary confidence at specific positions in the video, such as BMN. The frame action degree-based method generates proposals by evaluating the action degree scores of each frame, such as SSN. These methods usually rely on a well-trained classifier to classify each proposal and predict the action category.
[0004] The single-stage method simplifies the process of action proposal generation and classification through end-to-end training, reducing the complexity of training and inference. The single-stage method does not rely on traditional temporal anchors or sliding windows, but directly predicts the temporal span and category of actions at the video frame level. For example, the anchor-free TAL model directly predicts the action category for each video frame and simultaneously estimates the corresponding action start and end times. In addition, single-stage models based on transformers such as ActionFormer enhance the modeling ability for actions at different scales through multi-scale self-attention mechanisms. PBBNet improves the accuracy of action localization by gradually refining the prediction, and TriDet uses a granularity-aware layer to optimize temporal modeling.
[0005] Although existing two-stage and single-stage methods have made significant progress in the task of temporal action localization, they generally face the challenge of capturing complex temporal dependencies. Especially when dealing with actions with diverse time spans and complex temporal dependencies, how to effectively balance short-term and long-term dependencies remains the key to improving the accuracy of temporal action localization. Existing methods often have biases in modeling short-term and long-term dependencies, resulting in the inability to maintain high localization accuracy under the temporal characteristics of different actions. Short-term dependencies usually involve the details and local changes of actions and are suitable for precise localization at the start moment of an action; while long-term dependencies focus on the continuation and integrity of actions and are particularly crucial for the end moment and duration of an action. Traditional methods either overemphasize short-term information and neglect the overall evolution of actions, or focus on long-term information, leading to a decrease in sensitivity to local action changes. Therefore, how to adaptively adjust the balance between short-term and long-term dependencies in actions of different scales is the key to improving the accuracy of temporal action localization.
[0006] Temporal action localization (TAL) is a challenging visual task aiming to accurately classify and localize all actions in untrimmed videos. Due to the multi-scale nature of action intervals, short-term actions rely on local features while long-term actions rely on global information, which makes alleviating complex temporal dependency problems a core challenge that has long existed in the field of TAL. Against this background, the quality of features is crucial for improving the performance of temporal action localization. However, existing TAL methods usually rely on pre-trained features as input, and the expressive power of these features is limited by the pre-trained models, making it difficult to comprehensively capture the complex temporal dependencies of actions. Therefore, in the TAL task, how to enhance the representation ability of features through feature enhancement to more efficiently support complex temporal modeling has become a key issue that cannot be ignored. At the same time, to better model short-term and long-term dependencies, it is necessary to further explore effective mechanisms for temporal modeling, including convolutional-based methods and Transformer-based methods.
[0007] (1)Convolution-based TAL method: In the past few years, convolutional neural networks (CNNs) have been the key driving force in the development of computer vision. Since the landmark emergence of AlexNet, a series of highly influential CNN architectures have successively emerged, and they have performed excellently in many tasks in the field of image understanding, significantly improving performance metrics. In recent years, as videos have become a key data source in many real-world scenarios, due to the excellent timeliness of CNNs themselves, researchers have actively explored ways to extend them to the video field. Especially in the TAL task, the convolution neural network (CNN)-based TAL method, with its local connectivity and parameter sharing characteristics, has shown significant advantages in reducing redundancy and accelerating the computing process. Although CNNs perform excellently in dealing with short-term actions with clear boundaries, when faced with complex action sequences spanning a long time, their limitations make it difficult to effectively model global temporal dependencies, resulting in insufficient performance in dealing with long temporal dependencies. Therefore, later methods capture long-term information by expanding the receptive field of convolutional kernels. Although they alleviate the temporal dependency problem to a certain extent, they either lack sufficient dynamics and flexibility or fail to fully consider the balance between global and local features, resulting in unstable performance of the model in action localization at different scales. However, these methods still have limitations. Although they alleviate the temporal dependency problem to a certain extent, they either lack sufficient dynamics and flexibility or fail to fully consider the balance between global and local features, resulting in unstable performance of the model in action localization at different scales.
[0008] (2) Transformer-based TAL methods: With its powerful modeling capabilities, Transformer has achieved remarkable results in computer vision and natural language processing tasks in recent years. This advantage has also led to the widespread application of Transformer in the field of temporal action localization. For TAL tasks, Transformer significantly improves the performance of action localization and classification by capturing long-distance dependencies between video frame features. On this basis, Transformer-based TAL methods continue to emerge and show a diversified development trend. Some end-to-end modeling methods simplify the complexity of traditional TAL processes by building a trainable framework. For example, TALLFormer combines short-term Transformer encoders and long-term memory mechanisms to efficiently model video actions while effectively reducing GPU memory overhead; TadTR is based on the Transformer with a deformable attention mechanism, which selectively focuses on key sparse context subsets in the video, thereby improving efficiency and performance. In addition, some methods further optimize the use of global information by enhancing context-aware capabilities. For example, ActionFormer uses a multi-head self-attention mechanism to directly capture the global temporal context and achieves efficient modeling of global temporal relationships; while SAFormer introduces a self-attention mechanism for global channel feature responses and classification refinement modification loss to build an efficient one-stage Transformer model for optimizing action localization and classification. Although these Transformer-based methods have made significant progress in TAL tasks, they still have limitations in handling local redundant information and balancing long-term and short-term dependencies in complex video data. The key problem is the lack of an effective two-way interaction mechanism to simultaneously capture local and global features, so as to better adapt to complex action temporal relationships.
[0009] (3)The Importance of Feature Enhancement in TAL: Although significant progress has been made in temporal action localization (TAL) methods in recent years, there is still considerable room for improvement at present. Existing TAL methods mostly rely on annotated and untrimmed video materials during the training process. Unfortunately, the existing TAL datasets are relatively small in scale, and this limitation severely restricts the training effectiveness and generalization potential of the model, making it difficult to perform accurately in a wider range of scenarios. Compared with image datasets, video data exhibits more complex and diverse characteristics due to the additional time dimension, which undoubtedly adds many difficult problems to the TAL task. Based on this, how to effectively expand the existing data to help the model performance achieve a leap has become a key point that needs to be overcome urgently. Although data augmentation is a simple and feasible strategy, video data often has a long duration. If conventional data augmentation methods are directly applied, it is very likely to cause high computational costs, greatly reducing the efficiency of the entire process. In addition, most current TAL methods rely on pre-trained features as input, which makes feature-level enhancement particularly important in the TAL task. Summary of the Invention
[0010] To solve the technical problem of how to adaptively adjust the balance between short-term and long-term dependencies in actions of different scales in the prior art, an embodiment of the present invention provides a temporal action localization method and device based on bidirectional interaction and dynamic feature enhancement. The technical solution is as follows:
[0011] On the one hand, a temporal action localization method based on bidirectional interaction and dynamic feature enhancement is provided, which is characterized in that the method includes:
[0012] S1. Obtain an uncropped video, process the uncropped video through a pre-trained network, and extract initial features;
[0013] S2. Construct an adaptive feature enhancement strategy, and perform hierarchical network architecture integration dynamic modeling and local and global temporal interaction modeling through the adaptive feature enhancement strategy; the adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module and a global and local adaptive bidirectional interaction module;
[0014] S3. Introduce the adaptive feature enhancement strategy into the encoder; perform feature enhancement on the initial features through the encoder to generate enhanced temporal features;
[0015] S4. Transmit the enhanced temporal features to the classification and regression head to predict the action category and time boundary.
[0016] Optionally, in S2, the multi-scale dynamic temporal modeling module includes;
[0017] The multi-scale dynamic time series modeling module includes three dynamic local DynL affines; through the dynamic adjustment of the receptive field of the multi-scale dynamic time series modeling module, multi-scale temporal features are captured.
[0018] Optionally, through the dynamic adjustment of the receptive field of the multi-scale dynamic time series modeling module, multi-scale temporal features are captured, including:
[0019] Through the dynamic local affine transformation DynL affine, the dynamic changes of actions at different time scales in the TAL task are alleviated;
[0020] Given the input features, DynL affine processes the input features using two different-scale deep convolutional branches to obtain multi-scale temporal information;
[0021] Through the hierarchical convolution method, the first branch performs depth convolution with a smaller kernel to capture local temporal information; the second branch first convolves with kernel k1 and then with a larger kernel k2 to gradually expand the receptive field to capture a wider context;
[0022] The outputs of the two branches are fused to obtain a multi-scale sequence feature representation;
[0023] A learnable mask generation module is constructed, and the multi-scale sequence feature representation is input into the mask generation module to generate a mask sequence, which assigns different weights to each position in the sequence, highlighting the features of key frames while suppressing background information;
[0024] The convolutional layer is used to extract local patterns from the sequence and then refined by batch normalization BN; the sigmoid activation function is used to compress the weights into the range of [0, 1] to obtain the final mask sequence; the mask sequence is multiplied element-wise with the multi-scale sequence features to generate weighted sequence features weighted.
[0025] Optionally, in S2, the global and local adaptive bidirectional interaction module includes:
[0026] The dynamic local affine block DynL affine and the global Global affine are combined through the adaptive interaction feature fusion sub-module AIFF for bidirectional interaction to form the global and local adaptive bidirectional interaction module;
[0027] The global features are aggregated through the global affine module; the local features are aggregated through the dynamic local affine block DynL affine; the global and local information are fused through the adaptive interaction feature fusion sub-module AIFF.
[0028] Optionally, aggregating global features through the global affine module includes:
[0029] Incorporate the self-attention mechanism into the global affine module to obtain attention-based global feature aggregation; aggregate global features through the global affine module and enhance the model's ability to recognize long-term relationships within the sequence to obtain robust action boundary localization.
[0030] Optionally, fuse global and local information through the Adaptive Interaction Feature Fusion sub-module AIFF, including:
[0031] Dynamically assign weights through the Adaptive Interaction Feature Fusion sub-module for feature fusion, balancing and fusing long-term and short-term dependencies in time series information.
[0032] Optionally, dynamically assign weights through the Adaptive Interaction Feature Fusion sub-module for feature fusion, balancing and fusing long-term and short-term dependencies in time series information, including:
[0033] Given two input feature tensors;
[0034] Connect the two input feature tensors along the channel dimension through the Adaptive Interaction Feature Fusion sub-module to form a fused feature tensor. The connected features will undergo a series of transformations to generate dynamic weights. Use temporal average pooling to reduce the temporal dimension;
[0035] Extract key context information from the merged features through a convolutional layer;
[0036] Apply the sigmoid activation function to generate adaptive weights for local and global features;
[0037] Refine the obtained weights through one-dimensional convolution to reduce the dimension and obtain the final weights; weight the obtained weights and the original features to obtain the final output.
[0038] On the other hand, a temporal action localization device based on bidirectional interaction and dynamic feature enhancement is provided. This device is applied to the temporal action localization method based on bidirectional interaction and dynamic feature enhancement. The device includes:
[0039] An initial feature extraction module for obtaining an uncropped video, processing the uncropped video through a pre-trained network, and extracting initial features;
[0040] A feature enhancement strategy construction module for constructing an adaptive feature enhancement strategy, performing hierarchical network architecture integration dynamic modeling and local and global temporal interaction modeling through the adaptive feature enhancement strategy; The adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module and a global and local adaptive bidirectional interaction module;
[0041] A temporal feature enhancement module for introducing the adaptive feature enhancement strategy into the encoder; enhancing the initial features through the encoder to generate enhanced temporal features;
[0042] A prediction module for passing the enhanced temporal features to the classification and regression heads to predict the action categories and time boundaries.
[0043] On the other hand, there is provided a temporal action localization device based on bidirectional interaction and dynamic feature enhancement. The temporal action localization device based on bidirectional interaction and dynamic feature enhancement includes: a processor; a memory storing computer-readable instructions thereon, and when the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned temporal action localization method based on bidirectional interaction and dynamic feature enhancement is implemented.
[0044] On the other hand, there is provided a computer-readable storage medium storing at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned temporal action localization method based on bidirectional interaction and dynamic feature enhancement.
[0045] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0046] In the embodiments of the present invention: 1. The present invention proposes a brand-new adaptive temporal enhancement framework, aiming to solve the complex temporal dependency problem in the temporal action localization (TAL) task through the bidirectional dynamic interaction and balance of short-term and long-term features;
[0047] 2. The present invention designs a multi-scale dynamic temporal modeling module (Multi-Scale Dynamic Temporal Modeling, MS-DyTM), which integrates a multi-scale adaptive convolution kernel selection and a learnable mask mechanism, and can dynamically adjust the receptive field to efficiently capture multi-scale temporal features, thereby significantly improving the discriminability of features and enhancing the accuracy and flexibility of the model in different action scenarios;
[0048] 3. The present invention proposes an adaptive interaction feature fusion sub-module (Adaptive Interaction Feature Fusion, AIFF), which dynamically adjusts the feature weights through an attention allocation mechanism, realizes the efficient balance and fusion of short-term and long-term temporal dependencies, and thus significantly improves the overall performance of the temporal action localization task;
[0049] 4. Extensive experiments of the present invention on multiple TAL benchmark datasets show that the proposed method can efficiently achieve the current state-of-the-art performance, verifying its powerful ability and effectiveness in capturing different temporal patterns in untrimmed videos. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0051] Figure 1 It is a schematic flowchart of a temporal action localization method based on bidirectional interaction and dynamic feature enhancement provided by an embodiment of the present invention;
[0052] Figure 2 It is the overall model architecture diagram provided by an embodiment of the present invention;
[0053] Figure 3 It is a block diagram of a temporal action localization device based on bidirectional interaction and dynamic feature enhancement provided by an embodiment of the present invention;
[0054] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments
[0055] The following will describe the technical solutions in the present invention with reference to the accompanying drawings.
[0056] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0057] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning to be expressed is the same.
[0058] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0059] The embodiments of the present invention provide a temporal action localization method based on bidirectional interaction and dynamic feature enhancement. This method can be implemented by a temporal action localization device based on bidirectional interaction and dynamic feature enhancement. The temporal action localization device based on bidirectional interaction and dynamic feature enhancement can be a terminal or a server. As Figure 1 shown in the flowchart of the temporal action localization method based on bidirectional interaction and dynamic feature enhancement, as Figure 1As shown in the figure, the proposed temporal action localization method based on bidirectional interaction and dynamic feature enhancement, the processing flow of this method can include the following steps:
[0060] S1. Obtain the uncropped video, process the uncropped video through the pre-trained network, and extract the initial features;
[0061] S2. Construct an adaptive feature enhancement strategy, and perform hierarchical network architecture integrated dynamic modeling and local and global temporal interaction modeling through the adaptive feature enhancement strategy; the adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module, a global and local adaptive bidirectional interaction module.
[0062] In a feasible implementation, the adaptive feature enhancement strategy performs hierarchical network architecture integrated dynamic modeling and local and global temporal interaction modeling, solving the complex temporal dependence problem in temporal action localization. This strategy combines the core modules - the multi-scale dynamic temporal modeling module (MS-DyTM), dynamic local affine (DynL affine), global affine (Global affine), and adaptive interactive feature fusion (AIFF) module - to dynamically model and effectively balance short-term and long-term dependencies.
[0063] The overall architecture of the model is as Figure 2 shown, aiming to solve the challenge of complex temporal dependencies in temporal action localization through the "adaptive feature enhancement strategy". First, the uncropped video is processed through the pre-trained network to extract the initial feature representation form Fpre, and then the features are enhanced through the encoder to generate the enhanced temporal features FEnh. Finally, the rich features FEnh are passed to the classification and regression head to predict the action category and time boundaries. To achieve dynamic modeling and effectively capture the balance between short-term and long-term dependencies, the present invention introduces the "adaptive feature enhancement strategy" into the encoder, which mainly includes: the MS-DyTM module, the DynL affine module, the Global affine module, and the AIFF module. Next, the core modules of the model will be described in detail.
[0064] In a feasible implementation, in S2, the multi-scale dynamic temporal modeling module includes;
[0065] The multi-scale dynamic temporal modeling module includes three dynamic local DynL affines; allowing dynamic adjustment of the receptive field to effectively capture multi-scale temporal features, improve the discriminability of the features, and enhance the accuracy and flexibility of the model in different action scenarios.
[0066] In a feasible implementation, the adaptive feature enhancement strategy integrates dynamic local and global temporal modeling through a hierarchical network architecture, solving the complex temporal dependency problem in temporal action localization. This strategy combines three core modules - multi-scale dynamic temporal modeling (MS-DyTM) (dynamic local (DynL affine), global affine (Global affine), and adaptive interactive feature fusion (AIFF) modules - to effectively balance short-term and long-term dependencies, as Figure 2 shown in (a).
[0067] To improve efficiency and effectiveness, the present invention assigns different modeling strategies to different layers.
[0068] In the shallow layer, MS-DyTM consists of three layers of DynL affine, focusing on dynamically capturing fine-grained local temporal features using adaptive convolutional kernels, ensuring accurate modeling of short-term dependencies and enhancing action boundary detection. By focusing on local temporal segments, the shallow layer retains detail-oriented features while avoiding premature consideration of global context, thus improving efficiency and accuracy.
[0069] In the deep layer study, global affine captures the overall temporal relationship and long-term dependencies through a self-attention mechanism, providing a comprehensive understanding of global action semantics. The AIFF module acts as a bridge, supporting two-way interaction between the DynL affine and global affine outputs. This module adaptively fuses localized details with global context, dynamically balancing short-term and long-term information, thus enhancing feature integration. This hierarchical design takes advantage of the shallow layer for efficient local modeling and the deep layer for global context understanding to ensure comprehensive temporal feature representation across different time scales. By combining dynamic local modeling with adaptive global integration, this framework significantly improves the accuracy and localization efficiency of action detection and localization in the TAL task. This design ensures that the model adapts to actions of different time scales, improving the accuracy of action detection in the TAL task.
[0070] In a feasible implementation, through the dynamic adjustment of the receptive field of the multi-scale dynamic temporal modeling module, multi-scale temporal features are captured, including:
[0071] By the dynamic local affine transformation DynL affine, the dynamic changes of actions at different time scales in the TAL task are alleviated;
[0072] Given the input features, DynL affine processes the input features using two different-scale deep convolutional branches to obtain multi-scale temporal information;
[0073] Through the hierarchical convolution method, the first branch performs depth convolution with a smaller kernel to capture local temporal information; the second branch first convolves with kernel k1 and then with a larger kernel k2, gradually expanding the receptive field to capture a wider context;
[0074] Fuse the outputs of the two branches to obtain a multi-scale sequence feature representation;
[0075] Construct a learnable mask generation module, input the multi-scale sequence feature representation into the mask generation module to generate a mask sequence, assign different weights to each position in the sequence, and highlight the features of key frames while suppressing background information;
[0076] Use a convolutional layer to extract local patterns from the sequence and then refine them through batch normalization (BN) to enhance expressiveness;
[0077] Use the sigmoid activation function to compress the weights into the range [0, 1] to obtain the final mask sequence;
[0078] Element-wise multiply the mask sequence with the multi-scale sequence features to generate weighted sequence features weighted.
[0079] In a feasible implementation, the present invention introduces a multi-scale dynamic temporal modeling module (MS-DyTM), which effectively alleviates the challenges brought by the dynamic changes of actions at different time scales in the TAL task by using dynamic local affine transformation (DynL affine). This module integrates multi-scale adaptive kernel selection and a learnable mask mechanism, allowing for dynamic adjustment of the receptive field to effectively capture multi-scale temporal features, improve the discriminability of features, and make its detection of the start and end moments of actions more accurate.
[0080] The implementation of DynL affine will be introduced below, as Figure 2 shown in (b).
[0081] DynL affine consists of multi-scale convolution and a mask mechanism. Given the input features, the DynL module first processes the input features using two depth convolution branches with different scales to obtain multi-scale temporal information. The first branch performs depth convolution with a smaller kernel to capture local temporal information; the second branch first convolves with kernel k1 and then with a larger kernel k2, gradually expanding the receptive field to capture a wider context. Through this hierarchical convolution method, the model can maintain smoothness while capturing a larger temporal context, effectively obtaining a wider range of temporal information. Then, the outputs of the two branches are fused to obtain a multi-scale sequence feature representation.
[0082] To adaptively weigh the key positions in the sequence, the DynL module designs a learnable mask generation module. The input of this module is the fused multi-scale sequence features, and its goal is to generate a mask sequence that assigns different weights to each position in the sequence, highlighting the features of key frames while suppressing background information. A convolutional layer is used to extract local patterns from the sequence, which are then refined by batch normalization (BN) to enhance expressiveness; the sigmoid activation function is used to compress the weights into the range [0, 1] to obtain the final mask sequence; then, the mask sequence is multiplied element-wise with the multi-scale sequence features to generate weighted sequence features. Through this weighting method, the model can emphasize the feature expressions of key positions and suppress redundant background information. Then, the weighted features are processed by depth convolution to adaptively adjust the receptive field, generating two attention maps, and then activating the original features respectively to obtain the output features of the DynL module.
[0083] Compared with the static feature learning mechanism used in most existing TAL models, the DyTM module introduces an innovative dynamic feature learning method. Traditional TAD models usually adopt a fixed receptive field in the convolutional kernel or attention mechanism, assigning equal weights to all temporal features. In contrast, the DyTM module adaptively selects the receptive field through multi-scale convolution to capture features in different time ranges. In addition, DyTM uses a learnable mask to assign different weights to each frame, dynamically focusing on the critical moments when actions occur while suppressing irrelevant frames. This improves the accuracy of feature representation.
[0084] In a feasible implementation manner, in S2, the global and local adaptive bidirectional interaction module includes:
[0085] The dynamic local affine block DynL affine and the global Global affine interact bidirectionally through the adaptive interaction feature fusion sub-module AIFF, and are combined to form the global and local adaptive bidirectional interaction module;
[0086] Aggregate global features through the global affine module; aggregate local features through the dynamic local affine block DynL affine; fuse global and local information through the adaptive interaction feature fusion sub-module AIFF.
[0087] In a feasible implementation manner, aggregating global features through the global affine module includes:
[0088] Incorporate the self-attention mechanism into the global affine module to obtain attention-based global feature aggregation; aggregate global features through the global affine module and enhance the model's ability to identify long-term relationships within the sequence to obtain robust action boundary localization.
[0089] In a feasible implementation manner, in TAL, capturing the stochastic correlation across time series is crucial for accurately identifying the start and end of an action. The global affine module addresses this challenge by aggregating global features and enhancing the model's ability to recognize long-term relationships within the sequence, ensuring robust action boundary localization. To better capture the long-term dependencies in the time series data, the present invention combines an attention mechanism to implement global affine transformation. By incorporating attention into the global affine module, the global affine formula is modified to include attention-based feature aggregation. By combining self-attention with affine transformation, the global affine module can effectively capture global feature interactions, enabling the model to better understand the temporal relationships between the entire sequences. This makes more precise action localization possible, especially in scenarios where the action spans multiple time steps or exhibits complex temporal dependencies.
[0090] In a feasible implementation manner, the global and local information is fused through the Adaptive Interactive Feature Fusion sub-module AIFF, including:
[0091] By dynamically assigning weights to achieve effective feature fusion, it effectively balances and fuses the long-term and short-term dependencies in the time series information, improving the overall performance of the TAL task.
[0092] In a feasible implementation manner, through the Adaptive Interactive Feature Fusion sub-module to dynamically assign weights for feature fusion, balancing and fusing the long-term and short-term dependencies in the time series information, including:
[0093] The Adaptive Interactive Feature Fusion sub-module (AIFF) is adopted to effectively balance and fuse the long-term and short-term dependencies in the time series information. This module utilizes the combination of convolutional operations and attention mechanism to achieve effective feature fusion by dynamically assigning weights, improving the overall performance of the Temporal Action Localization (TAL) task. Obtain the DynL local output feature and the Global module output feature as the input feature tensors; connect the two input feature tensors (local and global feature tensors) along the channel dimension through the Adaptive Interactive Feature Fusion sub-module to form a fused feature tensor, and the connected features will undergo a series of transformations to generate dynamic weights. Use temporal average pooling to reduce the temporal dimension; extract key context information from the merged features through a convolutional layer; apply the sigmoid activation function to generate the adaptive weights of the local and global features; refine the obtained weights through one-dimensional convolution to reduce the dimension and obtain the final weights; weight the obtained weights and the original features to obtain the final output. This adaptive fusion process ensures that AIFF can effectively weigh the short-term and long-term dependencies based on the task context, enabling the model to better capture the temporal dynamics inherent in action localization. As Figure 2 shown in (c).
[0094] S3. Introduce an adaptive feature enhancement strategy into the encoder; enhance the initial features through the encoder to generate enhanced temporal features;
[0095] S4. Pass the enhanced temporal features to the classification and regression heads to predict the action class and time boundary.
[0096] In a feasible implementation, as shown in Table 1 below, the performance differences between the method of the present invention and the current state-of-the-art methods on the THUMOS14 dataset are compared. To comprehensively evaluate the effectiveness of the method of the present invention, the present invention conducts experiments under various pre-trained features, including I3D, InterVideo, and VideoMAE, to verify the applicability and robustness of the method in different backbone architectures.
[0097] Table 1: Comparison results with advanced methods on the THUMOS14 dataset
[0098]
[0099] On the classic I3D backbone network, the method of the present invention achieves an average mAP of 69.4%, which is 2.6% higher than ActionFormer, and its performance is sufficient to compete with TriDet. On the stronger InterVideo backbone, the method of the present invention further increases the average mAP to 74.6%, surpassing all existing methods and reaching the SOTA performance. Among them, it is 3.0% higher than ActionFormer, 2.2% higher than DyFADet, and 1.9% higher than ActionMamba. Based on the VideoMAE backbone, the method of the present invention also performs excellently, with the average mAP being 1.5% higher than ActionFormer, 1.3% higher than MFAM, and 1.0% higher than TriDet.
[0100] These results show that the method of the present invention performs particularly significantly on InterVideo and VideoMAE features, effectively overcoming their limitations from the perspective of feature enhancement. For InterVideo features, the method of the present invention adaptively enhances the expressive ability of the time span and complex feature representation through a dynamic modeling function, and fully excavates the rich semantic and temporal information across videos. For VideoMAE features, the method of the present invention further strengthens the interaction between global and local features, and ensures accurate temporal action modeling by dynamically balancing short-term and long-term dependencies, effectively overcoming the inherent static limitations of pre-trained features. These performance improvements fully demonstrate the superiority and wide applicability of the method of the present invention in complex temporal modeling tasks.
[0101] Table II: Comparison Results with Advanced Methods on the HACS Dataset
[0102]
[0103] As shown in Table II, the performance of the method on the HACS dataset reports the average mAP at [0.5, 0.75, 0.95] tIoU thresholds, and the best results are marked in bold. It can be seen that the method of the present invention achieved an impressive 45.1% mAP on the HACS dataset, outperforming other state-of-the-art methods. Notably, even at higher tIoU thresholds, the method of the present invention still exhibits significant advantages. At tIoU = 0.95, the method of the present invention is 1.5% higher than DyFADet, 2.3% higher than ActionMamba, and 2.5% higher than TriDet, fully demonstrating its strong robustness in accurate action boundary localization. Especially under the challenging high-threshold conditions, the consistent advantages shown by the method of the present invention further prove its excellent ability in modeling complex temporal dependencies and adaptively balancing short-term and long-term interactions.
[0104] Table III: Comparison Results with Advanced Methods on the ActivityNet-1.3 Dataset
[0105]
[0106] As shown in Table III, the experimental results of the method of the present invention on the ActivityNet-1.3 dataset verify its effectiveness under I3D and InterVideo feature backbones. When using I3D features, the method of the present invention achieved a significant improvement in terms of average mAP compared to the ActionFormer method, with a performance gain of 1.4%. This indicates that even in the case of relatively limited feature representation, the method of the present invention can still effectively improve performance. When using InterVideo features, the method of the present invention achieved stable performance improvements at all tIoU thresholds, especially at higher thresholds (such as tIoU = 0.95), which fully reflects its strong ability to accurately capture action boundaries. In terms of average mAP, the method of the present invention is 1.1% higher than TriDet. These experimental results strongly prove the excellent performance and robustness of the method of the present invention in dealing with different feature representation challenges.
[0107] In summary, extensive experiments were conducted on the THUMOS14, HACS, and ActivityNet-1.3 datasets. The results show that the proposed method comprehensively outperforms existing state-of-the-art methods in terms of performance. The proposed algorithm provides an effective solution to solve the complex temporal dependence problem, and at the same time opens up new directions and ideas for further research in the field of temporal action localization.
[0108] In the embodiments of the present invention, a novel adaptive temporal enhancement framework is proposed, aiming to solve the complex temporal dependence problem in the temporal action localization (TAL) task through the bidirectional dynamic interaction and balance of short-term and long-term features;
[0109] The present invention designs a multi-scale dynamic temporal modeling module (MS-DyTM), which integrates multi-scale adaptive convolutional kernel selection and a learnable mask mechanism, and can dynamically adjust the receptive field to efficiently capture multi-scale temporal features, thereby significantly improving the discriminability of features and enhancing the accuracy and flexibility of the model in different action scenarios;
[0110] The present invention proposes an adaptive interaction feature fusion sub-module (Adaptive Interaction Feature Fusion, AIFF), which dynamically adjusts the feature weights through an attention allocation mechanism, realizes the efficient balance and fusion of short-term and long-term temporal dependencies, and thus significantly improves the overall performance of the temporal action localization task;
[0111] Extensive experiments of the present invention on multiple TAL benchmark datasets show that the proposed method can efficiently achieve the current state-of-the-art performance, verifying its strong ability and effectiveness in capturing different temporal patterns in untrimmed videos.
[0112] Figure 3 It is a block diagram of a temporal action localization device 300 based on bidirectional interaction and dynamic feature enhancement shown according to an exemplary embodiment. The device 300 is used for the temporal action localization method based on bidirectional interaction and dynamic feature enhancement. Refer to Figure 3 In this, the device includes a signal initial feature extraction module 310, a feature enhancement strategy construction module 320, a temporal feature enhancement module 330, and a prediction module 340. Among them:
[0113] The initial feature extraction module 310 is used to obtain an uncropped video, process the uncropped video through a pre-trained network, and extract initial features;
[0114] The feature enhancement strategy construction module 320 is used to construct an adaptive feature enhancement strategy, and perform hierarchical network architecture integration dynamic modeling and local and global temporal interaction modeling through the adaptive feature enhancement strategy; the adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module and a global and local adaptive bidirectional interaction module;
[0115] The temporal feature enhancement module 330 is used to introduce the adaptive feature enhancement strategy into the encoder; the encoder is used to perform feature enhancement on the initial features to generate enhanced temporal features;
[0116] The prediction module 340 is used to transfer the enhanced temporal features to the classification and regression heads to perform prediction of action categories and time boundaries.
[0117] Optionally, the multi-scale dynamic temporal modeling module includes:
[0118] The multi-scale dynamic temporal modeling module includes three dynamic local DynL affines; through the dynamic adjustment of the receptive field of the multi-scale dynamic temporal modeling module, multi-scale temporal features are captured.
[0119] Optionally, through the dynamic adjustment of the receptive field of the multi-scale dynamic temporal modeling module, multi-scale temporal features are captured, including:
[0120] Through the dynamic local affine transformation DynL affine, the dynamic changes of actions at different time scales in the TAL task are alleviated;
[0121] Given the input features, the DynL module processes the input features using two different-scale deep convolutional branches to obtain multi-scale temporal information;
[0122] Through the hierarchical convolution method, the first branch performs deep convolution with a smaller kernel to capture local temporal information; the second branch first convolves with kernel k1 and then convolves with a larger kernel k2 to gradually expand the receptive field to capture a wider context;
[0123] The outputs of the two branches are fused to obtain a multi-scale sequence feature representation;
[0124] A learnable mask generation module is constructed, and the multi-scale sequence feature representation is input into the mask generation module to generate a mask sequence, which assigns different weights to each position in the sequence, highlighting the features of key frames while suppressing background information; a convolutional layer is used to extract local patterns from the sequence, and then refined through batch normalization BN to enhance the expressiveness; the mask sequence is multiplied element-wise with the multi-scale sequence features to generate weighted sequence features weighted.
[0125] Optionally, the global and local adaptive bidirectional interaction module includes:
[0126] The dynamic local affine block DynL affine and the global Global affine are combined through two-way interaction of the adaptive interaction feature fusion sub-module AIFF to form a global and local adaptive two-way interaction module;
[0127] The global features are aggregated through the global affine module; the local features are aggregated through the dynamic local affine block DynL affine; the global and local information is fused through the adaptive interaction feature fusion sub-module AIFF.
[0128] Optionally, aggregating the global features through the global affine module includes:
[0129] Incorporating the self-attention mechanism into the global affine module to obtain attention-based global feature aggregation; aggregating the global features through the global affine module and enhancing the model's ability to identify long-term relationships within the sequence to obtain robust action boundary localization.
[0130] Optionally, fusing the global and local information through the adaptive interaction feature fusion sub-module AIFF includes:
[0131] Dynamically assigning weights through the adaptive interaction feature fusion sub-module for feature fusion to balance and fuse the long-term and short-term dependencies in the time series information.
[0132] Optionally, dynamically assigning weights through the adaptive interaction feature fusion sub-module for feature fusion to balance and fuse the long-term and short-term dependencies in the time series information includes:
[0133] Given two input feature tensors; the input feature tensors include: obtaining the output features of DynL affine and the output features of Global affine;
[0134] Connecting the two input feature tensors along the channel dimension through the adaptive interaction feature fusion sub-module to form a fused feature tensor, and the connected features will undergo a series of transformations to generate dynamic weights. Temporal average pooling is used to reduce the temporal dimension;
[0135] Extracting key context information from the merged features through a convolutional layer;
[0136] Applying the sigmoid activation function to generate the adaptive weights of the local and global features;
[0137] Refining the obtained weights through one-dimensional convolution to reduce the dimension and obtain the final weights; weighting the obtained weights and the original features to obtain the final output.
[0138] In the embodiments of the present invention, a novel adaptive temporal enhancement framework is proposed, aiming to solve the complex temporal dependency problem in the temporal action localization (TAL) task through two-way dynamic interaction and balance of short-term and long-term features;
[0139] The present invention designs a multi-scale dynamic time series modeling module (MS-DyTM). This module integrates a multi-scale adaptive convolution kernel selection and a learnable mask mechanism, which can dynamically adjust the receptive field to efficiently capture multi-scale time series features, thereby significantly improving the discriminability of features and enhancing the accuracy and flexibility of the model in different action scenarios.
[0140] The present invention proposes an adaptive interactive feature fusion sub-module, which dynamically adjusts feature weights through an attention allocation mechanism to achieve an efficient balance and fusion of short-term and long-term time series dependencies, thereby significantly improving the overall performance of the time series action localization task.
[0141] Extensive experiments of the present invention on multiple TAL benchmark datasets show that the proposed method can efficiently achieve the current state-of-the-art performance, verifying its powerful ability and effectiveness in capturing different time series patterns in untrimmed videos.
[0142] Figure 4 It is a schematic structural diagram of a time series action localization device based on bidirectional interaction and dynamic feature enhancement provided by an embodiment of the present invention. As Figure 4 shown, the time series action localization device based on bidirectional interaction and dynamic feature enhancement may include the above Figure 3 shown time series action localization device based on bidirectional interaction and dynamic feature enhancement. Optionally, the time series action localization device 410 based on bidirectional interaction and dynamic feature enhancement may include a first processor 2001.
[0143] Optionally, the time series action localization device 410 based on bidirectional interaction and dynamic feature enhancement may further include a memory 2002 and a transceiver 2003.
[0144] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.
[0145] Next, in combination with Figure 4 each component of the time series action localization device 410 based on bidirectional interaction and dynamic feature enhancement will be specifically introduced:
[0146] Among them, the first processor 2001 is the control center of the temporal action localization device 410 based on two-way interaction and dynamic feature enhancement, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0147] Optionally, the first processor 2001 can execute various functions of the temporal action localization device 410 based on two-way interaction and dynamic feature enhancement by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0148] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 shown in
[0149] In a specific implementation, as an embodiment, the temporal action localization device 410 based on two-way interaction and dynamic feature enhancement may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in
[0150] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0151] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through an interface circuit ( Figure 4 not shown in) of the timing action positioning device 410 based on two-way interaction and enhanced dynamic characteristics. The embodiments of the present invention do not make specific limitations on this.
[0152] The transceiver 2003 is used to communicate with a network device or communicate with a terminal device.
[0153] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0154] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through an interface circuit ( Figure 4 not shown in) of the timing action positioning device 410 based on two-way interaction and enhanced dynamic characteristics. The embodiments of the present invention do not make specific limitations on this.
[0155] It should be noted that Figure 4 the structure of the timing action positioning device 410 based on two-way interaction and enhanced dynamic characteristics shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0156] In addition, the technical effects of the timing action positioning device 410 based on two-way interaction and enhanced dynamic characteristics may refer to the technical effects of the timing action positioning method based on two-way interaction and enhanced dynamic characteristics described in the above method embodiments, and will not be elaborated here.
[0157] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0158] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0159] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable sensors. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0160] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0161] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0162] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0163] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0165] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0166] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A temporal action localization method based on two-way interaction and dynamic feature enhancement, characterized in that: The method comprises: S1. Obtain an uncropped video, process the uncropped video through a pre-trained network, and extract initial features; S2. Construct an adaptive feature enhancement strategy, and perform hierarchical network architecture integrated dynamic modeling and local and global temporal interaction modeling through the adaptive feature enhancement strategy; the adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module and a global and local adaptive bidirectional interaction module; S3, introducing the adaptive feature enhancement strategy into the encoder; performing feature enhancement on the initial features through the encoder to generate enhanced time series features; S4. Pass the enhanced temporal features to the classification and regression heads to predict action categories and temporal boundaries.
2. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 1 is characterized in that: In said S2, the multi-scale dynamic time series modeling module comprises: The multi-scale dynamic time series modeling module includes three dynamic local DynL affines; and multi-scale time features are captured by dynamically adjusting the receptive domain of the multi-scale dynamic time series modeling module.
3. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 2 is characterized in that: The method captures multi-scale temporal features by dynamically adjusting the receptive domain of the multi-scale dynamic time series modeling module, including: The dynamic changes of actions at different time scales in the TAL task are alleviated through the dynamic local affine transformation DynL affine; Given the input features, DynL affine uses two deep convolution branches of different scales to process the input features and obtain multi-scale temporal information; Through a layered convolutional approach, the first branch performs depthwise convolution with a smaller kernel to capture local temporal information; the second branch first performs convolution with kernel k1 and then with a larger kernel k2, gradually expanding the receptive field to capture a wider context; The outputs of the two branches are fused to obtain multi-scale sequence feature representation; Constructing a learnable mask generation module, inputting the multi-scale sequence feature representation into the mask generation module, generating a mask sequence, assigning different weights to each position in the sequence, and highlighting the features of key frames while suppressing background information; The convolutional layer is used to extract local patterns from the sequence, which are refined by batch normalization (BN) and the weights are compressed to the range of [0, 1] using the S-type activation function to obtain the final mask sequence. The mask sequence is element-wise multiplied with the multi-scale sequence features to generate weighted sequence features.
4. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 3 is characterized in that: In S2, the global and local adaptive two-way interaction module includes: The dynamic local affine block DynL affine and the global affine Global affine interact bidirectionally through the adaptive interactive feature fusion submodule AIFF to form a global and local adaptive bidirectional interactive module; The global features are aggregated by the global Global affine; the local features are aggregated by the dynamic local affine block DynL; and the global and local information are fused by the adaptive interactive feature fusion submodule AIFF.
5. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 4 is characterized in that: The aggregating global features through the global affine module includes: The self-attention mechanism is incorporated into the global affine module to obtain attention-based global feature aggregation; global features are aggregated through the global affine module, and the ability of the model to recognize long-term relations in a sequence is enhanced to obtain robust action boundary positioning.
6. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 5 is characterized in that: The method of fusing global and local information through the adaptive interactive feature fusion submodule AIFF includes: The adaptive interactive feature fusion submodule dynamically assigns weights to perform feature fusion, balance and fuse long-term and short-term dependencies in time series information.
7. The temporal action positioning method based on two-way interaction and dynamic feature enhancement according to claim 6 is characterized in that: The method of dynamically assigning weights through the adaptive interactive feature fusion submodule to perform feature fusion, balance and fuse long-term and short-term dependencies in time series information, includes: Given two input feature tensors; the input feature tensors include: obtaining output features of DynL affine and output features of Global affine; The adaptive interactive feature fusion submodule connects two input feature tensors along the channel dimension to form a fused feature tensor. The connected features undergo a series of transformations to generate dynamic weights. Time average pooling is used to reduce the time dimension. Pass the merged features through the convolutional layer to extract key contextual information; Apply sigmoid activation function to generate adaptive weights of local and global features; The weights obtained by one-dimensional convolution are refined to reduce the dimension and obtain the final weights; the obtained weights are weighted with the original features to obtain the final output.
8. A temporal action positioning device based on two-way interaction and dynamic feature enhancement, the temporal action positioning device based on two-way interaction and dynamic feature enhancement is used to implement the temporal action positioning method based on two-way interaction and dynamic feature enhancement as claimed in any one of claims 1 to 7, characterized in that: The device comprises: An initial feature extraction module, used to obtain an uncropped video, process the uncropped video through a pre-trained network, and extract initial features; A feature enhancement strategy building module is used to build an adaptive feature enhancement strategy, through which hierarchical network architecture integrated dynamic modeling and local and global temporal interaction modeling are performed; the adaptive feature enhancement strategy includes: a multi-scale dynamic temporal modeling module and a global and local adaptive two-way interaction module; A time series feature enhancement module, used for introducing the adaptive feature enhancement strategy into the encoder; performing feature enhancement on the initial feature through the encoder to generate enhanced time series features; The prediction module is used to pass the enhanced temporal features to the classification and regression heads to predict action categories and temporal boundaries.
9. A time sequence action positioning device based on two-way interaction and dynamic feature enhancement, the time sequence action positioning device based on two-way interaction and dynamic feature enhancement comprising: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, any one of the temporal action localization methods based on two-way interaction and dynamic feature enhancement as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any of the temporal action localization methods based on two-way interaction and dynamic feature enhancement as described in any of claims 1-7.
Citation Information
Patent Citations
Video interaction action detection method based on multi-modal time perception and attention
CN114842559A
Self-adaptive perception video time sequence action positioning system and method thereof
CN116052034A
Timing sequence action nomination generation method and system based on coarse time granularity
CN117292307A
Time sequence action detection method and device based on potential action interval feature integration
CN118053107A
Video time sequence action positioning method, system and equipment based on proxy attention and multi-scale Transform and medium
CN118351475A
Cited By
Intestinal capsule endoscopy video ulcer fragment automatic positioning method, computer equipment and system
CN121616808A
A long short-term parallel double-branch network-based timing behavior detection method
CN122510971A