Time sequence action positioning method and system based on bidirectional clue enhancement

By introducing a bidirectional clue enhancement module in timing action positioning, using bidirectional feature extraction and enhancement technology, the problems of inaccurate action boundary positioning and loss of action dependence are solved, and higher positioning accuracy and model performance are achieved.

CN119992410APending Publication Date: 2025-05-13UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510004451.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing timing action positioning methods have insufficient accuracy in action boundary positioning, and the model is prone to lose action dependence in long video sequences, resulting in inaccurate positioning.

Method used

The timing action positioning method based on bidirectional clue enhancement is adopted to extract and enhance video features from both directions through bidirectional multi-head self-attention branch, bidirectional state space branch and bidirectional convolution branch to capture action details and long-term dependency information.

Benefits of technology

The accuracy of action boundary positioning and the model's ability to retain action-dependent memory is improved, and the performance of timing action positioning is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992410A_ABST
    Figure CN119992410A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence action positioning method and system based on bidirectional clue enhancement, and belongs to the technical field of computer vision, and the method comprises the steps: obtaining to-be-processed video data; inputting the acquired video data into a time sequence action positioning model; wherein the time sequence action positioning model comprises a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used for extracting video features corresponding to the video data; the bidirectional clue enhancement module is used for performing feature extraction and enhancement on the video features extracted by the video feature extraction module by adopting a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used for completing classification and positioning of action instances according to the enhanced features; and positioning and classifying actions in an input video by using the time sequence action positioning model. According to the scheme, the action positioning accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a temporal action positioning method and system based on bidirectional cue enhancement. Background Art

[0002] Temporal Action Localization (TAL) is one of the important research directions in the field of computer vision. Its task is to locate and classify the actions in a given input video, that is, to locate the start and end time of the action in the video and predict the corresponding action category. Including video analysis, intelligent monitoring and film editing, it has promoted the widespread application of video understanding technology in practical applications.

[0003] At present, the task of temporal action localization is mainly divided into two-stage and single-stage methods:

[0004] The two-stage method completes action detection through two key stages: the generation and classification of action candidate regions. In the first stage, the model first predicts candidate regions where actions may exist from the video, and then classifies the candidate action regions in the second stage, while correcting the boundaries of the candidate regions to obtain the final action category and start and end time predictions. Some mainstream two-stage methods include TSN, BSN, BMN, LGN, TCANet, etc. These methods have achieved remarkable achievements in the field of temporal action localization. However, the positioning accuracy of the two-stage method is heavily dependent on the action candidate regions generated in the first stage. At the same time, the two-stage method involves more parameters and complex structures, has high requirements for computing resources, and is difficult to deploy end-to-end.

[0005] The single-stage method does not have an independent process for generating candidate regions. It directly extracts video frames from the video for prediction. Compared with the two-stage method, it has faster calculation speed and lower model complexity. Some mainstream two-stage methods include SSAD, GTAN, ASFD, RCL, A2Net, ActionFormer, Tridet, etc. This method omits the additional candidate proposal generation step, simplifies the action localization process and computational complexity, and is more conducive to end-to-end deployment.

[0006] However, regardless of whether they adopt a two-stage or a single-stage approach, most methods use a simple one-way process to extract action features and classify and locate actions in videos. This simple one-way process often misses some action features and lacks the extraction of spatial information in video features, resulting in the offset of the localization boundary, which is particularly common in videos with multiple sub-actions. Therefore, achieving accurate boundary localization remains a challenging task.

[0007] For a long time, a core problem faced by TAL is the fuzziness of action boundaries, which leads to inaccurate boundary prediction. Although existing methods have tried to improve this problem in many aspects, most methods still rely on simple one-way designs to utilize the characteristics of video actions. This simple design often ignores some details that help the model to locate the action, making it particularly difficult to achieve high-precision positioning. The existing technology can be summarized as follows:

[0008] (1) Sliding window-based method: This method presets a set of sliding windows of different lengths and slides them step by step along the time dimension to determine the action categories within the sliding windows one by one, and at the same time counts the determination results of sliding windows of different lengths. This method can effectively capture action clips of different lengths and adapt to the diversity of action durations by presetting sliding windows of different lengths. At the same time, the dense sliding of the sliding window covers most of the timeline to ensure that the action clips are not missed. However, this method cannot handle video actions of different lengths. In addition, the fixed length of the sliding window may lead to inaccurate action boundaries, affecting positioning accuracy, and redundant calculations of non-action areas may lead to a high false detection rate.

[0009] (2) Boundary prediction-based method: This method transforms the positioning problem into a boundary detection problem. First, the boundary points of the video clip are identified at each time position, and the start and end probabilities of the action are predicted. Then, the start and end points with high prediction probabilities are directly combined to generate candidate segments of different lengths. Next, all candidate segments are evaluated and action segments with high confidence are selected as the final prediction results. This method can achieve more accurate action boundary positioning by predicting the start and end times of the action separately. It is not limited by the fixed-length sliding window and can flexibly adapt to actions of different lengths, thereby improving the adaptability to diverse action durations. However, when there are multiple overlapping actions in the video, the boundary prediction-based method may find it difficult to distinguish the boundaries of each action. Especially in complex scenes, when multiple actions are intertwined, the boundary prediction may not be accurate enough.

[0010] (3) Method based on temporal anchors: This method first defines a set of temporal anchors of fixed or variable length on the timeline of the video. These anchors represent time segments that may contain actions. Next, the model evaluates each anchor, predicts whether the time segment contains an action, and regresses the action category and the start and end time of the action. Finally, the model generates multiple candidate action segments based on the evaluation results, removes redundant prediction segments, and finally retains the prediction results with high confidence as the final results. This method can flexibly deal with actions of different time lengths, while avoiding the phenomenon of inaccurate prediction boundaries caused by the existence of multiple overlapping actions in the boundary prediction problem. However, this method has great problems in locating shorter action segments, and the anchors may not be located in shorter action durations.

[0011] In summary, in TAL, since it processes one-dimensional time series data that changes over time, the information available is relatively simple compared to traditional image or video processing tasks. Specifically, temporal action localization cannot rely on rich spatial structural information like static image analysis, and thus lacks spatial performance capabilities. Most of the existing temporal action localization algorithms use unidirectional action feature processing, which does not fully utilize spatial information, thereby ignoring possible action details. In this case, it is likely to cause inaccurate model positioning boundaries. In addition, long-term dependency in temporal modeling is also a big difficulty. Existing methods are prone to losing relevant action dependencies when modeling long sequences. This is reflected in the fact that in a long video, the model will gradually lose its learning memory of previous actions as it learns. Summary of the invention

[0012] The present invention provides a temporal action localization method based on bidirectional cue enhancement to solve the technical problems that the existing method has inaccurate localization boundaries and the model gradually loses the learning memory of previous actions as it learns.

[0013] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0014] On the one hand, the present invention provides a temporal action localization method based on bidirectional cue enhancement, comprising:

[0015] Obtaining video data to be processed;

[0016] The acquired video data is input into a temporal action localization model; wherein the temporal action localization model comprises: a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used to extract video features corresponding to the video data; the bidirectional clue enhancement module is used to extract and enhance the video features extracted by the video feature extraction module using a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used to complete the classification and positioning of the action instance according to the enhanced features;

[0017] The temporal action localization model is used to locate and classify actions in an input video.

[0018] Furthermore, the video feature extraction module includes: a pre-trained backbone network and a linear layer; wherein, the video data is input into the pre-trained backbone network, and the pre-trained backbone network extracts video features, and the extracted video features are sent to the linear layer, and the linear layer enlarges the video feature dimension through a linear mapping, and the enlarged video features are split into two parts, respectively serving as forward video features and backward video features.

[0019] Furthermore, the bidirectional clue enhancement module includes: a bidirectional multi-head self-attention branch, a bidirectional state space branch and a bidirectional convolution branch; wherein the inputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch are all forward video features and backward video features output by the video feature extraction module;

[0020] The bidirectional multi-head self-attention branch is used to capture global context information, so that the model focuses on the action area features and ignores irrelevant background feature information;

[0021] The bidirectional state space branch is used to solve the dependency problem existing in long video sequences;

[0022] The bidirectional convolution branch is used to capture local adjacent action features in the video and enhance action understanding.

[0023] Furthermore, the action detection head is specifically used to: fuse the outputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch to obtain fused features; and complete the classification and positioning of the action instance according to the fused features.

[0024] On the other hand, the present invention also provides a temporal action localization system based on bidirectional cue enhancement, comprising:

[0025] A data acquisition module, used for acquiring video data to be processed;

[0026] Data processing module for:

[0027] The acquired video data is input into a temporal action localization model; wherein the temporal action localization model comprises: a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used to extract video features corresponding to the video data; the bidirectional clue enhancement module is used to extract and enhance the video features extracted by the video feature extraction module using a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used to complete the classification and positioning of the action instance according to the enhanced features;

[0028] The temporal action localization model is used to locate and classify actions in an input video.

[0029] Furthermore, the video feature extraction module includes: a pre-trained backbone network and a linear layer; wherein, the video data is input into the pre-trained backbone network, and the pre-trained backbone network extracts video features, and the extracted video features are sent to the linear layer, and the linear layer enlarges the video feature dimension through a linear mapping, and the enlarged video features are split into two parts, respectively serving as forward video features and backward video features.

[0030] Furthermore, the bidirectional clue enhancement module includes: a bidirectional multi-head self-attention branch, a bidirectional state space branch and a bidirectional convolution branch; wherein the inputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch are all forward video features and backward video features output by the video feature extraction module;

[0031] The bidirectional multi-head self-attention branch is used to capture global context information, so that the model focuses on the action area features and ignores irrelevant background feature information;

[0032] The bidirectional state space branch is used to solve the dependency problem existing in long video sequences;

[0033] The bidirectional convolution branch is used to capture local adjacent action features in the video and enhance action understanding.

[0034] Furthermore, the action detection head is specifically used to: fuse the outputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch to obtain fused features; and complete the classification and positioning of the action instance according to the fused features.

[0035] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0036] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the above method.

[0037] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0038] 1. The present invention introduces Mamba into temporal action localization, alleviates the modeling constraints of convolutional neural networks through global receptive field and dynamic weighting, and solves the long dependency existing in long video sequences;

[0039] 2. The present invention designs a bidirectional feature processing scheme to capture spatial information in time series and utilize video features from two directions at the same time. The bidirectional feature processing scheme enhances the model's extraction of action details, thereby alleviating the problem of insufficient action feature learning in the video;

[0040] 3. The present invention introduces a hybrid bidirectional cue enhancement module, which strengthens temporal modeling from two directions through three branches and captures feature long-term dependency information, thereby further improving the accuracy of action localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0042] Figure 1 It is an overall framework diagram of the temporal action positioning model provided by an embodiment of the present invention;

[0043] Figure 2 It is a comparison chart of the positioning effects of the temporal action positioning method based on bidirectional cue enhancement provided by an embodiment of the present invention and the existing ActionMamba method on the THUMOS14 dataset;

[0044] Figure 3 It is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0046] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0047] First embodiment

[0048] This embodiment provides a method for temporal action localization based on bidirectional cue enhancement, with the aim of establishing a single-stage temporal action detection solution. In view of the fact that existing methods have difficulty in solving the long dependencies existing in long video sequences and in capturing the correlations in sequence features, we introduce Mamba's state space model into temporal action localization, and alleviate the modeling constraints of convolutional neural networks through global receptive fields and dynamic weighting. At the same time, in view of the fact that most of the current methods do not perform a one-way processing process on video features and lack the capture of spatial information due to the one-dimensional characteristics of video sequences, resulting in the model ignoring the details of the action in the video, we design a bidirectional feature learning mechanism to extract spatial feature information, perform a bidirectional enhancement processing operation on video features, capture clues in the action, and thus help the model to better locate and classify. The method can be implemented by an electronic device (terminal or server), and the execution process of the method includes:

[0049] S1, obtaining video data to be processed;

[0050] S2, input the acquired video data into the temporal action localization model;

[0051] S3, using the temporal action localization model to locate and classify actions in the input video.

[0052] The overall architecture of the temporal action localization model is as follows: Figure 1 As shown, it includes three parts: a video feature extraction module, a bidirectional cue enhancement module (Bidirectional Cue Enhancement, BCE) and an action detection head; the video feature extraction module includes: a pre-trained backbone network and a linear layer.

[0053] The workflow of the temporal action localization model can be described as follows: first, the video is extracted through a pre-trained backbone network; then, the extracted video features are sent to the linear layer, and the video feature dimensions are enlarged through a simple linear mapping. The enlarged video features are divided into two parts, which are used for forward and backward feature processing in the subsequent process; then, the forward and backward video features are sent to the BCE module for feature extraction and enhancement; finally, the enhanced video features are sent to the action detection head composed of convolutional layers and linear layers to complete the classification and positioning of action instances. Among them, the pre-trained backbone network can use feature extraction networks such as I3D, InterVideo, SlowFast, etc. that have been trained on large-scale data sets; the action detection head can use action classification and localization heads such as ActionFormer, TriDet, etc.

[0054] The BCE module uses a bidirectional feature extraction mechanism to extract video features to capture the missing details that may exist in unidirectional extraction, solve the problems of insufficient action detail extraction, lack of spatial information, and long sequence dependencies, and thus improve the detection performance of the model. First, it encodes the video features and, like most temporal action localization algorithms, downsamples the video features through maximum pooling to obtain a feature pyramid to form a multi-scale feature representation. Specifically, Figure 1 As shown in the figure, the module includes: Bidirectional Multi-Head Self-Attention (Bi-MHSA), Bidirectional State Space Model (Bi-SSM) and Bidirectional Convolution (Bi-CONV). The above three branches are used to refine and extract action features from three dimensions. Among them, the Bi-MHSA branch is mainly used to capture global context information, so that the model focuses on the action area features and ignores irrelevant background feature information; the Bi-SSM branch is mainly used to solve the dependency problem in long video sequences, that is, the model will ignore distant feature information; the Bi-CONV branch is mainly used to capture local adjacent action features in the video and enhance action understanding. The following is a detailed explanation.

[0055] Bidirectional Multi-Head Self-Attention Branch (Bi-MHSA): In order to more effectively obtain contextual information in video sequences, so that the model can focus more on important areas with discrimination and ignore some irrelevant background information, we introduce the multi-head self-attention mechanism in the Transformer model to accurately capture features. Specifically, the input forward and reverse features are first mapped through a linear layer to obtain the Q, K, and V matrices. Then the similarity between Q and K is calculated and multiplied with the V matrix to obtain the attention score. Then the results of each head are concatenated and mapped back to the original dimension to obtain the forward and reverse outputs. Finally, the reverse features are flipped and concatenated with the forward features to obtain the final output. Using multiple parallel attention heads, information representation can be learned from different subspaces, which enhances the model's ability to focus on key features and improves its ability to capture global dependencies. Different from the original multi-head self-attention mechanism, here we introduce bidirectional features and use SiLU, a simple activation function, to achieve the effect of gating, thereby enhancing and weakening video features. Through the Bi-MHSA module, the model can more comprehensively understand the internal structure of the input data, thereby improving the model's positioning accuracy for action clips.

[0056] Bidirectional State Space Branch (Bi-SSM): In order to overcome the dependency problem of the model in long sequence modeling and enable the model to fully grasp the relationship between feature sequences, we use state space equations to model the video feature sequence. Inspired by the excellent performance of Mamba in long sequences, we introduce the SSM module in Mamba for dependency modeling. However, since the extracted video sequence is a one-dimensional sequence, it lacks spatial feature information. Therefore, we divide the feature sequence into two groups of forward and reverse symmetry to simulate the extraction of spatial information on the image. Specifically, for the two groups of forward and reverse video features, firstly, feature extraction is performed through forward and backward convolution respectively, and then they are sent to the forward and backward SSM modules for sequence dependency modeling respectively. Then, the SiLU activation function is used to achieve the gating effect to enhance and suppress the features. Finally, the reverse features are flipped and spliced ​​with the forward features to obtain the final output. Through Bi-SSM, spatial information is enhanced and supplemented, which helps the model better understand the content of video actions.

[0057] Bidirectional convolution branch (Bi-CONV): After obtaining the global and semantic features extracted from Bi-MHSA and Bi-SSM, in order to further enhance the model's ability to capture fine-grained local information in video sequences, we introduced basic convolution operations to extract local features on video sequences. Specifically, we first linearly map the forward and reverse video features to unify the video feature dimensions, then use convolution operations to capture fine-grained information in video sequences, then use SiLU, an activation function, to achieve the effect of gating to enhance the use of video features, and finally flip the reverse video features and concatenate them with the forward video features to obtain the final output. By limiting the size of the convolution kernel in Bi-CONV, the convolution operation can focus on only a small area of ​​the video sequence each time, thereby effectively capturing the changing pattern of local information. This approach makes up for the problem of local details that may be ignored when Bi-MHSA and Bi-SSM focus on global features, while further enhancing the diversity and expression of features, enabling the model to better capture the dynamic changes of details in the sequence, and provide richer and more accurate feature information for subsequent feature fusion and task decision-making.

[0058] Based on the above, the implementation process of the scheme of the present invention is as follows: the features of the pre-trained model are expanded to a higher dimension so as to be divided into forward and backward features; the forward and backward features are sent to the BCE composed of a bidirectional multi-head self-attention branch, a bidirectional state space branch, and a bidirectional convolution branch for feature information enhancement. The features are downsampled using a maximum pooling with a step size of 2 to obtain a feature pyramid to capture the action content of the fixed dimension.

[0059] The training and inference process of the above model is as follows:

[0060] Model training: During the training phase, the model determines the action type and start and end time of all action clips appearing in the video, and uses tIoU to calculate the intersection over union (IoU) of the predicted interval and the true interval.

[0061] Model reasoning: In the reasoning stage, the action scores above the set threshold are selected as the final predicted action intervals, and Soft-NMS is used to eliminate the prediction of repeated action instances.

[0062] Next, the superiority of the method of the present invention is verified by means of comparative experiments.

[0063] Figure 2 The positioning effect of the method of the present invention and the existing Mamba-based temporal action positioning method ActionMamba on the THUMOS14 dataset is demonstrated. The results show that the method proposed in the present invention has higher positioning accuracy than ActionMamba, demonstrating the reliability of the method of the present invention.

[0064] Table 1 provides the comparison results between the performance of the proposed method and the most advanced methods on the THUMOS14 dataset. Compared with other methods that directly generate target intervals, the bidirectional cue enhancement model of the proposed method achieves an average mean average precision (mAP) of 73.7%, surpassing all previous methods. The experimental results verify the effectiveness of the bidirectional feature design of the proposed method and prove its effectiveness in extracting feature details.

[0065] Table 1 Performance of the proposed method (BCE) on the THUMOS14 dataset

[0066]

[0067] Table 2 shows that for the HACS dataset, the proposed method outperforms all previous methods, with an average mAP of 45.0%. These results highlight the excellent performance of the proposed method on this challenging dataset.

[0068] Table 2 Performance of the proposed method (BCE) on the HACS dataset

[0069]

[0070] Table 3 shows that on the ActivityNet-1.3 dataset, the proposed method achieved an average mAP of 42.1%, ranking first, and also achieved improvements in different tIoUs.

[0071] Table 3 Performance of the proposed method BCE on the ActivityNet-1.3 dataset

[0072]

[0073] In summary, extensive experiments on THUMOS14, HACS and ActivityNet-1.3 datasets show that the proposed temporal action localization method based on bidirectional cue enhancement outperforms the existing methods. The bidirectional feature algorithm proposed in this paper provides a feasible solution for action boundary generation and provides new directions and ideas for further research on temporal action localization.

[0074] Second embodiment

[0075] This embodiment provides a temporal action localization system based on bidirectional cue enhancement, the system comprising:

[0076] A data acquisition module, used for acquiring video data to be processed;

[0077] Data processing module for:

[0078] The acquired video data is input into a temporal action localization model; wherein the temporal action localization model comprises: a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used to extract video features corresponding to the video data; the bidirectional clue enhancement module is used to extract and enhance the video features extracted by the video feature extraction module using a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used to complete the classification and positioning of the action instance according to the enhanced features;

[0079] The temporal action localization model is used to locate and classify actions in an input video.

[0080] Among them, it should be noted that the temporal action positioning system based on two-way clue enhancement of the present embodiment corresponds to the temporal action positioning method based on two-way clue enhancement mentioned above; wherein, the functions implemented by each functional module in the temporal action positioning system based on two-way clue enhancement of the present embodiment correspond one by one to each process step in the temporal action positioning method based on two-way clue enhancement mentioned above; therefore, they will not be repeated here.

[0081] Third embodiment

[0082] This embodiment provides an electronic device, such as Figure 3 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment. In addition, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0083] Next, combine Figure 3 The following is a detailed introduction to the various components of the electronic device:

[0084] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor may execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0085] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 3 The CPU0 and CPU1 shown in the figure are, of course, only exemplary.

[0086] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0087] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 3 The processor is coupled to the processor (not shown), which is not specifically limited in this embodiment of the present invention.

[0088] The transceiver may include a receiver and a transmitter ( Figure 3 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 3 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0089] In addition, it should be noted that Figure 3 The structure of the electronic device shown in the figure does not constitute a limitation on the device, and the actual device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they are not repeated here.

[0090] Fourth embodiment

[0091] This embodiment provides a computer-readable storage medium, which stores at least one instruction, and the instruction is loaded and executed by a processor to implement the method of the first embodiment. The computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein may be loaded by a processor in a terminal to execute the method.

[0092] In addition, it should be noted that the present invention can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present invention can be in the form of a full or partial hardware embodiment, a full or partial software embodiment or an embodiment combining software and hardware. Moreover, when implemented using software, the embodiment of the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more available media sets. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.

[0093] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0094] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0095] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. In addition, the term "and / or" is only an association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or plural.

[0096] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0097] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0098] In several embodiments provided by the present invention, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of functional modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0099] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0100] Finally, it should be noted that the above is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiment of the present invention has been described, for ordinary technicians in this technical field, once the basic creative concept of the present invention is known, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. Therefore, the attached claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A temporal action localization method based on bidirectional cue enhancement, characterized in that: include: Obtaining video data to be processed; The acquired video data is input into a temporal action localization model; wherein the temporal action localization model comprises: a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used to extract video features corresponding to the video data; the bidirectional clue enhancement module is used to extract and enhance the video features extracted by the video feature extraction module using a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used to complete the classification and positioning of the action instance according to the enhanced features; The temporal action localization model is used to locate and classify actions in an input video.

2. The temporal action localization method based on bidirectional cue enhancement as claimed in claim 1, characterized in that: The video feature extraction module includes: a pre-trained backbone network and a linear layer; wherein, video data is input into the pre-trained backbone network, and the pre-trained backbone network extracts video features, and the extracted video features are sent to the linear layer, and the linear layer enlarges the video feature dimension through a linear mapping, and the enlarged video features are split into two parts, respectively serving as forward video features and backward video features.

3. The temporal action localization method based on bidirectional cue enhancement as claimed in claim 2, characterized in that: The bidirectional clue enhancement module includes: a bidirectional multi-head self-attention branch, a bidirectional state space branch and a bidirectional convolution branch; wherein the inputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch are all forward video features and backward video features output by the video feature extraction module; The bidirectional multi-head self-attention branch is used to capture global context information, so that the model focuses on the action area features and ignores irrelevant background feature information; The bidirectional state space branch is used to solve the dependency problem existing in long video sequences; The bidirectional convolution branch is used to capture local adjacent action features in the video and enhance action understanding.

4. The method for temporal action localization based on bidirectional cue enhancement as claimed in claim 3, characterized in that: The action detection head is specifically used to: fuse the outputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch to obtain fused features; and complete the classification and positioning of the action instance according to the fused features.

5. A temporal action localization system based on bidirectional cue enhancement, characterized in that: include: A data acquisition module, used for acquiring video data to be processed; Data processing module for: The acquired video data is input into a temporal action localization model; wherein the temporal action localization model comprises: a video feature extraction module, a bidirectional clue enhancement module and an action detection head; the video feature extraction module is used to extract video features corresponding to the video data; the bidirectional clue enhancement module is used to extract and enhance the video features extracted by the video feature extraction module using a bidirectional feature extraction mechanism to obtain enhanced features; the action detection head is used to complete the classification and positioning of the action instance according to the enhanced features; The temporal action localization model is used to locate and classify actions in an input video.

6. The temporal action localization system based on bidirectional cue enhancement as claimed in claim 5, characterized in that: The video feature extraction module includes: a pre-trained backbone network and a linear layer; wherein, video data is input into the pre-trained backbone network, and the pre-trained backbone network extracts video features, and the extracted video features are sent to the linear layer, and the linear layer enlarges the video feature dimension through a linear mapping, and the enlarged video features are split into two parts, respectively serving as forward video features and backward video features.

7. The temporal action localization system based on bidirectional cue enhancement as claimed in claim 6, characterized in that: The bidirectional clue enhancement module includes: a bidirectional multi-head self-attention branch, a bidirectional state space branch and a bidirectional convolution branch; wherein the inputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch are all forward video features and backward video features output by the video feature extraction module; The bidirectional multi-head self-attention branch is used to capture global context information, so that the model focuses on the action area features and ignores irrelevant background feature information; The bidirectional state space branch is used to solve the dependency problem existing in long video sequences; The bidirectional convolution branch is used to capture local adjacent action features in the video and enhance action understanding.

8. The temporal action localization system based on bidirectional cue enhancement as claimed in claim 7, characterized in that: The action detection head is specifically used to: fuse the outputs of the bidirectional multi-head self-attention branch, the bidirectional state space branch and the bidirectional convolution branch to obtain a fusion feature; The action instances are classified and located based on the fused features.