A Weakly Supervised Video Temporal Action Localization Method Based on Action-Correlated Attention

Through the combination of the action-associated attention model and the Transformer architecture, the problem of insufficient long-term time segment relationship simulation in weak-supervised timing action positioning is solved, and high-precision action segment positioning and classification are achieved.

CN114898259BActive Publication Date: 2025-07-04BEIJING UNION UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210481400.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-07-04
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

The existing weak-supervised timing action positioning technology cannot simulate the relationship between long-term time segments, resulting in the inability to capture the dependence information of the subsequent actions on the previous actions when some actions are separated by other actions, resulting in large errors in the positioning of the corresponding timing action in the final predicted.

Method used

The action-associated attention model is adopted, and weakly supervised pre-training is established using the query mechanism, and the output of the query mechanism is input into the decoder of the Transformer architecture. The classification and positioning loss functions are combined for joint training. The relationship between the features of the video clips is determined through the Transformer architecture encoder, so as to achieve accurate positioning and classification of the action clips.

Benefits of technology

It improves the accuracy of video timing action positioning, can accurately capture the dependence information between long-term time segments, reduces prediction errors, and improves the classification and positioning accuracy of action segments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898259B_ABST
    Figure CN114898259B_ABST
Patent Text Reader

Abstract

The present application relates to a weakly supervised video temporal action localization method based on action correlation attention. An action correlation attention model is used to establish the relationship between action segments in a video, and then the localization and classification of action segments are realized. Among them, the action correlation attention model uses a query mechanism to establish weakly supervised pre-training, and inputs the output of the query mechanism into the decoder of the Transformer architecture to realize the temporal localization of the query set; the encoder of the Transformer architecture is used to determine the relationship between video segment features to realize the classification of action segments in the video. The present application realizes the temporal action localization of videos by using a weakly supervised method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video understanding in artificial intelligence, and in particular, to a weakly supervised video temporal action localization method based on action-correlated attention. Background Art

[0002] Temporal action localization (TAL) is a challenging task in video understanding and is widely used to quickly locate action segments in different time ranges, that is, to locate the start and end times of an action in a video and classify the action. In the prior art, temporal action localization is usually achieved under supervised or weakly supervised settings. For the supervised case, it is necessary to manually annotate the frame-level labels of each action and the start and end times of the action for the training video, which thus wastes a large amount of time. In contrast, the weakly supervised method only needs to annotate the video-level labels of the action, that is, the label indicating only whether the action is in the video, to classify the action and localize the time. Thus, this weakly supervised temporal action localization provides a labor-saving but more challenging solution.

[0003] In the absence of frame-level annotations, weakly supervised temporal action localization uses the similarity of the same action to determine its entire segment and uses the distinctiveness of different actions to classify the labels. Therefore, the two models of W-TALC and Autoloc use the co-activity similarity loss with feature similarity for localization and use the multi-instance learning loss with feature dissimilarity for classification. However, the above methods cannot simulate the relationship between long-term time segments. When some actions are separated by other actions, since the dependence information of the subsequent action on the previous action cannot be captured, the corresponding temporal action localization finally predicted has a large error. For example, "open the wardrobe" and "close the wardrobe" share information, but are separated by the long-term action of "fold the clothes" in the middle. Then, when predicting the temporal localization of the "close the wardrobe" action, the above method cannot capture the dependence on the information of "open the wardrobe", resulting in a large error in the finally predicted action temporal localization. Summary of the Invention

[0004] In order to solve the problem that the weakly supervised temporal action localization technology in the prior art cannot simulate the relationship between long-term time segments, when some actions are separated by other actions, since the dependence information of the subsequent action on the previous action cannot be captured, resulting in a large error in the corresponding temporal action localization finally predicted, this application provides a weakly supervised video temporal action localization method based on action-correlated attention.

[0005] In a first aspect, a weakly supervised video temporal action localization method based on action-correlated attention provided by this application adopts the following technical solution:

[0006] A weakly supervised video temporal action localization method based on action-correlated attention uses an action-correlated attention model to establish the relationship between action segments in a video, and then realizes the localization and classification of action segments. Among them, the action-correlated attention model uses a query mechanism to establish weakly supervised pre-training, and inputs the output of the query mechanism into the decoder of the Transformer architecture to realize the temporal localization of the query set; uses the encoder of the Transformer architecture to determine the relationship between video segment features, and realizes the classification of action segments in the video.

[0007] By adopting the above technical solutions, especially the action-correlated attention model solves the problem of no ground-truth supervision training in weakly supervised training by using the query mechanism to establish weakly supervised pre-training, and then inputs the output of the query mechanism into the decoder of the Transformer architecture to realize the temporal localization of the query set; at the same time, uses the encoder of the Transformer architecture to determine the relationship between video segment features, realizes the classification of action segments in the video, and finally realizes the temporal action localization of the video by using the weakly supervised method. Therefore, for the situation where some actions are separated by other actions, it can also accurately localize the corresponding action segments through the action segment classification of the present application, capture the dependency information of the subsequent action on the previous action, and make the final predicted temporal action localization accuracy relatively high.

[0008] Preferably, the action-correlated attention model uses a query mechanism to establish weakly supervised pre-training, specifically including:

[0009] Establish a pre-training task, randomly crop M video segments S = {S1,..., S M}, and record their timestamps as ground truth; extract the features of the M video segments to obtain a query segment set Randomly generate N time regions including start and end timestamps, where N is much larger than M;

[0010] Encode each time region as an action query Then the action query set containing N action queries is

[0011] Divide the query set Q u equally among the feature set F S , that is, N / M action queries correspond to one Obtain a query set with a corresponding relationship

[0012] Input the query set with the corresponding relationship into the decoder of the Transformer architecture, and use the query segment set F with timestamps S, the learning of the supervised action query set (that is, for the query set Q of randomly generated time regions initially u , q i continuously adjusts its start and end timestamp positions to match the query segment set F with known timestamps S ), so that u there are M action query time regions in Q that correspond one-to-one to the timestamps recorded in F S (that is, continuously training so that the features corresponding to the action queries match the extracted features).

[0013] By adopting the above method, the weakly supervised pre-training model established by using the query mechanism is enabled to have the ability to locate the start and end timestamps of this feature by giving any feature. In particular, by randomly cropping out M video segments S = {S1,..., S M} and recording their timestamps as the ground truth for model training, a method of query shuffling is adopted to achieve the randomness of query allocation in the input decoder.

[0014] More preferably, the I3D* network with frozen parameters is used to extract the features of the M video segments. By using the query mechanism in combination with the I3D* network with frozen parameters to extract the features of the M video segments in this application, the different preferences for features between classification and localization can be effectively balanced, making the simultaneously obtained video temporal action localization and classification data more accurate.

[0015] Preferably, when allocating the query set Q u , the mask matrix is added to the attention layer of the decoder, that is, the attention mask matrix is used to control the interaction between different object queries; the attention mask is:

[0016]

[0017] where X i,j determines whether the action query interacts with the action query. By the above method, the independence of each action query q i can be satisfied, and the overall accuracy and stability of video temporal action localization and classification are further improved.

[0018] Preferably, when allocating the query set Q u , randomly shuffle the permutation of all action query encodings; and / or during pre-training, randomly mask 10% of the action query segments as zeros. Thus, the generalization ability of the model can be improved, and the universality of the model for different data sets can be solved.

[0019] Preferably, the action correlation attention model is trained by the following method:

[0020] Input the video containing actions as training data; preprocess the training data to obtain video frames and optical flow frames of the video, and extract the I3D features of the video segments;

[0021] Encode the video temporal information of the video segment into positional encoding;

[0022] Input the video temporal positional encoding and I3D features of the video segment into the encoder of the action-related attention model to determine the relationship between video segment features and achieve the classification of action segments; input the video temporal positional encoding of the video segment into the decoder of the action-related attention model, and at the same time use the query mechanism to establish weakly supervised pre-training, and input the output of the query mechanism into the decoder of the action-related attention model to achieve the temporal localization of the query set;

[0023] Adopt a classification loss function and a localization loss function to jointly train the action-related attention model, where the classification loss function is used to supervise the feature classification effect; the localization loss function is used to locate the start and end times of the given feature the video segment loss that measures proximity in;

[0024] Merge the query sets output by the encoder and decoder to obtain the localization and classification of action segments in the video.

[0025] Preferably, the action-related attention model is trained by a global matching loss algorithm and achieves a unique prediction through bipartite matching; specifically, adopt a classification loss function feature reconstruction loss function and a localization loss function for joint training; where the classification loss function is used to supervise the feature classification effect; the feature reconstruction loss is used to balance the different preferences of classification and localization for features; the localization loss function is used to locate the start and end times of the given feature the video segment loss that measures proximity in.

[0026] Preferably, assume is the set of action ground truths, where a (i) =(c (i) , s (i) , r (i) ), for the i-th action a (i) , c (i) is the action category, s(i) is the start and end time of the action, r (i) It is a characteristic of the action; is the action prediction set; N is much larger than M (so at least NM action queries are predicted as no action); the Hungarian algorithm is used as the matching loss function to calculate the prediction set and truth value a (i) The same loss between ; the Hungarian loss "L H Defined as:

[0027]

[0028] in, is a calculated by the optimal bipartite matching (i) The action number; when the predicted label is in the set is 1, otherwise it is 0; is for c (i) Probability of class prediction Loss function. In this application, the Hungarian algorithm is used as the matching loss function to calculate the prediction set and truth value a (i) The same loss between , can reduce non-maximum suppression, and further improve the accuracy of temporal action localization and classification.

[0029] Preferably, the positioning loss function Indicates the start and end time of locating a given feature The video segment loss that measures proximity is defined as the L1 loss between the prediction and the ground truth (sensitive to instance duration) and A weighted combination of the losses (invariant to instance duration), namely:

[0030]

[0031] in, λ iou and is a preset hyperparameter (which can be set according to conventional principles). Thus, the idea of ​​binary matching reduces non-maximum suppression and further improves the accuracy of temporal action localization and classification.

[0032] In a second aspect, the present application provides an electronic device that adopts the following technical solution:

[0033] An electronic device comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute any of the above methods.

[0034] In a third aspect, a computer-readable storage medium provided by the present application adopts the following technical solution:

[0035] A computer-readable storage medium stores a computer program that can be loaded and executed by a processor to perform any of the foregoing methods.

[0036] In summary, the present application includes at least one of the following beneficial technical effects:

[0037] The present application proposes a weakly supervised video temporal action localization model based on action-correlated attention (W-ART, i.e., the entire method of the present application) to establish the relationship between action segments with a long time interval, and finally obtain accurate video temporal action localization and classification data. Description of the Drawings

[0038] Figure 1 It is a flowchart of the training and testing method of the action-correlated attention model in the present application.

[0039] Figure 2 It is a schematic diagram of preprocessing data in the present application.

[0040] Figure 3 It is a flowchart of the codec module in the present application.

[0041] Figure 4 It is a schematic diagram of the core module of the encoder in the present application.

[0042] Figure 5 It is a schematic diagram of the core module of the decoder in the present application.

[0043] Figure 6 It is a schematic diagram of the training process of the action-correlated attention model in the present application. Detailed Embodiments

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will combine the Figure 1 — Figure 6 in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0045] When faced with the problem in the background art that the relationship between long-term time segments cannot be simulated, resulting in some actions being separated by other actions, and due to the inability to capture the dependency information of subsequent actions on previous actions, a large error exists in the corresponding temporal action localization predicted finally. The inventor has found through research that although 3D convolution and graph convolution are used to model the relationship of segments due to the advantages of local connectivity and translational invariance, these convolutions are usually only designed to capture short-range information and cannot capture long-distance dependency information beyond the convolution receptive field. Additionally, although D3d expands the receptive field, it cannot capture long-term dependencies by aggregating information at shorter distances. Furthermore, the separate temporal set recurrent network uses a recurrent neural network to capture the relationship of time segments, yet this method cannot ensure that all time segments are treated equally.

[0046] Therefore, the present application proposes a weakly supervised video temporal action localization method based on action association attention.

[0047] An embodiment of the present application discloses a weakly supervised video temporal action localization method based on action association attention. A weakly supervised video temporal action localization method based on action association attention uses an action association attention model to establish the relationship between action segments in a video, and further realizes the localization and classification of action segments. Among them, the action association attention model uses a query mechanism to establish weakly supervised pre-training, and inputs the output of the query mechanism into the decoder of the Transformer architecture to achieve the temporal localization of the query set; uses the encoder of the Transformer architecture to determine the relationship between video segment features and realizes the classification of action segments in the video.

[0048] Specifically, the action association attention model is trained by the following method, as Figure 1 shown:

[0049] S11, input a video containing actions as training data; preprocess the training data to obtain the video frames and optical flow frames of the video, and extract the I3D features of the video segments;

[0050] As Figure 2 shown, in specific implementation, in order to obtain the I3D feature F of a video V containing t frames V , t can be divided into segments of 8 frames with an overlap of 4 frames, thereby obtaining segments Since the size of T V varies with the duration of the video, here set to unify the dimensions, where T V1 represents the first video. can be composed of video frames and optical flow frames Composed of the sum of the feature cascades.

[0051] S12, encoding the video timing information of the video segment into positional encoding;

[0052] S13, inputting the video timing positional encoding and I3D features of the video segment into the encoder of the action correlation attention model to determine the relationship between video segment features and achieve the classification of action segments;

[0053] Specifically, the I3D feature F output by the data preprocessing of this application V Design a learnable positional encoding To retain the time position information of each segment. However, P does not carry time information at the beginning but is randomly initialized. P and F V After adding bit by bit, the input of the encoder is obtained. As Figure 3 Shown, the encoder-decoder contains an encoder and a decoder module. The encoder input Encoder output

[0054] Specifically, when implemented, the encoder in the Transformer architecture can be composed of L e Encoding blocks, and the l-th encoding block is denoted as E l , where l ∈ [1, L e . E l The input and output of are defined as And Therefore, as a cascaded structure, And Are equal.

[0055] The core of each encoding block is the multi-head self-attention (abbreviated as mh_s_attn) module, as Figure 4 Shown. It is composed of h single-head self-attention (abbreviated as sh_s_attn) modules in cascade. For the l-th encoding block, its input Is regarded as the query (Q), key (K), and value (V). For the i-th single-head self-attention module, its query, key, and value are defined as:

[0056]

[0057] Where Are the single-head transformation weights of the query, key, and value respectively. Obtain The formula for the i-th single-head self-attention of the l-th encoding block is:

[0058]

[0059] Among them, it is usually defaulted that After the normalization operation (softmax), the obtained attention map is Multiplied by V i to obtain The multi-head self-attention module of the l-th encoding block is defined as:

[0060]

[0061] Because each head is independent, the final result contains 1 / h, so After passing through the fully connected module (FFN) and the layer normalization module (add&normal), the dimensions remain unchanged.

[0062] S14, input the video temporal position encoding of the video segment into the decoder of the action-related attention model. At the same time, use the query mechanism to establish weakly supervised pre-training, and input the output of the query mechanism into the decoder of the action-related attention model to achieve the temporal localization of the query set;

[0063] Specifically, the decoder in the Transformer architecture can be composed of L d decoding blocks. The l-th decoding block is denoted as D l , where l ∈ [1, L d . The input and output of D l are defined as and Therefore, as a cascaded structure, and are equal. Similar to the encoder, its core is the multi-head self-attention (abbreviated as mh_s_attn) module and the multi-head cross-attention (abbreviated as mh_c_attn) module as Figure 5 shown.

[0064] Among them, the multi-head self-attention module is a cascade of h single-head self-attention (abbreviated as sh_s_attn) modules. The input of the l-th decoding block is which is regarded as the query (Q), key (K), and value (V). By analogy with the encoder, the formula for the multi-head self-attention of the decoder can be obtained as:

[0065]

[0066] where

[0067] The multi-head cross-attention module is a series connection of h single-head cross-attention (sh_c_attn) modules. Let serve as the query (Q) of the multi-head cross-attention; the output of the encoder is assigned to each decoding block of the decoder, that is, the input serves as its value (V); the positional encoding P and are added bit by bit to obtain which serves as the key (K). In the l-th decoding block, after being processed by the steps of formula (1), by analogy with formula (2), the definition of single-head cross-attention is obtained:

[0068]

[0069] After the normalization operation (softmax), the attention map is which is multiplied by to obtain The multi-head self-attention is a series connection of h single-head attentions, defined as:

[0070] mh_c_attn (l) = [sh_c_attn1;...; sh_c_attn h (6)

[0071] where

[0072] S15, the classification loss function and the localization loss function are used to jointly train the action-associated attention model. Among them, the classification loss function is used to supervise the feature classification effect; the localization loss function is used to localize the start and end times of the given feature and the video segment loss for measuring proximity in

[0073] S16, the query sets output by the encoder and the decoder are merged to obtain the localization and classification of the action segments in the video.

[0074] The testing process of this model is similar to the training process and consists of three parts (excluding joint training). Input a video with action labels, and the output is a set of action queries with timestamps and labels.

[0075] In the above method, the I3D* network with frozen parameters can be used to extract the features of the M video segments.

[0076] To balance the feature preferences of localization and classification and improve the accuracy of final video temporal action localization and classification, the action correlation attention model is trained through the global matching loss algorithm and achieves unique prediction through bipartite matching. Specifically, the classification loss function feature reconstruction loss function and the localization loss function are used for joint training. Among them, the classification loss function is used to supervise the feature classification effect; the feature reconstruction loss is used to balance the different preferences of classification and localization for features; the localization loss function is used to localize the start and end times of the given features in the video segment loss that measures proximity.

[0077] The training process is as shown in Figure 6 . The localization loss function is used to localize the start and end times of the given features in the video segment loss that measures proximity. In each round of training, the action query set is a set of random query segments F with ground truth S in F i to find the best-matched action query, which means that during the training process, the model gradually acquires the ability to localize the start and end time points of any given feature.

[0078] The features F of the query segment S and the final prediction P r are input into which retains the feature recognition for classification. It is used for classification and thus can be used for different videos. Usually, the model uses and to find the segments containing actions, and then uses the codec and to localize them. Finally, the adjacent temporal action queries with the same label are merged into one segment to obtain the final prediction P r . In the test, the unclipped video without labels is sent to the model to predict its action classification and localization.

[0079] Formally, assume is the set of action ground truth, where a (i) =(c (i) , s (i) , r (i) ). For the i-th action a (i) , c (i) is the action category, s (i) is the start and end time of the action occurrence, and r (i) is the feature of the action; is the action prediction set; N is much larger than M (so at least N - M action queries are predicted as no action); the Hungarian algorithm is used as the matching loss function to calculate the and the ground truth a (i) between the same loss; the Hungarian loss "L H of all matching pairs is defined as:

[0080]

[0081] where, is the action sequence number of a (i) calculated by the optimal bipartite matching; when the predicted label is in the set is 1, otherwise 0; is for the probability of class c (i) prediction loss function.

[0082] Video-level labels are used to classify actions. The model averages the top k of each class to obtain a c-dimensional video-level prediction. It is defined as follows:

[0083]

[0084] The positioning loss function represents the video segment loss that measures the proximity in locating the start and end times of the given feature. The segment loss is defined as a weighted combination of the L1 loss (sensitive to instance duration) between the prediction and the ground truth and the loss (invariant to instance duration), that is:

[0085]

[0086] where,

[0087]

[0088] λ iou and are preset hyperparameters (which can be set according to general principles).

[0089] is the mean squared error between two normalized segment features extracted by the I3D backbone, which is defined as

[0090] This model freezes the backbone I3D and proposes the feature reconstruction I3D* to retain the feature recognition for classification.

[0091] In the above method, the action correlation attention model uses a query mechanism to establish weakly supervised pre-training, specifically including:

[0092] S21, establish a pre-training task, randomly crop M video segments S = {S1,..., S M}, and record their timestamps as the ground truth;

[0093] S22, extract the features of the M video segments to obtain a query segment set

[0094] S23, randomly generate N time regions including start and end timestamps, where N is much larger than M;

[0095] S24, encode each time region into an action query Then the action query set containing N action queries is

[0096] S25, evenly distribute the query set Q u to the feature set F S , that is, N / M action queries correspond to one Obtain a query set with a corresponding relationship

[0097] S26, input the query set with the corresponding relationship into the decoder of the Transformer architecture, and use the query segment set F with timestamps S , to supervise the learning of the action query set (that is, for the query set Q u , q i continuously adjusts its start and end timestamp positions to match the query segment set F with known timestamps S ), so that there are M action queries in Q u whose time regions correspond one-to-one to the timestamps recorded in F S (that is, continuously train so that the features corresponding to the action queries match the extracted features). (In this way, the model has the ability to give any feature and locate the start and end timestamps of this feature)

[0098] To satisfy the independence of each action query q i , when distributing the query set Q u , add a mask matrix to the attention layer of the decoder, that is, use the attention mask matrix to control the interaction between different object queries; the attention mask is:

[0099]

[0100] Among them, X i,j Determine whether the action query interacts with the action query interaction.

[0101] In addition, in the action localization task, there is no explicit group assignment among the action queries. Therefore, in order to simulate the implicit group assignment among the action queries, the query set Q is assigned during pre-training u When, randomly shuffle the permutation of all action query encodings Figure 6 Shows pre-training of multi-query segments with attention masks and shuffled action query sets; in addition, in order to improve the generalization ability of the model and solve the generality of the model for different data sets, during pre-training, 10% of the action query segments can be randomly masked to zero.

[0102] The embodiments of the present application also disclose an electronic device. An electronic device includes a memory and a processor, and a computer program capable of being loaded and executed by the processor as any one of the foregoing methods is stored on the memory.

[0103] Among them, the electronic device can be a desktop computer, a laptop computer or a cloud server, etc. And the electronic device includes, but is not limited to, a processor and a memory. For example, the electronic device may further include input / output devices, network access devices, and a bus, etc.

[0104] The processor in the present application may include one or more processing cores. The processor runs or executes instructions, programs, code sets or instruction sets stored in the memory, calls data stored in the memory, and executes various functions of the present application and processes data. The processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above processor functions may also be others, and the embodiments of the present application do not make specific limitations.

[0105] Among them, the memory can be an internal storage unit of the electronic device, for example, the hard disk or memory of the electronic device, or can be an external storage device of the electronic device, for example, a plug-in hard disk, a smart memory card (SMC), a secure digital card (SD), or a flash memory card (FC) equipped on the electronic device, etc. Moreover, the memory can also be a combination of the internal storage unit and the external storage device of the electronic device. The memory is used to store computer programs and other programs and data required by the electronic device. The memory can also be used to temporarily store the data that has been output or will be output. This application does not make any restrictions on this.

[0106] The embodiments of the present application also disclose a computer-readable storage medium. A computer-readable storage medium stores a computer program that can be loaded and executed by a processor and is the same as any of the foregoing methods.

[0107] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0108] In order to verify the effect of the present application, the inventor also conducted the following comparative experiments:

[0109] To verify the performance of the action-related attention model of this application in localizing and classifying action segments in videos, the inventor compared the performance of the W-ART model (i.e., the action-related attention model of this application) with the SOTA (using the THUMOS14 dataset) method. The mean average precision (mAP) was used as the metric to evaluate the model, and specifically, the mean average precision at different intersection over union (mAP@tIoU) on t-IoU was used as the evaluation metric for the model (tIoU ∈ {0.3, 0.4, 0.5, 0.6, 0.7}, where ↑ indicates the higher the better). The comparison results are shown in Table 1:

[0110] Table 1

[0111]

[0112] As can be seen from Table 1: On the THUMOS14 dataset, the method model W-ART of this application achieved a 0.6% improvement in precision on UNT features and a 0.7% improvement in precision on I3D features compared to SOTA. That is, the action-related attention model of this application has better performance and higher precision in localizing and classifying action segments in videos; among them, SOTA is short for state of art, which means the highest precision in the model for this task.

[0113] In addition, when comparing with SOTA (Charades), the inventor calculated the mAP using the same settings in Table 2. The results are shown in Table 2:

[0114] Table 2

[0115]

[0116] Table 2 shows that: On the Charades dataset, W-ART (i.e., the technical solution of this application) achieved a 0.5% improvement in I3D features compared to SOTA. That is, the action-related attention model of this application has better performance and higher precision in localizing and classifying action segments in videos.

[0117] The above are all the preferred embodiments of this application, and the protection scope of this application is not limited by this. Therefore, all equivalent changes made according to the method and principle of this application should be covered within the protection scope of this application.

Claims

1. A weakly supervised video temporal action localization method based on action-correlated attention, characterized in that: An action-correlated attention model is adopted to establish the relationships between action segments in a video, and further achieve the localization and classification of action segments. Among them, for the action-correlated attention model, a weakly supervised pre-training is established using a query mechanism, and the output of the query mechanism is input into the decoder of the Transformer architecture to achieve the temporal localization of the query set. The encoder of the Transformer architecture is used to determine the relationships between video segment features to achieve the classification of action segments in the video. Among them, the action-associated attention model is trained by a global matching loss algorithm and achieves a unique prediction through bipartite matching; a classification loss function feature reconstruction loss function and a localization loss function are used for joint training; among them, the classification loss function is used to supervise the feature classification effect; the feature reconstruction loss is used to balance the different preferences of classification and localization for features; the localization loss function is used to localize the start and end times of a given feature the video segment loss for measuring proximity in; among them, represents the start time, represents the end time, s (i) represents the start and end times when the action occurs; Hypothesis is the set of action true values, where a (i) =(c (i) , s (i) , r (i) ), for the i-th action a (i) , c (i) is the action category, s (i) is the start and end time of the action occurrence, r (i) is the feature of the action; is the action prediction set; N mentioned above is much larger than M; the Hungarian algorithm is used as the matching loss function to calculate the similarity loss between the prediction set and the true value a (i) ; the Hungarian loss L H of all matching pairs is defined as: Among them, is α calculated by optimal bipartite matching (i) is the action sequence number of ; it is 1 when the predicted label is in the set, otherwise it is 0; is for c of (i) the probability of class prediction Loss function.

2. The weakly supervised video temporal action localization method based on action-correlated attention according to claim 1, wherein For the action-correlated attention model, establishing a weakly supervised pre-training using a query mechanism specifically includes: Establish a pre-training task, randomly crop out M video segments S = {S1,..., S M}, and record their timestamps as the ground truth; Extract the features of the M video segments to obtain a query segment set Randomly generate N time regions each containing start and end timestamps, where N is much larger than M, and D represents the number of feature values of each video segment. Encode each time region as an action query Then, the action query set containing N action queries is Divide the query set Qu evenly among the query fragment sets F S , that is, N / M action queries correspond to one Obtain a query set with a corresponding relationship Input the query set with the corresponding relationship into the decoder of the Transformer architecture, and use the set F of query segments with timestamps S to supervise the learning of the action query set so that the time regions of M action queries in Qu correspond one-to-one to the timestamps S recorded in F.

3. The weakly supervised video temporal action localization method based on action correlation attention according to claim 2, characterized in that Use the I3D* network with frozen parameters to extract the features of the M video segments; and / or, during pre-training, randomly mask 10% of the action query segments as zeros.

4. The weakly supervised video temporal action localization method based on action correlation attention according to claim 2, characterized in that, When allocating the query set Qu, the mask matrix is added to the attention layer of the decoder, that is, the attention mask matrix is used to control the interaction between different object queries; the attention mask is as follows: Among them, X i,j Determine whether the action query interacts with the action query; Or, When allocating the query set Qu, randomly shuffle the permutations of all action query encodings.

5. The weakly supervised video temporal action localization method based on action-correlated attention according to claim 1, wherein The action-correlated attention model is trained by the following method: Input a video containing actions as training data; preprocess the training data to obtain video frames and optical flow frames of the video, and extract the I3D features of video segments. Encode the video temporal information of the video segments into position encodings. Input the video temporal position encodings and I3D features of the video segments into the encoder of the action-correlated attention model to determine the relationships between video segment features and achieve the classification of action segments. Input the video temporal position encodings of the video segments into the decoder of the action-correlated attention model. At the same time, establish a weakly supervised pre-training using the query mechanism, and input the output of the query mechanism into the decoder of the action-correlated attention model to achieve the temporal localization of the query set. Merge the query sets output by the encoder and decoder to obtain the localization and classification of action segments in the video.

6. The weakly supervised video temporal action localization method based on action-correlated attention according to claim 1, wherein The positioning loss function For positioning the start and end times of a given feature The video segment loss for measuring proximity in, defining the segment loss as the L1 loss between the prediction and the ground truth and A weighted combination of losses, i.e.: Among them, λ iou and are preset hyperparameters.

7. An electronic device, characterized in that, It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor as described in any one of the methods in claims 1 to 6 is stored on the memory.

8. A computer-readable storage medium, characterized in that, A computer program capable of being loaded and executed by the processor as described in any one of the methods in claims 1 to 6 is stored.

Citation Information

Patent Citations

  • Weak supervision behavior positioning method and device based on action fragment sorting

    CN114049581A

  • Video action clip generation method, system and device and readable storage medium

    CN114339403A