A structured attention synthesis method for temporal action localization
By introducing a structured attention synthesis method in timing action positioning, using optimal transmission theory combined with modal and frame attention, the problem of failure to effectively capture frame-modal structure in the prior art is solved, and higher positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202210500127.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-06
AI Technical Summary
When the existing timing action positioning method integrates appearance characteristics and motion characteristics, it fails to effectively capture the frame-modal structure, resulting in suboptimal model performance.
A structured attention synthesis method is proposed, combining modal attention and frame attention through optimal transmission theory, learning the encoding of frame-modal structures and regularizing it to improve the accuracy of action positioning.
This method improves the accuracy of timing action positioning by accurately inferring frame-modal attention and shows better robustness when processing video data in different scenarios.
Smart Images

Figure CN115240097B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision algorithm research and relates to a structured attention synthesis method for temporal action localization. Specifically, it relates to a method for completing the task of temporal action localization in videos by combining optimal transmission theory with an action localization network in the context of video understanding. Background Art
[0002] In recent years, the rapid development of social media and video sharing websites has led to an increasingly strong demand for video processing. The research on temporal action localization methods has significant application value.
[0003] Since Karen Simonyan proposed two-stream networks in her 2014 work, Two-stream convolutional networks for action recognition in videos, researchers such as Joao Carreira have verified the effectiveness and necessity of simultaneously utilizing appearance and motion features in works such as QuoVadis, Action Recognition A New Model and the Kinetics Dataset. Regarding current action localization research, Xin Li et al. treated these two feature modalities equally in their 2019 work, Deep Concept-wise Temporal Convolutional Networks for Action Localization, and concatenated the two features. The complementarity of the two modalities was mined by a subsequent network in a data-driven manner. In their 2018 work, Rethinking the Faster R-CNN Architecture for Temporal Action Localization, Yu-Wei Chao et al. learned two action localization models from two feature modalities, respectively, and then fused the localization results using a trade-off coefficient.
[0004] While temporal action localization has been well developed, the difference between appearance and motion modalities has been inadvertently ignored, leaving the learned models suboptimal. Recently, Weiyao Wang et al. (2020) mentioned in their work WhatMakes Training Multi-modal Classification Networks Hard that integrating appearance and motion features can improve or degrade performance, depending on whether an effective integration strategy can be designed. Appearance and motion modalities exhibit different advantages for action instances. However, insufficient modeling of the impact of each modality and each frame has limited the performance of existing works. Summary of the Invention
[0005] Technical problems to be solved
[0006] In order to avoid the shortcomings of the existing technology, the present invention proposes a structured attention synthesis method for temporal action localization, which is used to perform multimodal feature learning and complete the temporal action localization task.
[0007] The basic idea of the present invention is: input an unedited video and its appearance features and motion features. This method first infers modal attention and frame attention respectively. The former uses global features to judge the pros and cons of the modality, and the latter uses local features to estimate potential motion frames, and then synthesizes frame-modal attention. In order to capture the frame-modal structure, the learning process between modal attention and frame attention is regarded as a structure-guided attention allocation process. Specifically, modal attention is regarded as a supplier and frame attention is regarded as a receiver. Each allocation from supplier to receiver will incur a specific cost. The purpose of the allocation process is to find an optimal allocation scheme to minimize the overall allocation cost. Based on the optimal transmission theory, such a process can prove the adaptability between the supplier and the receiver by considering the close or distant relationship between each pair of supplier-receiver. Finally, this module can obtain the video appearance features and motion features that have been adjusted by attention.
[0008] Technical Solution
[0009] A structured attention synthesis method for temporal action localization is characterized by the following steps:
[0010] Step 1: Extract video features:
[0011] For an unedited video, its appearance characteristics are The movement characteristics are Where D represents the feature latitude, and T represents the length of the feature sequence;
[0012] First, F a and F mInput two convolutional layers respectively, and the output results are Θ a (F a ) and Θ m (F m ), and splice the output results to get the video features
[0013] Step 2: Infer frame-modality attention:
[0014] Step a: Calculate F V The Euclidean norm of each eigenvector in , reorder the eigenvectors in descending order of norm size, and obtain the sorting index
[0015] Set the ratio k and extract F V The top N with the largest mean norm k feature vectors, where Indicates rounding down and extracting the sort index χ v The first N k The items form a new index χ = χ v [1:N k ];
[0016] From F V N extracted from k Perform average pooling on the feature vectors to obtain action-perception features in After reordering according to the size of the Euclidean norm, the feature vector with index t is in the original video feature F V The index in
[0017] f w Enter a convolutional layer with Softmax activation to obtain modality attention
[0018] Step b: Video feature F V Input a temporal convolution layer with Sigmoid activation to obtain frame attention
[0019] Step c: Attention to modality a m and frame attention a f Do matrix multiplication to get the frame-modality attention A=a m ×a f ,
[0020] Step 3: Evaluate attention allocation results based on optimal transfer theory:
[0021] Define two 2×T matrices Ψ and S, ψ m,t and sm,t denote the elements in the mth row and tth column of Ψ and S respectively, where the matrix S can be called the structure matrix;
[0022] Given a modality attention a m and frame attention a f , initialize the structure matrix S and input it into a convolutional layer, then use the Sinkhorn-Knopp algorithm proposed by Marco Cuturi in his 2013 work Sinkhorn distances: Lightspeed computation of optimal transport to calculate the dual-Sinkhorn divergence and optimize the following problem:
[0023]
[0024]
[0025] m∈[1,2],t∈[1,2,...,T],
[0026] Get the optimal solution Ψ * and in As the loss function, the structure matrix S is optimized in an iterative manner to optimize the modal attention a m and frame attention a f ;
[0027] Step 4. Loss function:
[0028] Step a: Structural matrix Perform the maximum pooling operation by column to obtain the structure vector Then calculate the adjacent difference vector d, where the elements Set the ratio η and extract the first N smallest values among the adjacent difference vectors d η Item element, where Then calculate this N η The average value q of the item elements d :
[0029]
[0030] Then define the loss function
[0031] Step b: In order to prevent the structure matrix S from having zero solutions, define the loss function
[0032] Step c: Calculation where λ s and λ F is the equalization coefficient; Used as a loss function for training the structured attention synthesis module to obtain the frame-modality attention A;
[0033] Step 5: Set Θ a (F a ) and Θ m (F m ) and the frame-modality attention A to calculate the dot product and obtain the attention-adjusted video feature H a and H m , as the overall output of the structured attention synthesis module.
[0034] The ratio k is 0.125.
[0035] Beneficial effects
[0036] The present invention proposes a structured attention synthesis method for temporal action localization, and constructs a novel structured attention synthesis module. By converting the relationship between modal attention and frame attention into an attention allocation process, the module can learn the encoding of the frame-modal structure and use it to regularize the frame attention and modal attention respectively according to the optimal transmission theory. The final frame-modal attention is obtained by synthesizing two separate attentions. The structured attention synthesis module proposed in the present invention can be deployed as a plug-and-play module in the existing action localization framework, has higher positioning accuracy, and exhibits better robustness when processing video data from different scenes.
[0037] This paper studies temporal action localization from the perspective of multimodal feature learning, a point previously overlooked but crucial for localization performance. We propose a structured attention synthesis module to accurately infer frame-modality attention. Under optimal transmission, structural information is captured by a learnable matrix, demonstrating the goodness of fit between frame and modality attention.
[0038] Compared with existing temporal action localization methods, the method of the present invention has higher positioning accuracy and shows better robustness when processing video data of different scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flow chart of the method of the present invention
[0040] Figure 2 This is a visualization of some training data
[0041] Figure 3 This is the experimental result diagram of the method of the present invention DETAILED DESCRIPTION
[0042] The present invention will now be further described with reference to the embodiments and accompanying drawings:
[0043] The computer hardware used for this implementation is an Intel Xeon E5-2600 v3 2.6GHz 8-core CPU processor, 128GB of RAM, and a GeForce GTX TITAN 2080Ti GPU. The software environment is Linux 16.04 64-bit. We implemented the proposed method using PyTorch 1.5.
[0044] Reference Figure 1 The method flow chart of the present invention is specifically implemented as follows:
[0045] 1. Construct a training dataset. In this embodiment, the THUMOS14 dataset is used for experiments. The dataset is sourced from http: / / crcv.ucf.edu / THUMOS14 / . This training dataset contains 20 action categories, and each training video contains multiple action instances. Figure 2 All videos are processed individually using the method of the present invention.
[0046] 2. Extract video features. Use the I3D model proposed by Joao Carreira in his 2017 work Quo vadis, action recognition a new model and the kinetics dataset [C] / / proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 6299-6308. to extract appearance features for each video. and motion characteristics Where D represents the feature dimension and T represents the length of the feature sequence.
[0047] The I3D model is pre-trained on the Kinetics-400 dataset and does not require fine-tuning when used. The optical flow is calculated using the TV-L1 algorithm mentioned in the 2007 work A duality based approach for realtime TV-L1 optical flow [C] / / Joint pattern recognition symposium. Springer, Berlin, Heidelberg, 2007: 214-223. The Kinetics-400 dataset is from https: / / deepmind.com / research / open-source / kinetics .
[0048] 3. Train the structured attention synthesis module. The module parameters are trained on the PyTorch platform. In this embodiment, the values of the parameters are set as follows: learning rate 0.0001, number of iterations 6000, weight decay coefficient 0.0005, regularization equalization coefficient λ s =0.10,λ F =0.01.
[0049] Based on appearance feature F a and motion characteristics F m As input, it is sent to two temporal convolution layers respectively, and the output result Θ a (F a ) and Θ m (F m ), and concatenate the output results row by row to get the video features
[0050] The frame-modality attention of each video is calculated as follows:
[0051] Step a: Calculate F V The Euclidean norm of each eigenvector in , reorder the eigenvectors in descending order of norm size, and obtain the sorting index
[0052] In this embodiment, the ratio k is set to 0.125, and the F V The top N with the largest mean norm k feature vectors, where Indicates rounding down. Extract sort index χ v The first N k The items form a new index χ = χ v [1:N k ]. V N extracted from k Perform average pooling on the feature vectors to obtain action-perception features in After reordering according to the size of the Euclidean norm, the feature vector with index t is in the original video feature F V The index in .
[0053] f w Input a temporal convolution layer with Softmax activation to obtain modality attention
[0054] Step b: Video feature F V Input a temporal convolution layer with Sigmoid activation to obtain frame attention
[0055] Step c: Attention to modality am and frame attention a f Do matrix multiplication to get the frame-modality attention A=a m ×a f ,
[0056] Given a modality attention a m and frame attention a f , define two 2×T matrices Ψ and S, ψ m,t and s m,t Denote the elements in the mth row and tth column of Ψ and S respectively, where the matrix S can be called the structure matrix. Initialize the structure matrix S.
[0057] The loss function is calculated as follows during each training:
[0058] Step a: Input the structure matrix S into a temporal convolutional layer, and then use the Sinkhorn-Knopp algorithm proposed by Marco Cuturi in 2013 (Sinkhorn distances: Lightspeed computation of optimal transport [J]. Advances in neural information processing systems, 2013, 26: 2292-2300) to calculate the dual-Sinkhorn divergence and optimize the following problem:
[0059]
[0060]
[0061] m∈[1,2],t∈[1,2,...,T],
[0062] Get the optimal solution Ψ * and loss function During training, the structure matrix S is optimized iteratively to optimize the modal attention a m and frame attention a f .
[0063] Step b: Structure matrix Perform the maximum pooling operation by column to obtain the structure vector Then calculate the adjacent difference vector d, where the elements In this embodiment, the ratio η is set to 80%, and the first N adjacent differential vectors d with the smallest values are extracted. η Item element, where Then calculate this N η The average value q of the item elementsd :
[0064]
[0065] Then calculate the smooth regularization loss function
[0066] Step c: To prevent the structure matrix S from having zero solutions, calculate the F-norm regularization loss function
[0067] Step d: Calculate where λ s and λ F is the regularization equalization coefficient, in this embodiment λ s = 0.10 and λ F =0.01.
[0068] As the overall loss function, it is iteratively used to train the structured attention synthesis module, continuously optimizing the structure matrix S, modal attention a m and frame attention a f , obtain the frame-modality attention A.
[0069] 4. Set Θ a (F a ) and Θ m (F m ) and the frame-modality attention A to calculate the dot product and obtain the attention-adjusted video feature H a and H m , as the overall output of the structured attention synthesis module. The structured attention synthesis module proposed in this method is used as a plug-and-play module and combined with the existing temporal action localization framework. The temporal action localization results can be obtained as follows: Figure 3 shown.
Claims
1. A structured attention synthesis method for temporal action localization, characterized by Here are the steps: Step 1: Extract video features: For an unedited video, its appearance characteristics are The movement characteristics are Where D represents the feature latitude, and T represents the length of the feature sequence; First, F a and F m Input two convolutional layers respectively, and the output results are Θ a (F a ) and Θ m (F m ), and concatenate the output results to obtain video features Step 2: Infer frame-modality attention: Step a: Calculate F V The Euclidean norm of each eigenvector in , reorder the eigenvectors in descending order of norm size, and obtain the sorting index Set the ratio k and extract F V The top N with the largest median norm k feature vectors, where Indicates rounding down and extracting the sort index χ v The first N k The items form a new index χ = χ v [1:N k ]; From F V N is extracted from k The feature vectors are averaged and pooled to obtain the action-perception feature. in After reordering according to the size of the Euclidean norm, the feature vector with index t in the original video feature F V The index in ; f w Enter a convolutional layer with Softmax activation to obtain modality attention Step b: Transform the video feature F V Enter a temporal convolutional layer with Sigmoid activation to obtain frame attention Step c: Attention to modality a m and frame attention a f Do matrix multiplication and get the frame-modality attention A=a m ×a f , Step 3: Evaluate the attention allocation results based on the optimal transfer theory: Define two 2×T matrices Ψ and S, ψ m,t and m,t denote the elements in the mth row and tth column of Ψ and S respectively, where the matrix S can be called the structure matrix; Given a modality attention a m and frame attention a f , initialize the structure matrix S and input it into a convolutional layer, and then use the Sinkhorn-Knopp algorithm proposed by Marco Cuturi in his 2013 work Sinkhorn distances: Lightspeed computation of optimal transport to calculate the dual-Sinkhorn divergence to optimize the following problem: m∈[1,2],t∈[1,2,...,T], Get the optimal solution * and in As the loss function, the structure matrix S is optimized in an iterative manner to optimize the modal attention a m and frame attention a f ; Step 4. Loss function: Step a: For the structure matrix Perform the maximum pooling operation by column to get the structure vector Then calculate the adjacent difference vector d, where the elements Set the ratio η and extract the first N smallest values among the adjacent difference vectors d. η Item element, where Then calculate this N η The average value of the item elements q d : Then define the loss function Step b: In order to prevent the structure matrix S from having zero solutions, define the loss function Step c: Calculation where λ s and λ F is the equalization coefficient; It is used as the loss function for training the structured attention synthesis module to obtain the frame-modality attention A. Step 5: Set Θ a (F a ) and Θ m (F m ) is multiplied with the frame-modality attention A to obtain the attention-adjusted video feature H a and H m , as the overall output of the structured attention synthesis module.
2. The structured attention synthesis method for temporal action localization according to claim 1, characterized in that: The ratio k is 0.125.
Citation Information
Patent Citations
Video bullet screen emotion analysis method based on multi-scale attention convolutional coding network
CN111144448A
Action recognition method based on double-flow convolution attention
CN112926396A