Multi-time span context modeling action recognition method based on multi-head self-attention

By adopting a multi-time span context modeling method of multi-head self-attention in action recognition, video feature streams with different time axis sampling rates are processed, and the problem of insufficient time span dependence information capture in the prior art is solved, and higher action video classification accuracy and action positioning performance are achieved.

CN116524402BActive Publication Date: 2025-06-06HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310471296.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-27
Publication Date
2025-06-06
Estimated Expiration
2043-04-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture dependency information of different time spans in action recognition, resulting in unclear action boundaries and difficult to distinguish between start and end boundaries, which affects the accuracy of action video classification.

Method used

A multi-time span context modeling method based on multi-head self-attention is adopted to obtain the feature sequence extracted by video frames and sample them using different time axis sampling rates respectively to form high-frequency and low-frequency time axis sampling streams. These sampling streams go through the multi-time span context aggregation module and the data enhancement module for feature extraction and enhancement, aggregating context information of diversity and robustness, and finally input into the post-processing module for action recognition.

Benefits of technology

It significantly improves the accuracy of action video classification, reduces the interference of video background on action recognition, and enhances the robustness of the model and action positioning performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524402B_ABST
    Figure CN116524402B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-time span context modeling action recognition method based on multi-head self-attention, comprising: obtaining a video to be recognized and extracting video frames to form a first feature sequence; sampling the first feature sequence to obtain a high-frequency time axis sampling stream and a low-frequency time axis sampling stream; extracting features from the high-frequency time axis sampling stream through a first multi-time span context aggregation module and a data enhancement module respectively, correspondingly forming a first context extraction feature and a first enhancement feature, and extracting features from the low-frequency time axis sampling stream through a second multi-time span context aggregation module to form a second context extraction feature; adding and averaging the two context extraction features and then adding and aggregating them with the first enhancement feature to form a first aggregation feature; inputting the first aggregation feature into a post-processing module to obtain an action recognition result. The method can have excellent temporal action positioning performance and high action recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning video understanding, and specifically relates to a multi-time span context modeling action recognition method based on multi-head self-attention. Background Art

[0002] In real life, there are a large number of actions that may occur simultaneously in different time spans and can only provide limited background information. For example, the action of "kicking a football" can obtain the corresponding long-term and short-term action correlation relationships from the two actions of "running" and "passing the ball". In addition, the appearance of the two actions of "running" and "kicking the ball" provides context for the compound action of "shooting". This example shows that an effective time-related modeling technology is needed to detect actions in densely labeled videos. The modeling of the entire network model also needs to take into account the time differences of different actions and personalize the model.

[0003] In the prior art, for temporal relationship modeling in videos without spatial relationship processing, the general approach is to first divide the video into groups of continuous image frames as input. Yeung et al. (Yeung S, Russakovsky O, Mori G, et al. End-to-End Learning of Action Detection from Frame Glimpses in Videos [J]. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.) use LSTM to perform temporal modeling and generate action suggestions. Dai et al. (Dai R, Das S, Bremond F. CTRN: Class-Temporal Relational Network for Action Detection [J]. arXiv preprint arXiv: 2110.13473, 2021.) use one-dimensional convolution to locally model temporal relationships. However, although one-dimensional temporal convolution performs better than LSTM in simulating the long-term structure of actions, due to the limitation of the kernel size of one-dimensional convolution, convolution is difficult to obtain information features of long time spans. Different actions have different speeds. For slow motion, its boundary is usually not a clear time position, but an excessive interval. Therefore, if the local temporal context is not effectively aggregated, there is not enough semantic information to predict these boundaries. In recent years, due to the success of the Transformer model in natural language processing and image tasks, Tan et al. (Tan J, Tang J, Wang L, et al. Relaxed transformer decoders for direct action proposal generation [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 13526-13535.) tried to apply the attention mechanism to the task of temporal action localization. The advantage of this mechanism is that it can establish a one-to-one global relationship between each time period of the video, which can be used to detect highly related actions. However, since each video segment required by the Transformer requires only a few frames, it is often too short compared to the duration of the action instance.Therefore, the traditional Transformer is not suitable for the task of temporal action localization because it has no specific mechanism to enhance the relationship between adjacent video frames. The start and end boundaries of some actions are very similar, and many invalid matches will be generated at these positions. Aggregating adjacent matches of different time scales and semantic densities will destroy the internal semantic representation of the matches. In addition, adjacent matches are highly overlapped, and the semantic information is too similar, so it is impossible to obtain sufficient semantic supplementation after aggregation. Therefore, it is necessary to design an effective action recognition method to address the problems such as unclear boundaries of some actions, difficulty in distinguishing the start and end boundaries, and insufficient capture of dependency information of different time spans by existing network algorithm models. Summary of the invention

[0004] The purpose of the present invention is to address the above-mentioned problems and propose a multi-time span context modeling action recognition method based on multi-head self-attention, which can improve the accuracy of action video classification.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] The multi-time span context modeling action recognition method proposed in the present invention based on multi-head self-attention includes the following steps:

[0007] S1, obtaining a video to be identified and extracting video frames from the video to be identified to form a first feature sequence;

[0008] S2, sampling the first feature sequence using the first time axis sampling rate and the second time axis sampling rate respectively, and correspondingly obtaining a high-frequency time axis sampling stream and a low-frequency time axis sampling stream;

[0009] S3, respectively extracting features from the high-frequency time axis sampling stream through the first multi-time span context aggregation module and the data enhancement module to form the first context extraction feature and the first enhancement feature respectively, and extracting features from the low-frequency time axis sampling stream through the second multi-time span context aggregation module to form the second context extraction feature, each multi-time span context aggregation module includes four attenuation modules, four aggregation modules and a first Concat function, each attenuation module is connected in series in sequence and the output end of each attenuation module is connected one-to-one with the input end of the aggregation module, the first Concat function is used to aggregate the output features of each aggregation module to form the corresponding context extraction feature, and the attenuation modules of the first multi-time span context aggregation module and the second multi-time span context aggregation module use different multiples of the time axis sampling rate for attenuation;

[0010] S4, adding and averaging the first context extraction feature and the second context extraction feature, and then adding and aggregating the feature with the first enhanced feature to form a first aggregated feature;

[0011] S5. Input the first aggregated features into a post-processing module to obtain an action recognition result.

[0012] Preferably, the first time axis sampling rate is T A , the second time axis sampling rate is T A / 2.

[0013] Preferably, the high-frequency time axis sampling stream or the low-frequency time axis sampling stream is obtained as follows:

[0014] The first feature sequence is divided into several segments, and the video frames of each segment are averaged and then superimposed along the time axis to form a C×T-dimensional video representation;

[0015] Wherein, T is the first time axis sampling rate or the second time axis sampling rate, T=t / σ, t is the total number of video frames, σ is the number of video frames of each segment, and C is the feature dimension of each segment.

[0016] Preferably, the attenuation module includes a fusion module, a global module and a local module, wherein:

[0017] The fusion module is used to attenuate the time axis sampling rate of the corresponding input data by a first preset multiple, then increase the feature dimension of the attenuated input data by a second preset multiple, and use the first time domain convolution layer to perform an attenuation operation on the input data after the feature dimension is increased, so as to obtain a first fusion feature;

[0018] The global module is used to extract the global temporal context information of the first fusion feature through a multi-head self-attention layer to obtain the first extracted feature;

[0019] The local module is used to extract the first extracted feature through the second time domain convolution layer, the third time domain convolution layer and the fourth time domain convolution layer in sequence to form a second extracted feature, and add and average the first extracted feature and the second extracted feature to obtain the first local time context information;

[0020] Preferably, the aggregation module is used to perform the following operations:

[0021] The corresponding first local temporal context information is resized and upsampled through the third linear layer to obtain the upsampled feature g n (M n ), the formula is as follows:

[0022] g n (M n )=UpSamoling(M n );

[0023] The upsampled feature g of the fourth attenuation module 4 (M 4) After being resized by the fourth linear layer, it is combined with the current upsampled feature g n (M n ) is added to obtain the corresponding third extracted feature M′ n , the formula is as follows:

[0024]

[0025] Among them, M n represents the corresponding first local temporal context information, n∈{1,…,4}, UpSampling represents the upsampling operation, Indicates adding in element order.

[0026] Preferably, the kernel sizes of the second time domain convolution layer, the third time domain convolution layer and the fourth time domain convolution layer are 1, 3 and 1 respectively.

[0027] Preferably, the first preset magnification of the fusion module of the first multi-time span context aggregation module is 2, the second preset magnification is 1.5, and the kernel size and step size of the time domain convolution layer are both 2; the first preset magnification of the fusion module of the second multi-time span context aggregation module is 1.5, the second preset magnification is 1.5, and the kernel size and step size of the time domain convolution layer are both 2.

[0028] Preferably, the data enhancement module includes eight serially connected enhancement units, each of which includes an expansion module and a fixed module connected in sequence, wherein:

[0029] The expansion module is used to divide the corresponding input data into three paths. The first path is sequentially processed by the fifth time domain convolution layer with a kernel size of 1 and normalization, and the second path is sequentially processed by the kernel size of 2. k The sixth time domain convolution layer and normalization processing, k = l + 1, l is the number of the current expansion module, that is, k ∈ {1, 2, ..., 8}, the third path is normalized, and the input data after each normalization processing is added, and then the fourth extraction feature is formed through the first activation function;

[0030] A fixed module is used to divide the corresponding fourth extracted features into three paths, the first path is sequentially subjected to the seventh time domain convolution layer with a kernel size of 1 and normalization processing, the second path is sequentially subjected to the eighth time domain convolution layer with a kernel size of a preset value and normalization processing, the third path is normalized, and the fourth extracted features after each normalization processing are added, and then subjected to the second activation function to form a fifth extracted feature;

[0031] Then the first enhanced feature F IE It is expressed as follows:

[0032]

[0033] Among them, E P (x z ) l Represents x z The corresponding output feature of the expansion module numbered l, F d (x z ) l Represents x z The corresponding output feature of the fixed module numbered l, l∈[0,7], x z is the zth video frame of the high-frequency timeline sampling stream.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The present application first extracts video frames from the video to be identified to form a first feature sequence, and samples the first feature sequence using different time axis sampling rates to form a low-frequency time axis sampling stream and a high-frequency time axis sampling stream with different time scales, and naturally performs time series modeling on the acquired image sequence; the low-frequency time axis sampling stream and the high-frequency time axis sampling stream are respectively passed through a multi-time span context aggregation module (MRAC) based on a multi-head self-attention mechanism to obtain local and global time context information, and the high-frequency time axis sampling stream is further processed by a data enhancement module (IE) as a supplement to the deficiencies of the multi-head self-attention mechanism in the multi-time span context aggregation module (MRCA), thereby enhancing the aggregation of long time span and short time span information and increasing the diversity and robustness of context information. Specifically, different The time axis sampling stream of the time scale is input into the multi-time span context aggregation module (MRAC) for difference operation, which can significantly reduce the interference of the video background on the accuracy of action recognition without increasing the amount of calculation. The video as a whole can be globally modeled through the attention timing module (the global module of the attenuation module). In this process, the robustness can be continuously enhanced, so that in the aggregation module, the feature fusion of each layer can represent the characteristics extracted from its own layer, which can greatly improve the accuracy of the model in classifying action videos. Finally, the data processed by the data enhancement module and the data processed by the two multi-time span context aggregation modules are aggregated and sent to the post-processing module for classification to obtain the action recognition result, which has excellent temporal action positioning performance and high action recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flowchart of a method for action recognition based on multi-time span context modeling of multi-head self-attention in the present invention;

[0037] Figure 2 It is a schematic diagram of the structure of the multi-time span context aggregation module of the present invention;

[0038] Figure 3This is a schematic diagram of the attenuation module structure of the present invention;

[0039] Figure 4 It is a schematic diagram of the structure of the aggregation module of the present invention;

[0040] Figure 5 It is a structural schematic diagram of a data enhancement module of the present invention;

[0041] Figure 6 It is a schematic diagram of the expansion module structure of the present invention;

[0042] Figure 7 It is a schematic diagram of the fixed module structure of the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0044] It should be noted that when a component is referred to as being "connected" to another component, it may be directly connected to the other component or there may be an intermediate component. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the technical field of this application. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0045] The present invention discloses a multi-time span context modeling action recognition method based on multi-head self-attention, comprising: sampling the pre-extracted video frames into two feature streams of different time scales, called low-frequency time axis sampling stream and high-frequency time axis sampling stream. They are independently input into the multi-time span context aggregation module, structured modeling of time relations on different time scales, and global time context relations are aggregated through the multi-head self-attention mechanism. Secondly, in order to enhance the aggregation of long time span and short time span information and increase the diversity and robustness of context information, a data enhancement module is designed. Finally, the data processed by the data enhancement module is aggregated with the data processed by the two multi-time span context aggregation modules, and then the final prediction result is output through the post-processing module. After multiple experiments, it has been proved that the method has achieved good recognition accuracy on three public data sets (ActivityNet-1.3, Charades, TSU).

[0046] like Figure 1-7 As shown in FIG, a multi-time span context modeling action recognition method based on multi-head self-attention includes the following steps:

[0047] S1. Obtain a video to be identified and extract video frames from the video to be identified to form a first feature sequence.

[0048] S2. Sample the first feature sequence using the first time axis sampling rate and the second time axis sampling rate respectively, and obtain a high-frequency time axis sampling stream and a low-frequency time axis sampling stream accordingly.

[0049] In one embodiment, the first time axis sampling rate is T A , the second time axis sampling rate is T A / 2.

[0050] In one embodiment, the high-frequency time axis sampling stream or the low-frequency time axis sampling stream is obtained as follows:

[0051] The first feature sequence is divided into several segments, and the video frames of each segment are averaged and then superimposed along the time axis to form a C×T-dimensional video representation;

[0052] Wherein, T is the first time axis sampling rate or the second time axis sampling rate, T=t / σ, t is the total number of video frames, σ is the number of video frames of each segment, and C is the feature dimension of each segment.

[0053] Duplicate the first feature sequence twice and adjust their time axis sampling rates to T = T through convolution operations. A 、T=T A / 2, the time axis sampling rate is called T A The characteristic stream is a high-frequency time axis sampling stream, and the time axis sampling rate is T A The feature stream of / 2 is a low-frequency time axis sampling stream. The specific operation is to divide the first feature sequence into several segments, and each segment performs feature averaging of σ frames (that is, σ frames are regarded as a video frame after pixel averaging), and each segment contains σ frames. After the features of each segment are averaged, they are superimposed along the time axis to form a C×T dimensional video representation. For example, the first feature sequence tensor is C×Tσ, which becomes C×T frames after feature averaging of σ frames.

[0054] S3. Feature extraction is performed on the high-frequency time axis sampling stream through the first multi-time span context aggregation module and the data enhancement module respectively to form the first context extraction feature and the first enhancement feature respectively, and feature extraction is performed on the low-frequency time axis sampling stream through the second multi-time span context aggregation module to form the second context extraction feature. Each multi-time span context aggregation module includes four attenuation modules, four aggregation modules and a first Concat function. Each attenuation module is connected in series in sequence and the output end of each attenuation module is connected one by one to the input end of the aggregation module. The first Concat function is used to aggregate the output features of each aggregation module to form the corresponding context extraction feature. The attenuation modules of the first multi-time span context aggregation module and the second multi-time span context aggregation module use different multiples of the time axis sampling rate for attenuation.

[0055] In one embodiment, the attenuation module includes a fusion module, a global module and a local module, wherein:

[0056] The fusion module is used to attenuate the time axis sampling rate of the corresponding input data by a first preset multiple, then increase the feature dimension of the attenuated input data by a second preset multiple, and use the first time domain convolution layer to perform an attenuation operation on the input data after the feature dimension is increased, so as to obtain a first fusion feature;

[0057] The global module is used to extract the global temporal context information of the first fusion feature through a multi-head self-attention layer to obtain the first extracted feature;

[0058] The local module is used to extract the first extracted feature through the second time domain convolution layer, the third time domain convolution layer and the fourth time domain convolution layer in sequence to form a second extracted feature, and add and average the first extracted feature and the second extracted feature to obtain the first local time context information;

[0059] In one embodiment, the aggregation module is used to perform the following operations:

[0060] The corresponding first local temporal context information is resized and upsampled through the third linear layer to obtain the upsampled feature g n (M n ), the formula is as follows:

[0061] g n (M n )=UpSampling(M n );

[0062] The upsampled feature g of the fourth attenuation module 4 (M 4 ) After being resized by the fourth linear layer, it is combined with the current upsampled feature g n (M n) is added to obtain the corresponding third extracted feature M′ n , the formula is as follows:

[0063]

[0064] Among them, M n represents the corresponding first local temporal context information, n∈{1,…,4}, UpSampling represents the upsampling operation, Indicates adding in element order.

[0065] In one embodiment, the kernel sizes of the second time domain convolution layer, the third time domain convolution layer, and the fourth time domain convolution layer are 1, 3, and 1, respectively.

[0066] In one embodiment, the first preset magnification of the fusion module of the first multi-time span context aggregation module is 2, the second preset magnification is 1.5, the kernel size and step size of the time domain convolution layer are both 2, and the first preset magnification of the fusion module of the second multi-time span context aggregation module is 1.5, the second preset magnification is 1.5, and the kernel size and step size of the time domain convolution layer are both 2.

[0067] In one embodiment, the data enhancement module includes eight enhancement units connected in series, each enhancement unit includes an expansion module and a fixed module connected in sequence, wherein:

[0068] The expansion module is used to divide the corresponding input data into three paths. The first path is sequentially processed by the fifth time domain convolution layer with a kernel size of 1 and normalization, and the second path is sequentially processed by the kernel size of 2. k The sixth time domain convolution layer and normalization processing, k = l + 1, l is the number of the current expansion module, that is, k ∈ {1, 2, ..., 8}, the third path is normalized, and the input data after each normalization processing is added, and then the fourth extraction feature is formed through the first activation function;

[0069] A fixed module is used to divide the corresponding fourth extracted features into three paths, the first path is sequentially subjected to the seventh time domain convolution layer with a kernel size of 1 and normalization processing, the second path is sequentially subjected to the eighth time domain convolution layer with a kernel size of a preset value and normalization processing, the third path is normalized, and the fourth extracted features after each normalization processing are added, and then subjected to the second activation function to form a fifth extracted feature;

[0070] Then the first enhanced feature F IE It is expressed as follows:

[0071]

[0072] Among them, E P (x z ) lRepresents x z The corresponding output feature of the expansion module numbered l, F d (x z ) l Represents x z The corresponding output feature of the fixed module numbered l, l∈[0,7], x z is the zth video frame of the high-frequency timeline sampling stream.

[0073] Specifically, the first multi-time span context aggregation module, the data enhancement module and the first multi-time span context aggregation module are processed in parallel:

[0074] 1) The high-frequency time axis sampling stream is input into the first multi-time span context aggregation module (MRCA1) for processing, and each attenuation module performs 2 times of time axis sampling rate attenuation and 1.5 times of feature dimension growth;

[0075] 2) Inputting the low-frequency time axis sampling stream into the second multi-time span context aggregation module (MRCA2) for processing, which is different from the first multi-time span context aggregation module in that the attenuation of the time axis sampling rate in each attenuation module is changed to 1.5 times;

[0076] 3) The high-frequency time axis sampling stream is processed through the data enhancement module.

[0077] Furthermore, each multi-time span context aggregation module (MRCA) is composed of four decay modules, four aggregation modules and a first Concat function. Specifically, it includes:

[0078] Enter the data That is, the high-frequency time axis sampling stream or the low-frequency time axis sampling stream passes through four attenuation blocks in turn. Each attenuation block has the same structure and can be subdivided into three main structures: fusion block, global module and local module. i is the i-th video frame of the high-frequency timeline sampling stream or the low-frequency timeline sampling stream.

[0079] First, we enter the fusion module, which attenuates the time axis sampling rate of the input data by a factor of 2, increases the feature dimension of the attenuated data by a factor of 1.5, and then uses a first time domain convolution layer with a kernel size and step size of 2 to implement the attenuation operation.

[0080] Secondly, the data processed by the fusion module is input into the global module (Global), and the multi-head self-attention layer is used to model the global action dependency and obtain the global temporal context of the action. The multi-head self-attention layer is an existing technology, such as the transformer self-attention module. The calculation process is as follows:

[0081] Enter the data The output of the fusion module is divided into 8 subsequences After being processed by eight self-attention layers, for each head j∈{1,…,8}, the input x ji The output of the fusion module is divided into eight equal parts. Input x i Mapped to and in, Representing the weight of the linear layer, the self-attention of head j can be expressed as follows:

[0082]

[0083] The outputs of different heads are combined and a linear layer is added to aggregate the outputs of the eight self-attention modules above into one output result. After the eight-head self-attention layer, the output feature size is the same as the input feature size, as shown in the following formula:

[0084] P i =W i O Concat(A 1i ,…,A 8i )+x ji

[0085] Among them, W i O ∈R C×C represents the weight of the linear layer, A 1i ,…,A 8i Represents the output results of eight self-attention modules, P i It is the output after the aggregation of the multi-head self-attention layer, and Concat is the aggregation operation of the Concat function.

[0086] Then, the data output by the Global module P i Send to Local module, such as Figure 3As shown in the figure, the Local module aggregates the intermediate results (similarity between input tensors) generated by adjacent Golbal modules by using a second time-domain convolution layer (Convs) with a kernel size of 1, a third time-domain convolution layer with a kernel size of 3, and a fourth time-domain convolution layer with a kernel size of 1, enhances local information by aggregating contextual information of adjacent blocks, obtains local time contextual information, and balances the local and global information distribution of features.

[0087] The output data of the four attenuation modules are respectively provided to the aggregation block. For each attenuation module output M n , n∈{1,…,4}, UpSampling represents the upsampling operation, g n (M n ) is M n The upsampling result can be expressed as follows:

[0088] g n (M n )=UpSampling(M n )

[0089] like Figure 4 As shown in the figure, in the aggregation module, the output of each attenuation module is first resized and upsampled by the linear layer, and finally combined with the fourth attenuation module (i.e. the last attenuation module g M (M N )) output results are added, in this embodiment N = 4. Since the action probability needs to be predicted, the final post-processing module needs to use the original time series length as input for prediction, so the aggregation module sets a linear layer and an upsampling layer. Since the time axis sampling rate between the first attenuation module and the last attenuation module differs by eight times, the time axis sampling rate difference is large. In order to balance the resolution and semantics, the output of the last attenuation module is set to be added to the output of each attenuation module respectively, and the formula is expressed as follows:

[0090]

[0091] in, It means adding in order of elements;

[0092] Finally, the data output by the four aggregation modules will be aggregated by the first Concat function and the output result is as shown in the following formula: HM Represents the high-frequency timeline sampling stream after being processed by the multi-time span context aggregation module:

[0093] F HM =Concat(M′ 1,…,M′ 4 ).

[0094] Furthermore, the data enhancement module (IE) improves the model's ability to aggregate long- and short-range temporal context information. The data enhancement module is composed of eight enhancement units connected in series. In each enhancement unit, two different modules are set: expansion block and fixed block. Specifically, it includes:

[0095] The expansion module divides the input data into three paths. The first path is processed by the fifth temporal convolution layer with a kernel size of 1, in order to aggregate short-range context information. The second path is processed by the fifth temporal convolution layer with a kernel size of 2. k The sixth time domain convolution layer is processed, where k is equal to the number of the corresponding module expansion plus one, that is, k∈{1,2,…,8}, in order to quickly expand the receptive field and aggregate long-distance temporal context information. The third path does not need any processing, in order to retain the original information, reduce the network degradation problem generated during the training process, and make the number of enhancements in the data enhancement module more selective. Finally, they are normalized (Normalization) and combined, and then processed by the first activation function (Relu) and output.

[0096] The only difference between the fixed module and the expansion module is the processing of the second channel data. The kernel size is variable from 2 k Adjust to a fixed value. The purpose of setting a fixed module is to set the convolution layer to 2 if only the expansion module is set. k If the expansion rate is exponential, this practice of expanding the receptive field too quickly may cause gridding artifacts, resulting in the inability of features at certain locations to participate in the calculation, because when the expansion rate is greater than 1, the adjacent units in the output are calculated from completely independent sets in the input. To avoid this phenomenon, the data enhancement module introduces a fixed module and connects it alternately with the expansion module. Figure 5-7 As shown, in this embodiment, the data enhancement module is provided with a total of eight expansion modules and fixed modules, and the specific number can be adjusted according to actual conditions.

[0097] The data enhancement module (IE) can be expressed as follows:

[0098]

[0099] Among them, F IE It represents the high-frequency time axis sampling stream after data enhancement processing, and l∈[0,7] represents eight expansion modules and fixed modules alternately connected.

[0100] S4: After adding and averaging the first context extraction feature and the second context extraction feature, the first enhanced feature is added and aggregated to form a first aggregate feature. BNL The formula is as follows:

[0101] F BNL =[(F HM +F LM ) / 2]+F IE

[0102] Among them, F HM represents the first context extraction feature, F IE represents the first enhancement feature, F LM Represents the second context extraction feature.

[0103] S5. Input the first aggregated feature into a post-processing module to obtain an action recognition result. The post-processing module (PostProcess, Norm & Location) is a technology well known to those skilled in the art. For example, the post-processing module performs the following operations:

[0104] The first aggregate feature F BNL The post-processing module uses a sub-graph alignment and positioning algorithm. The algorithm can be found in the literature: Xu M, Zhao C, Rojas DS, et al. G-TAD: Sub-Graph Localization for Temporal Action Detection[J]. 2019. (Website link: https: / / arxiv.org / abs / 1911.11462). The tuple Y formed by calculating the action category, start time, and end time of the h-th video frame is h ={ic h ,is h ,ie h}, h∈H, H is the number of actions, all tuples form an action sequence in chronological order, ic h is the action category of the hth video frame, is h represents the starting time point of the action category of the hth video frame, ie h Represents the end time point of the action category of the hth video frame. The loss function can be obtained through the training algorithm (also cited from Xu M, Zhao C, Rojas DS, et al. G-TAD: Sub-Graph Localization for Temporal Action Detection[J]. 2019.) to train the overall model of this application, and the high-frequency time axis sampling stream and low-frequency time axis sampling stream corresponding to the video to be identified are input into the trained model to obtain the action recognition result.

[0105] The present application first extracts video frames from the video to be identified to form a first feature sequence, and samples the first feature sequence using different time axis sampling rates to form a low-frequency time axis sampling stream and a high-frequency time axis sampling stream with different time scales, and naturally performs time series modeling on the acquired image sequence; the low-frequency time axis sampling stream and the high-frequency time axis sampling stream are respectively passed through a multi-time span context aggregation module (MRAC) based on a multi-head self-attention mechanism to obtain local and global time context information, and the high-frequency time axis sampling stream is further processed by a data enhancement module (IE) as a supplement to the deficiencies of the multi-head self-attention mechanism in the multi-time span context aggregation module (MRCA), thereby enhancing the aggregation of long time span and short time span information and increasing the diversity and robustness of context information. Specifically, different The time axis sampling stream of the time scale is input into the multi-time span context aggregation module (MRAC) for difference operation, which can significantly reduce the interference of the video background on the accuracy of action recognition without increasing the amount of calculation. The video as a whole can be globally modeled through the attention timing module (the global module of the attenuation module). In this process, the robustness can be continuously enhanced, so that in the aggregation module, the feature fusion of each layer can represent the characteristics extracted from its own layer, which can greatly improve the accuracy of the model in classifying action videos. Finally, the data processed by the data enhancement module and the data processed by the two multi-time span context aggregation modules are aggregated and sent to the post-processing module for classification to obtain the action recognition result, which has excellent temporal action positioning performance and high action recognition accuracy.

[0106] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0107] The above-described embodiments only express the more specific and detailed embodiments described in this application, but they cannot be understood as limiting the scope of the patent application. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of this application, which all belong to the protection scope of this application. Therefore, the protection scope of the patent application shall be based on the attached claims.

Claims

1. A multi-time span context modeling action recognition method based on multi-head self-attention, Features: The multi-time span context modeling action recognition method based on multi-head self-attention comprises the following steps: S1, obtaining a video to be identified and extracting video frames from the video to be identified to form a first feature sequence; S2, sampling the first feature sequence using the first time axis sampling rate and the second time axis sampling rate respectively, and correspondingly obtaining a high-frequency time axis sampling stream and a low-frequency time axis sampling stream; S3, respectively extracting features from the high-frequency time axis sampling stream through the first multi-time span context aggregation module and the data enhancement module to form the first context extraction feature and the first enhancement feature respectively, and extracting features from the low-frequency time axis sampling stream through the second multi-time span context aggregation module to form the second context extraction feature, each of the multi-time span context aggregation modules includes four attenuation modules, four aggregation modules and a first Concat function, each of the attenuation modules is connected in series in sequence and the output end of each attenuation module is connected one-to-one with the input end of the aggregation module, the first Concat function is used to aggregate the output features of each aggregation module to form the corresponding context extraction feature, and the attenuation modules of the first multi-time span context aggregation module and the second multi-time span context aggregation module use different multiples of the time axis sampling rate for attenuation; The attenuation module includes a fusion module, a global module and a local module, wherein: The fusion module is used to attenuate the time axis sampling rate of the corresponding input data by a first preset rate, then increase the feature dimension of the attenuated input data by a second preset rate, and use the first time domain convolution layer to perform an attenuation operation on the input data after the feature dimension is increased, so as to obtain a first fusion feature; The global module is used to extract the global temporal context information of the first fusion feature through a multi-head self-attention layer to obtain a first extracted feature; The local module is used to extract the first extracted feature through the second time domain convolution layer, the third time domain convolution layer and the fourth time domain convolution layer in sequence to form a second extracted feature, and add and average the first extracted feature and the second extracted feature to obtain the first local time context information; S4, adding and averaging the first context extraction feature and the second context extraction feature, and then adding and aggregating the feature with the first enhanced feature to form a first aggregated feature; S5. Input the first aggregated features into a post-processing module to obtain an action recognition result.

2. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The first time axis sampling rate is T A The sampling rate of the second time axis is T A / 2.

3. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The high-frequency time axis sampling stream and the low-frequency time axis sampling stream are obtained as follows: The first feature sequence is divided into several segments, and the video frames of each segment are averaged and then superimposed along the time axis to form a C×T-dimensional video representation; Wherein, T is the first time axis sampling rate or the second time axis sampling rate, T=t / σ, t is the total number of video frames, σ is the number of video frames of each segment, and C is the feature dimension of each segment.

4. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The aggregation module is used to perform the following operations: The corresponding first local temporal context information is resized and upsampled through the third linear layer to obtain the upsampled feature g n (M n ), the formula is as follows: g n (M n )=UpSampling(M n ); The upsampled feature g of the fourth attenuation module 4 (M 4 ) After being resized by the fourth linear layer, it is combined with the current upsampled feature g n (M n ) is added to obtain the corresponding third extracted feature M′ n , the formula is as follows: M′ n =g m (M n )⊕g 4 (M 4 ) Among them, M n Represents the corresponding first local time context information, n∈{1,…,4}, UpSampling represents the upsampling operation, and ⊕ represents adding in element order.

5. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The kernel sizes of the second time domain convolution layer, the third time domain convolution layer and the fourth time domain convolution layer are 1, 3 and 1 respectively.

6. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The first preset magnification of the fusion module of the first multi-time span context aggregation module is 2, the second preset magnification is 1.5, and the kernel size and step size of the time domain convolution layer are both 2. The first preset magnification of the fusion module of the second multi-time span context aggregation module is 1.5, the second preset magnification is 1.5, and the kernel size and step size of the time domain convolution layer are both 2.

7. The multi-time span context modeling action recognition method based on multi-head self-attention as claimed in claim 1, Features: The data enhancement module includes eight enhancement units connected in series, each of which includes an expansion module and a fixed module connected in sequence, wherein: The expansion module is used to divide the corresponding input data into three paths. The first path is sequentially processed by the fifth time domain convolution layer with a kernel size of 1 and normalization, and the second path is sequentially processed by the kernel size of 2. k The sixth time domain convolution layer and normalization processing, k = l + 1, l is the number of the current expansion module, that is, k ∈ {1, 2, ..., 8}, the third path is normalized, and the input data after each normalization processing is added, and then the fourth extraction feature is formed through the first activation function; The fixed module is used to divide the corresponding fourth extracted features into three paths, the first path is sequentially subjected to the seventh time domain convolution layer with a kernel size of 1 and normalization processing, the second path is sequentially subjected to the eighth time domain convolution layer with a kernel size of a preset value and normalization processing, the third path is normalized, and the fourth extracted features after each normalization processing are added, and then subjected to the second activation function to form a fifth extracted feature; Then the first enhanced feature F IE It is expressed as follows: Among them, E P (x z ) l Represents x z The corresponding output feature of the expansion module numbered l, F d (x z ) l Represents x z The corresponding output feature of the fixed module numbered l, l∈[0,7], x z is the zth video frame of the high-frequency timeline sampling stream.

Citation Information

Patent Citations

  • Action recognition algorithm based on human skeleton sequence

    CN113936333A