A Multi-scale Action Recognition Method Based on Hierarchical Ordered Residual Network Structure

By adopting a hierarchical orderly residual network structure and attention mechanism in action recognition, motion features are generated and multi-scale spatiotemporal modeling is performed, and a more efficient action recognition effect is achieved.

CN114220171BActive Publication Date: 2025-05-30CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111546471.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-05-30
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

Existing deep learning-based action recognition methods are difficult to effectively capture the diversity of actions in time and space, resulting in insufficient modeling accuracy in space-time modeling.

Method used

A multi-scale action recognition method based on hierarchical orderly residual network structure is adopted to generate motion features through feature segmentation and matrix subtraction, and the motion state and change attention score are calculated in combination with the attention mechanism, the characteristics are stimulated, and multi-scale spatiotemporal relationship modeling is realized through hierarchical orderly residual convolution module.

Benefits of technology

It improves the accuracy of space-time modeling, reduces the amount of calculation, improves the model's ability to understand the space-time of actions, and enhances the accuracy of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220171B_ABST
    Figure CN114220171B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-scale action recognition method based on a hierarchical ordered residual network structure, comprising the following steps: obtaining an original feature group by feature segmentation, calculating a corresponding reference feature group of the original feature group to obtain a motion feature group; calculating a motion state attention score, a motion state change attention score and a comprehensive feature attention score for the motion feature, exciting and outputting the motion feature according to the comprehensive feature attention score, and updating the attention weight sequence number; segmenting the excited motion feature, sorting the feature group according to the updated feature attention score, and implementing multi-scale spatio-temporal relationship modeling through a hierarchical ordered residual convolution module; updating the convolution kernel weight, and repeating the content of S1-S3 until convergence. The present invention can generate spatio-temporal fusion features by a method with less computational complexity, reduce the computational complexity of generating optical flow features to improve the model running speed, and effectively improve the accuracy of spatio-temporal modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a multi-scale action recognition method based on a hierarchical ordered residual network structure. Background Art

[0002] The purpose of video-based action recognition is to classify the actions of objects in a video, which is an important branch in the field of computer vision. This technology is widely applied in fields such as public security, medical treatment, and operational risk management. Due to the large amount of video data, the calculation process of traditional methods will become overly complex and time-consuming.

[0003] With the development of deep learning technology, deep learning-based methods not only exceed traditional methods in terms of accuracy but also significantly reduce the computational amount. Generally, deep learning-based action recognition methods do not fully consider the diversity of actions in time and space. Their model receptive fields are relatively fixed and cannot dynamically adapt to the differences in time and space of different actions, making it difficult for the model to understand. The current classic idea in the field of action recognition is to use spatial and temporal information simultaneously, and capture the information in different-sized (temporal, spatial) receptive fields through a large number of convolutional layers for processing, and use the logical relationship in space and time to infer the actions in the video. Existing methods mainly capture the temporal information between video frames by extracting optical flow information, and process the spatial information through a two-dimensional convolutional neural network, or directly achieve the fusion of temporal and spatial information and the modeling of spatio-temporal features and actions through three-dimensional convolution. These methods all require a large amount of computer resources.

[0004] Therefore, how to provide a multi-scale action recognition method based on a hierarchical ordered residual network structure that can effectively improve the accuracy of spatio-temporal modeling is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a multi-scale action recognition method based on a hierarchical ordered residual network structure that can effectively improve the accuracy of spatio-temporal modeling.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multi-scale action recognition method based on a hierarchical ordered residual network structure, comprising the following steps:

[0008] S1. Perform feature segmentation on each frame of input image data to obtain an original feature group Calculate the original feature group using hierarchical ordered residual convolution The corresponding reference feature group The reference feature group And the original feature group Perform matrix subtraction to obtain a set of motion features where m and f represent the attention weight serial number and the frame serial number respectively;

[0009] S2. For the motion features calculate the motion state attention score and the motion state change attention score for the motion information therein, and further calculate the comprehensive feature attention score of the motion features according to the motion state attention score and the motion state change attention score, stimulate and output the motion features based on the comprehensive feature attention score, and update the attention weight serial number;

[0010] S3. Segment the stimulated motion features sort the feature groups according to the updated feature attention score, and implement multi-scale spatio-temporal relationship modeling through a hierarchical ordered residual convolution module;

[0011] S4. Update the convolutional kernel weights, and repeat the content of S1-S3 until convergence.

[0012] Preferably, before performing S1, it further includes:

[0013] Perform sparse sampling on the video, and each video is segmented into video segments according to a preset time length, and a frame image is randomly selected from each video segment to form input image data.

[0014] Preferably, the specific content of obtaining the reference feature group in S1 includes:

[0015] Segment each frame of input image data to obtain the original feature group and sort the feature groups according to the attention score, and use two-dimensional convolution and residual connection to generate a reference feature group corresponding to the original feature group The reference feature hierarchical ordered residual convolution formula is as follows:

[0016]

[0017] where i(·) represents the mapping relationship between the attention score and the feature group, C i(m) (·) represents the convolutional layer corresponding to the feature group with the feature attention sorting serial number m, represents the original feature group with the feature attention sorting serial number m in the f-th frame, represents the convolutional result of the feature group with the feature attention sorting serial number m.

[0018] Preferably, the specific content of obtaining the motion feature group in S1 includes:

[0019] Use reference features Subtract the original feature group Perform matrix subtraction to obtain the difference between the original feature and the smoothed feature after a short time interval That is, the motion feature group Is:

[0020]

[0021]

[0022] Preferably, S1 also includes: reducing the dimension of the input image data;

[0023] Extract features from the input image data, use two-dimensional convolution to reduce the dimension of the extracted features and perform feature offset, and the dimension reduction method is:

[0024]

[0025] Among them, Represents the input feature of the f-th frame, C(·) represents the two-dimensional convolutional layer for dimension reduction, and S(·) represents the feature offset module.

[0026] Preferably, for the motion feature in S2 The specific content of calculating the motion state attention score and the motion state change attention score for the motion information in is as follows:

[0027] Use two-dimensional average pooling to capture the motion information in the motion feature And calculate the motion state attention score. The motion state attention score formula is as follows:

[0028]

[0029] Among them, avg(·) represents the two-dimensional average pooling layer, and C v (·) represents the two-dimensional convolutional layer for restoring the feature dimension;

[0030] Use max pooling and min pooling to capture the change information of the motion state in the optical flow feature And calculate the motion state change attention score. The motion state change attention score formula is as follows:

[0031]

[0032] Among them, max(·) represents the two-dimensional max pooling layer, min(·) represents the two-dimensional min pooling layer, and C a (·) represents the two-dimensional convolutional layer to make the weight dimension consistent with the input feature dimension.

[0033] Preferably, the motion features calculated in S2 The specific method for calculating the comprehensive feature attention score is as follows:

[0034] W = σ(δ(λ 1 ×W v +λ 2 ×W a ))

[0035] Where, δ represents a two-dimensional normalization layer, σ represents a sigmoid activation function, λ 1 and λ 2 both represent hyperparameters, W v represents the motion state attention score, and W a represents the motion state change attention score.

[0036] Preferably, the specific content of motivating the motion features according to the comprehensive feature attention score in S2 and outputting, and updating the attention weight sequence number includes:

[0037] Motivate the motion features using the comprehensive feature attention score, calculate the Hadamard product of the attention score and the motion features, amplify the features with higher attention scores, and suppress the features with lower attention scores. The feature motivation method is:

[0038]

[0039] Where, represents the Hadamard product, represents the input feature of the f-th frame, represents the motion feature of the f-th frame, represents the excitation feature of the f-th frame;

[0040] Update the attention weight sequence number, and use a two-dimensional average pooling layer and a one-dimensional average pooling layer to make the dimension of the attention score consistent with the number of feature groups. The specific method is:

[0041] i(m) = sort(avg 1d (avg 2d (W m )))

[0042] Where, avg 2d (·) represents a two-dimensional average pooling layer, and avg 1d (·) represents a one-dimensional average pooling layer, and sort(·) represents a sorting function for the feature groups and the attention scores.

[0043] Preferably, the specific content of S3 includes:

[0044] Segment the input features and perform hierarchical ordered residual convolution on the feature groups according to the attention scores:

[0045]

[0046] Among them, S i(m) (·) represents the feature offset module corresponding to the feature group, and C i(m) (·) represents the two-dimensional convolutional layer corresponding to the feature group;

[0047] Reorder and splice the feature groups according to the attention scores to obtain the output features

[0048]

[0049] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a multi-scale action recognition method based on a hierarchical ordered residual network structure. This method can generate spatio-temporal fusion features, that is, motion features, in a method with less computational complexity, reduce the computational complexity of generating optical flow features to improve the model running speed; introduce an attention mechanism to evaluate the motion state and the change of the motion state and stimulate the features accordingly, improve the accuracy of the model neural network for spatio-temporal modeling of actions; use the importance of features for action recognition to sort, improve the utilization rate of features with higher importance in hierarchical residual convolution, and further improve the accuracy of spatio-temporal modeling. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0051] Figure 1 The drawings are the flowcharts of a multi-scale action recognition method provided by the present invention based on a hierarchical ordered residual network structure;

[0052] Figure 2 The drawings are the framework diagrams of a multi-scale action recognition method provided by the present invention based on a hierarchical ordered residual network structure. Detailed Embodiments

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] An embodiment of the present invention discloses a multi-scale action recognition method based on a hierarchical ordered residual network structure, as Figure 1-2 shown, which includes the following steps:

[0055] S1. Perform feature segmentation on each frame of input image data to obtain the original feature group Calculate the original feature group using hierarchical ordered residual convolution The corresponding reference feature group Subtract the reference feature group from the original feature group to obtain the motion feature group where m and f respectively represent the attention weight serial number and the frame serial number;

[0056] S2. Calculate the motion state attention score and the motion state change attention score for the motion information in the motion feature , and further calculate the comprehensive feature attention score of the motion feature according to the motion state attention score and the motion state change attention score. Output the excitation of the motion feature based on the comprehensive feature attention score and update the attention weight serial number;

[0057] S3. Segment the excited motion feature , sort the feature group according to the updated feature attention score, and implement multi-scale spatio-temporal relationship modeling through a hierarchical ordered residual convolution module;

[0058] S4. Update the convolution kernel weight, and repeat the content of S1-S3 until convergence.

[0059] It should be noted that:

[0060] The convergence condition is to complete the iteration and the change of the loss function does not exceed the threshold.

[0061] To further implement the above technical solution, before performing S1, it further includes:

[0062] Perform sparse sampling on the video, and each video is segmented into video segments according to a preset time length, and a frame of image is randomly selected from each video segment to form the input image data.

[0063] To further implement the above technical solution, the specific content of obtaining the reference feature group in S1 is as follows:

[0064] Segment each frame of input image data to obtain the original feature group And sort the feature groups according to the attention scores, and use two-dimensional convolution and residual connection to generate the reference feature group corresponding to the original feature group The reference feature hierarchical ordered residual convolution formula is as follows:

[0065]

[0066] where i(·) represents the mapping relationship between the attention score and the feature group, and C i(m) (·) represents the convolutional layer corresponding to the feature group with the feature attention sorting number m, represents the original feature group with the feature attention sorting number m in the f-th frame, represents the convolutional result of the feature group with the feature attention sorting number m.

[0067] It should be noted that:

[0068] In this embodiment, C i(m) (·) is a 3×3 two-dimensional convolutional layer with a stride of 2 corresponding to the feature group.

[0069] To further implement the above technical solution, the specific content of obtaining the motion feature group in S1 is as follows:

[0070] Use the reference feature and the original feature group to perform matrix subtraction to obtain the difference between the original feature and the smoothed feature after a short time interval That is, the motion feature group is:

[0071]

[0072]

[0073] It should be noted that:

[0074] Using the reference feature and the original feature group to perform matrix subtraction can obtain the difference between the original feature and the smoothed feature after a short time interval According to the formula:

[0075]

[0076] The difference between a feature and the mean value of the features at the same position, divided by the time interval, can represent the magnitude of the state change within this time interval at the feature. At this time, the original feature represents position information, so this difference can mark the change information of the position, that is, the motion information. Therefore, this feature difference is called the motion feature, and the calculation method of the motion feature is as shown in the above formula.

[0077] To further implement the above technical solution, S1 further includes: reducing the dimension of the input image data;

[0078] Extracting features from the input image data, using two-dimensional convolution to reduce the dimension of the extracted features and perform feature offset. The dimension reduction method is: It should be noted that:

[0079]

[0080] Among them, represents the input feature of the f-th frame, C(·) represents the two-dimensional convolutional layer for dimension reduction, and S(·) represents the feature offset module.

[0081] It should be noted that:

[0082] In this embodiment, a 1×1 two-dimensional convolution is used to reduce the dimension of the features and perform feature offset, reducing the amount of data for subsequent processing, that is, C(·) is a 1×1 two-dimensional convolutional layer for dimension reduction.

[0083] To further implement the above technical solution, the specific content of calculating the motion state attention score and the motion state change attention score for the motion information in the motion feature in S2 includes: Using two-dimensional average pooling to capture the motion information in the motion feature

[0084] and calculating the motion state attention score. The formula for the motion state attention score is as follows:

[0085]

[0086] Among them, avg(·) represents the two-dimensional average pooling layer, and C v (·) represents the two-dimensional convolutional layer for restoring the feature dimension;

[0087] Using max pooling and min pooling to capture the change information of the motion state in the optical flow feature and calculating the motion state change attention score. The formula for the motion state change attention score is as follows:

[0088]

[0089] Among them, max(·) represents a two-dimensional max pooling layer, min(·) represents a two-dimensional min pooling layer, and C a (·) represents a two-dimensional convolutional layer, which makes the weight dimension consistent with the input feature dimension.

[0090] It should be noted that:

[0091] In this embodiment, C v (·) and C a (·) are both 1×1 two-dimensional convolutional layers with a stride of 1 used to restore the feature dimension.

[0092] To further implement the above technical solution, the specific method for calculating the comprehensive feature attention score of the motion feature obtained in S2 is: is:

[0093] W = σ(δ(λ 1 ×W v + λ 2 ×W a ))

[0094] Among them, δ represents a two-dimensional normalization layer, σ represents a sigmoid activation function, λ 1 and λ 2 both represent hyperparameters, W v represents the motion state attention score, and W a represents the motion state change attention score.

[0095] It should be noted that:

[0096] In this embodiment, the hyperparameters are all taken as 0.5.

[0097] To further implement the above technical solution, the specific content of exciting and outputting the motion feature according to the comprehensive feature attention score in S2 and updating the attention weight serial number includes: is:

[0098] Using the comprehensive feature attention score to excite the motion feature, calculating the Hadamard product of the attention score and the motion feature, amplifying the features with higher attention scores, and suppressing the features with lower attention scores. The feature excitation method is:

[0099]

[0100] Among them, represents the Hadamard product, represents the input feature of the f-th frame, represents the motion feature of the f-th frame, represents the excitation feature of the f-th frame;

[0101] Update the attention weight sequence number. Use a two-dimensional average pooling layer and a one-dimensional average pooling layer to make the dimension of the attention score consistent with the number of feature groups. The specific method is as follows:

[0102] i(m) = sort(avg 1d (avg 2d (W m )))

[0103] Among them, avg 2d (·) represents the two-dimensional average pooling layer, and avg 1d (·) represents the one-dimensional average pooling layer. sort(·) represents the sorting function for the feature group and the attention score.

[0104] To further implement the above technical solution, the specific content of S3 includes:

[0105] Segment the input features and perform hierarchical ordered residual convolution on the feature groups according to the attention score:

[0106]

[0107] Among them, S i(m) (·) represents the feature offset module corresponding to the feature group, and C i(m) (·) represents the two-dimensional convolution layer corresponding to the feature group;

[0108] Reorder and splice the feature groups according to the attention score to obtain the output features

[0109]

[0110] It should be noted that:

[0111] In this embodiment, C i(m) (·) is a 3×3 two-dimensional convolution layer with a stride of 2 corresponding to the feature group.

[0112] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, please refer to the description in the method part.

[0113] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-scale action recognition method based on a hierarchical ordered residual network structure, characterized in that, it includes the following steps: S1. Segment the feature of each input frame image data to obtain the original feature group Calculate the original feature group by using hierarchical ordered residual convolution The corresponding reference feature group The calculation method of the reference optical flow feature, use the reference feature group And the original feature group Perform matrix subtraction to obtain the motion feature group Where m and f respectively represent the attention weight serial number and the frame serial number; S2. Targeting motion features The motion information in the calculation is used to calculate the motion state attention score and the motion state change attention score, and the motion feature is further calculated based on the motion state attention score and the motion state change attention score. The comprehensive feature attention score is used to evaluate the motion feature Provide incentives and outputs, and update the attention weight sequence number; S3. For the stimulated motion features perform segmentation, sort the original feature group according to the updated comprehensive feature attention score, and implement multi-scale spatio-temporal relationship modeling through the hierarchical ordered residual convolution module; The specific content of S3 includes: Segment the input features and perform hierarchical ordered residual convolution on the feature groups according to the attention scores: Among them, S i(m) (·) represents a feature offset module corresponding to a feature group, C i(m) (·) represents a two-dimensional convolutional layer corresponding to a feature group; Reorder and splice the feature groups according to the attention scores to obtain the output features S4. Update the convolution kernel weights, and repeat the content of S1 - S3 until the network parameters converge.

2. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, before performing S1, it further includes: Sparse sampling is performed on the video, and each video is segmented into video segments according to a preset time length, and one frame image is randomly selected from each video segment to form the input image data.

3. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, Obtaining a reference feature group in S1 The specific content includes: Segment each frame of the input image data to obtain the original feature group And sort the feature groups according to the attention scores, and use two-dimensional convolution and residual connections to generate the same as the original feature group Corresponding reference feature group The reference feature hierarchical ordered residual convolution formula is as follows: Among them, i(·) represents the mapping relationship between the attention score and the feature group, and C i(m) (·) represents the convolutional layer corresponding to the feature group with the m-th feature attention sorting number, represents the original feature group with the m-th feature attention sorting number in the f-th frame, represents the convolutional result of the feature group with the m-th feature attention sorting number.

4. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, Obtain the motion feature group in S1 The specific content includes: Use reference features and the original feature group to perform matrix subtraction to obtain the difference between the original feature and the smoothed feature after a short time interval i.e., the motion feature group which is 5. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, In S1, it further includes: reducing the dimension of the input image data; Feature extraction is performed on the input image data, and the extracted features are reduced in dimension and feature offset is performed using two-dimensional convolution. The dimension reduction method is: Among them, represents the input feature of the f-th frame, C(·) represents the two-dimensional convolutional layer for dimensionality reduction, and S(·) represents the feature offset module.

6. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, For the motion features in S2 The specific content of calculating the motion state attention score and the motion state change attention score for the motion information in Capturing motion features using two-dimensional average pooling and calculating a motion state attention score. The formula for the motion state attention score is as follows: Among them, avg(·) represents a two-dimensional average pooling layer, and C v (·) represents a two-dimensional convolutional layer for restoring the feature dimension; Capture motion features using max pooling and min pooling The change information of the motion state in and calculate the motion state change attention score. The motion state change attention score formula is as follows: Among them, max(·) represents a two-dimensional max pooling layer, min(·) represents a two-dimensional min pooling layer, and C a (·) represents a two-dimensional convolutional layer, which makes the weight dimension consistent with the input feature dimension.

7. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, The motion features calculated in S2 The specific method for calculating the comprehensive feature attention score is as follows: W = σ(δ(λ 1 ×W v + λ 2 ×W a )) where, δ represents the two-dimensional normalization layer, σ represents the sigmoid activation function, λ 1 and λ 2 both represent hyperparameters, W v represents the motion state attention score, W a represents the motion state change attention score.

8. The multi-scale action recognition method based on a hierarchical ordered residual network structure according to claim 1, characterized in that, In S2, according to the comprehensive feature attention score, the motion features are stimulated and output, and the specific content of updating the attention weight serial number includes: The motion features are excited using the comprehensive feature attention scores, the Hadamard product of the attention scores and the motion features is calculated, the features with higher attention scores are amplified, and the features with lower attention scores are suppressed. The feature excitation method is: Among them, represents the Hadamard product, represents the input feature of the f-th frame, represents the motion feature of the f-th frame, represents the excitation feature of the f-th frame; Update the attention weight sequence numbers, and use a two-dimensional average pooling layer and a one-dimensional average pooling layer to make the dimension of the attention scores consistent with the number of feature groups. The specific method is: i(m) = sort(avg 1d (avg 2d (W m ))) Among them, avg 2d (·) represents a two-dimensional average pooling layer, avg 1d (·) represents a one-dimensional average pooling layer, and sort(·) represents a sorting function for the feature group and the attention score.