Fine-grained action recognition and posture detection method based on attention mechanism

By introducing an attention mechanism in human body movement recognition, the problem of insufficient accuracy and robustness of action recognition in the prior art is solved, and efficient identification and rapid detection of fine-grained movements and body postures is achieved.

CN117409482BActive Publication Date: 2025-05-02国家体育总局秦皇岛训练基地(中国足球学校) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311467857.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-02
Estimated Expiration
2043-11-07

AI Technical Summary

Technical Problem

The prior art has problems with insufficient recognition accuracy and robustness in human body movement recognition. Especially when the video scene changes greatly, interferes with content, and different shooting angles, it is difficult to effectively identify fine-grained actions and body postures.

Method used

The fine-grained action recognition and body posture detection method based on attention mechanism is adopted to automatically identify and classify tiny movements of the human body or object through deep learning and attention mechanism. The method includes video data preprocessing, bone point data extraction, action feature map extraction, image and timing attention network processing, and finally body recognition is performed through multiple classifiers.

Benefits of technology

It improves the accuracy and robustness of action recognition, can effectively identify actions in different scenes, angles and videos with long time, and realizes rapid body posture detection during action recognition, improving detection speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117409482B_ABST
    Figure CN117409482B_ABST
Patent Text Reader

Abstract

The present invention provides a fine-grained action recognition and posture detection method based on an attention mechanism, comprising: obtaining video data; intercepting and extracting human skeleton point data by frame from the pre-processed video data to form a temporal key point set; obtaining an action feature map based on the temporal key point set; inputting the action feature map into an image attention network, outputting action categories and action probabilities; processing the frame image through a temporal attention network to obtain action categories and action probabilities; obtaining final action categories and final action probabilities based on the action categories and action probabilities obtained from the image attention network and the temporal attention network; obtaining a fixed frame key point set based on the temporal key point set extraction, inputting the posture recognition network, and obtaining the final classification result of the posture. The present invention can effectively recognize target videos of different scenes, different angles, and different durations, with high recognition accuracy and fast detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and machine learning, and in particular to a fine-grained action recognition and posture detection method based on an attention mechanism. Background Art

[0002] As human action recognition applications become more and more popular, the accuracy and timeliness of human action recognition are also increasing. Accurate recognition of human actions has always been an important research direction in the fields of computer vision and machine learning, and fine-grained action recognition is an important issue in the fields of computer vision and machine learning.

[0003] In the prior art, the recognition of actions is mainly based on convolutional neural networks. However, in the actual application scenarios of human action recognition, due to the large changes in video scenes and the large amount of interference content in video data, the features automatically extracted and learned by convolutional networks can achieve poor recognition effects in action recognition, and the improvement that can be obtained through learning of a large amount of data is relatively small; and the method of human action recognition by combining convolutional networks with skeleton data in the prior art is mainly through, for example, converting skeleton sequence data into a series of three-dimensional coordinates, or artificially designing skeleton sequence data into pictures, or representing skeleton sequence data through graph structures, and then combining with traditional convolutional neural networks for learning. Such methods often cause the loss of key information in human actions, or the ambiguity of feature recognition, and cannot solve the accuracy and robustness problems in human action recognition. At the same time, in the case of uneven quality of existing videos and different shooting angles, it is difficult for existing methods to effectively recognize videos that are greatly different from the training set. Some actions are highly similar under the influence of shooting angles or scenes, and it is difficult for existing technologies to effectively distinguish them. At the same time, some actions are small in amplitude, and it is difficult for existing technologies to accurately recognize them under environmental changes, and fine-grained detection is urgently needed. When judging posture problems through videos, existing technologies mainly use convolutional neural networks or template matching for recognition, which consumes a lot of computing resources and a lot of time during the recognition process. Summary of the invention

[0004] In view of this, this scheme proposes a fine-grained action recognition and posture fast detection method based on attention mechanism, which can automatically recognize and classify the tiny movements of human body or objects, so as to be used in various application fields, including security monitoring, sports training motion recognition analysis, posture monitoring and medical diagnosis, etc. This scheme uses deep learning and attention mechanism to accurately capture the key features of fine-grained actions and improve the accuracy and robustness of action recognition.

[0005] Specifically, the present invention provides the following technical solutions:

[0006] On the one hand, the present invention provides a fine-grained action recognition and posture detection method based on an attention mechanism, the method comprising:

[0007] S1. Acquire video data, preprocess the video data, and obtain preprocessed video data;

[0008] S2, intercepting the pre-processed video data by frame to obtain a frame image; obtaining the target human skeleton point data based on the frame image, and merging them in frame sequence to form a time sequence key point set;

[0009] Based on the set of temporal key points, the action recognition backbone network is used to extract spatial features and temporal features to obtain an action feature map.

[0010] S3, compressing the action feature map and inputting it into the image attention network to obtain an attention matrix, and outputting the action category and action probability;

[0011] Processing the frame image through a temporal attention network to obtain action categories and action probabilities; the temporal attention network includes a local branch network and a global branch network;

[0012] Based on the action categories and action probabilities obtained by the image attention network and the temporal attention network, the final action categories and final action probabilities are obtained to complete fine-grained action recognition;

[0013] S4. Extract the time sequence key point set in S2 to obtain a fixed frame key point set, and input the fixed frame key point set into a posture recognition network; the posture recognition network includes multiple classifiers, and obtains a final classification result of the posture based on the correlation coefficients of the multiple classifiers.

[0014] Preferably, the temporal key point set includes a key point coordinate set, a key point connection edge set, and a key point category code.

[0015] Preferably, in S2, the backbone network extracts spatial features in the spatial domain in the following manner:

[0016]

[0017] Among them, f in is the input matrix of the graph before iteration, f out is the output matrix after iteration, p represents the aggregation of adjacent node parameters, v represents the adjacent node parameters, w represents the update weight, I represents the node label, t represents the current image frame time sequence label, and K represents the total number of nodes.

[0018] Preferably, the adjacent node parameters are aggregated in the following manner: the adjacent node parameters are represented by a degree matrix, an adjacency matrix and a Laplace matrix, and then aggregated by a normalized Laplace transform.

[0019] Preferably, in S2, the backbone network extracts the timing features in the following manner:

[0020] For the human skeleton point data in a single frame image, the feature S is obtained by convolution i , S i After splicing in the time domain, we get the feature set {S1, S2, S3, ...S i}, feature aggregation is performed through a one-dimensional convolutional network:

[0021] A=Softmax(ReLU(Conv1D({S i})))

[0022] Among them, A is the action category and Conv1D is the one-dimensional convolution kernel.

[0023] Preferably, when extracting temporal features, additional codes are added to key points of the same part to represent the same part, thereby increasing the dimension of the key point features; and pooling operations are performed in the time dimension to reduce the dimension of the image frame.

[0024] Preferably, in S3, the image attention network solves the action probability in the following manner:

[0025] The action feature map is subjected to average pooling and maximum pooling operations to obtain two one-dimensional vectors as spatial dimension features;

[0026] The spatial dimension features are processed by a multi-layer perceptron and a sigmoid function to obtain an attention matrix:

[0027] M c (F)=σ(MLP(AvgPool(F))+MLP(max Pool(F)))

[0028] Among them, σ represents the sigmoid function, F represents the action feature map, MLP represents the multi-layer perceptron, and M c Represents the attention matrix, and a linear layer is set after the sigmoid function layer to obtain the action probability.

[0029] Preferably, in S3, the local branch network is processed as follows:

[0030] The frame image is processed by the first temporal convolution to obtain the local field of view temporal features; the local field of view temporal features are batch normalized, input into the second temporal convolution for processing, and then output through the sigmoid function layer. The specific method is as follows:

[0031]

[0032] Among them, V sRepresents the probability of the local branch output, Conv1D represents one-dimensional temporal convolution, BN represents batch normalization, σ represents the sigmoid function, represents a short-term output sequence, which refers to a data sequence formed by storing the segmented video data.

[0033] Preferably, in S3, the global branch network is processed as follows:

[0034] Through the two connected fully connected layers, the frame image is dynamically convolved, and then connected to the sigmoid function layer output:

[0035]

[0036] Among them, V L represents the probability of global branch output, FC represents the fully connected layer, W1 represents the dynamic convolution kernel weight of the first fully connected layer, and W2 represents the dynamic convolution kernel weight of the second fully connected layer. represents a long-term output sequence, σ represents a sigmoid function; and the video data is used as a long-term output sequence.

[0037] Preferably based on the probability V of the local branch output s And the probability V of the global branch output L , weight discrimination is performed through voting to obtain the action probability of the temporal attention network.

[0038] Preferably, in S4, the correlation coefficient is calculated as follows:

[0039] Based on the output results of each classifier, a result matrix is ​​constructed;

[0040] Calculate the correlation coefficient using covariance and standard deviation:

[0041]

[0042] Among them, Cov represents covariance, σ represents variance, Model represents the output result set of each classifier test set, and Label represents the body shape label of the corresponding person to be tested;

[0043] The final classification results of body posture are:

[0044] W final =α1C1+α2C2+α3C3+…

[0045] Among them, C represents the output result of each classifier.

[0046] Preferably, in S4, when the posture recognition network is trained, the training data set is constructed in the following manner:

[0047] Extract the key point category code and key point coordinate set contained in the fixed frame key point set, and store each frame as text format data of [key point category, key point coordinate];

[0048] Label the posture problems in each frame using one-hot encoding, and construct a one-dimensional empty array as the label according to the number of posture problems.

[0049] The labeled dataset is used as the training dataset.

[0050] Preferably, the classifier of the posture recognition network adopts a random forest model, an XGBoost model and a support vector machine based on a tree model.

[0051] In a second aspect, the present invention further provides a fine-grained action recognition and posture detection system based on an attention mechanism, the system comprising:

[0052] A video data preprocessing module preprocesses the video data to obtain preprocessed video data;

[0053] The action feature map calculation module intercepts the preprocessed video data by frame to obtain frame images; obtains the target human skeleton point data based on the frame images, and forms a set of temporal key points after merging them in frame sequence; based on the set of temporal key points, extracts spatial features and temporal features through the action recognition backbone network to obtain an action feature map;

[0054] The attention network module compresses the action feature map and inputs it into the image attention network to obtain the attention matrix, and then obtains the action category and action probability based on the attention matrix by adding a linear layer; the frame image is processed by the temporal attention network to obtain the action category and action probability; the temporal attention network includes a local branch network and a global branch network; the final action category and final action probability are obtained based on the action category and action probability obtained by the image attention network and the temporal attention network, and fine-grained action recognition is completed;

[0055] The posture detection module extracts a set of time-series key points to obtain a set of fixed-frame key points, which is input into a posture recognition network; the posture recognition network includes multiple classifiers, and a final classification result of the posture is obtained based on the correlation coefficients of the multiple classifiers.

[0056] Preferably, the temporal key point set includes a key point coordinate set, a key point connection edge set, and a key point category code.

[0057] In a third aspect, the present invention further provides an electronic device comprising a memory and a processor, wherein the processor can call computer instructions in the memory to execute the fine-grained action recognition and posture detection method based on the attention mechanism as described above.

[0058] Compared with the prior art, the technical solution of the present invention has obvious advantages in recognition accuracy and application scenarios. The spatiotemporal graph convolutional neural network model based on skeleton points can effectively recognize target videos of different scenes, different angles and different durations. An attention mechanism is added to the model for fine-grained action recognition, and high-precision recognition can still be performed in cases where the similarity of multiple actions is high and the degree of action is small and difficult to recognize with the prior art. At the same time, this solution uses a machine learning algorithm to quickly detect the body posture standard of the target person during the action recognition process, and the detection speed is better than the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0060] Figure 1 A schematic diagram of a key point detection process according to an embodiment of the present invention;

[0061] Figure 2 A schematic diagram of the attention-based fine-grained action recognition process according to an embodiment of the present invention;

[0062] Figure 3 A schematic diagram of a rapid posture recognition process according to an embodiment of the present invention;

[0063] Figure 4 The figure is a flow chart of the overall solution of an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0065] Those skilled in the art should know that the following specific embodiments or specific implementations are a series of optimized settings listed by the present invention to further explain the specific content of the invention, and these settings can be combined or used in association with each other, unless the present invention clearly states that some or a specific embodiment or implementation cannot be associated or used together with other embodiments or implementations. At the same time, the following specific embodiments or implementations are only used as the most optimized settings, and are not to be understood as limiting the scope of protection of the present invention.

[0066] This solution involves a method for fast detection of fine-grained action recognition and posture based on an attention mechanism, which can automatically recognize and classify tiny movements of the human body or object, and thus be used in various application fields, including security monitoring, sports recognition and analysis, posture monitoring and medical diagnosis, etc. This solution uses deep learning and attention mechanisms to accurately capture the key features of fine-grained actions and improve the accuracy and robustness of action recognition.

[0067] The following, combined Figure 4 The specific implementation method of this solution is described in detail in the specific embodiment shown.

[0068] Step 1: Collection and construction of training data sets: Collect diverse fine-grained action data sets, including video data or image data of different action types and angles. Preprocess the data, including video frame extraction, posture estimation, human key point detection, etc., to facilitate subsequent analysis. Increasing data diversity and richness can improve the robustness of the model. Therefore, it is further preferred that data enhancement techniques, such as image rotation, flipping, scaling, and adding noise, can be used to generate more training samples.

[0069] Step 2: Constructing the action recognition backbone network and building a deep neural network model. Figure 1 , the video is captured frame by frame to obtain frame images, and then the frame images are processed by, for example, OpenPose to obtain the category and coordinates of the target human skeleton points, which are merged in frame sequence and saved as a temporal key point set. The temporal key point combination includes a key point coordinate set, a key point connection edge set, and a key point category code.

[0070] Use graph convolutional neural network to perform spatial domain graph convolution operation on human skeleton points and connections in the set of temporal key points to extract the spatial features of human skeleton points and connections. In the spatial domain, the network weights are updated by sampling the neighbor feature graph and multiplying it by the corresponding weights, aggregating the features of adjacent nodes, representing the graph node data through the degree matrix, adjacency matrix and Laplace matrix, and aggregating through normalized Laplace transform to achieve the effect of extracting spatial features. The formula is as follows:

[0071]

[0072] Among them, f in is the input matrix of the graph before iteration, f out is the output matrix after iteration, p represents the aggregation of adjacent node parameters, v is the adjacent node parameter, w is the update weight, I is the node label, t represents the current picture frame timing label, and K represents the total number of nodes. The adjacent node parameters include the node category and node coordinates of the adjacent nodes. The above-mentioned aggregation process can be implemented using a conventional convolutional network in the prior art, which is known to those skilled in the art and will not be repeated here.

[0073] At the same time, a temporal convolutional network is used to perform time domain convolution on the human skeleton points and connections in the data set to obtain temporal features superimposed on the skeleton point graph to learn the local change characteristics of the joints in time. For temporal feature extraction, the following method can be preferably used: for the skeleton features (i.e., human skeleton point data) in a single frame image, the feature S is obtained after graph convolution i , and then concatenate them in the time domain to obtain the feature set {S1, S2, S3, ...S i}, feature aggregation is performed through a one-dimensional convolutional network, and then the action category and occurrence probability are output.

[0074] A=Softmax(ReLU(Conv1D({S i})))

[0075] Where A is the action category, S is the time series set, and Conv1D is the one-dimensional convolution kernel.

[0076] In this process, additional codes are added to the key points of the same parts, such as the palm, elbow and shoulder are unified as the content of the character's hand, so as to increase the dimension of the joint features; pooling operation is performed on the time series to reduce the dimension of the time series key point set in the time series, and finally feature classification is performed through the average pooling layer, the fully connected layer and the subsequent SoftMax layer to further infer the corresponding human body movements.

[0077] Step 3: Attention mechanism design: Combination Figure 2 As shown, on the basis of the action recognition model, an attention mechanism is added to perform fine-grained action recognition. In this embodiment, attention mechanisms are added to the network model from two aspects: time series and image features. On the one hand, a time series score is performed on the recognition video frame sequence set in the overall video, and on the other hand, a channel attention mechanism is added to the video frame image to improve the model's utilization of detail features.

[0078] Step 3-1: Image attention mechanism: By adding a channel attention mechanism to learn semantically sensitive objects in fine-grained images. First, the action recognition backbone network is used to extract the action feature map of the image, and the main network expresses the overall information of the fine-grained features through a dual-channel fusion network; then the channel attention mechanism is added by compressing the input feature map and inputting it into a multi-layer perceptron, and the feature map channels are sorted in time sequence and put into the attention network to obtain the ability to represent detailed features.

[0079] First, the action feature map obtained after the action recognition backbone network is processed is compressed in the spatial dimension to obtain a one-dimensional vector before further operation. When compressing the input feature map in the spatial dimension, while performing averagepooling, max pooling is additionally introduced as a supplement. After two pooling functions, a total of two one-dimensional vectors can be obtained as the features of the feature map in the spatial dimension. After obtaining the two vectors after the pooling operation, the attention matrix with feature information is obtained through a trainable multi-layer perceptron and a sigmoid function.

[0080] The formula is as follows:

[0081] M c (F)=σ(MLP(AvgPool(F))+MLP(max Pool(F)))

[0082] V channel = Linner(σ(M c (F)))

[0083] Among them, σ represents the sigmoid function, F represents the action feature map, MLP represents the multi-layer perceptron, and M c Represents the attention matrix of the feature map, V channel Represents the action probability output by the image attention network, and Linner represents the linear layer. After obtaining the attention matrix, a linear layer is set after the sigmoid function layer to output the corresponding action probability features.

[0084] Step 3-2: Temporal attention mechanism: By annotating the video clips in the training video data, the action categories, action clip duration, and action clip positions contained in the video data are specifically annotated. The input video is divided by duration, and the video is divided into 3s to 5s segments, which are saved as short-term output sequences, and the entire video is used as a long-term output sequence.

[0085] The temporal attention mechanism consists of two branches: a local branch and a global branch, which aims to learn a position-sensitive importance map to enhance discriminative features, and then output position-invariant weights in a convolutional manner to aggregate temporal information.

[0086] The local branch part uses short-term temporal information aggregation to fully obtain the short-term context information of the action category to determine the probability of the action under the previous and next actions. In the local branch, the one-dimensional convolution module is mainly used for temporal convolution to extract short-term temporal information. A two-stage convolution method is used in the one-dimensional convolution module: the first stage of temporal convolution is used to extract the main feature matrix of the local field of view temporal information (i.e., the local field of view temporal features), and then batch normalization is performed for dimensionality reduction and then input into the second stage of temporal convolution to convert it into a feature matrix, followed by a sigmoid function layer for output. The formula is as follows:

[0087]

[0088] Among them, V s Represents the probability feature of the local branch output, Conv1D represents one-dimensional temporal convolution, BN represents batch normalization, σ represents the sigmoid function, Represents a short-term output sequence.

[0089] The global branch focuses on long-term temporal information and learns position sharing weights through aggregation operations to obtain global information of the target video, thereby judging the probability of action categories occurring in long-term videos from a global perspective. By setting dynamic convolution kernels in video clips, temporal information is aggregated in a convolutional manner. The model learns dynamic convolution kernels through two fully connected layers. The first fully connected layer integrates the overall features, and then passes through the second fully connected layer followed by sigmoid for feature output. The formula is as follows:

[0090]

[0091] Among them, V L Represents the probability feature of the global branch output, FC represents the fully connected layer, W1 represents the dynamic convolution kernel weight of the first layer, and W2 represents the dynamic convolution kernel weight of the second fully connected layer. Represents the long-term output sequence, σ represents the sigmoid function. The long-term output sequence refers to the input video data sequence.

[0092] Finally, the probability outputs of the local branch and the global branch are weighted by the Voting algorithm to complete the final probability after aggregation, that is, the final action probability feature.

[0093] Step 4: Quick body posture detection: Combined Figure 3As shown, posture rapid detection is different from action recognition, and posture detection does not require highly continuous image features and key point sets. By extracting the time series key point set in step Step2 for interval extraction, a fixed frame key point set is obtained, because the extraction here is based on the time series key point set, the extracted fixed frame key point set also includes a key point coordinate set, a key point connection edge set, and a key point category code, and the extracted key point set of size M is saved. Due to the need for rapid detection, a machine learning method is used in this embodiment for rapid posture classification, and an integrated learning method is used to ensure the accuracy and robustness of the model.

[0094] Step4-1: Data format conversion: extract the 17 key point categories and key point coordinates of the extracted interval key point set, and save each picture as text format data of [key point category, key point coordinates]. In a preferred embodiment, the above 17 key point categories include head, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right waist, right knee, right foot, left waist, left knee, left foot, right eye, left eye, right ear, and left ear. When constructing the data set, the key point coordinates of each data are filled in rows in the order of "head, left shoulder, right shoulder..." by Numpy, and a complete data set of size M×17 is obtained. At the same time, whether there is a problem of non-standard posture in each image is marked. If there is a problem of non-standard posture, one-hot encoding is performed according to posture problems such as high and low shoulders, high and low hips, knee valgus and valgus, and head malposition... Label annotation is performed. In the one-hot encoding process, a one-dimensional empty array of size P is constructed as the label according to the number of categories P of non-standard postures. If a certain non-standard posture occurs, the corresponding position in the empty array is set to 1 according to the category order, otherwise it is set to 0. For example, if only high and low shoulders and high and low hips exist at the same time, the label of the person is [1, 1, 0, 0, ...].

[0095] Step 4-2: Machine learning model construction: In this embodiment, the random forest model based on the tree model and the XGBoost model and the support vector machine are selected as the basic classifier model. The data set obtained in step 4-1 is divided into 80% as a training set and 20% as a test set for training. At the same time, in order to ensure the classification speed of the detection model, the number of decision trees of the random forest model and the XGBoost model is controlled within 500 in this embodiment.

[0096] Step 4-3: Comprehensive evaluation of multiple models: Take the classification accuracy of the model in the test set as the benchmark and save the test set results. Construct a 0.2M×3 result matrix and calculate the correlation matrix to obtain the correlation coefficient α of each classifier for the result. In the result matrix, we determine the linear correlation by the ratio of covariance to standard deviation. The correlation coefficient is calculated as follows:

[0097]

[0098] Among them, Cov represents covariance, σ represents variance, Model represents the output result set of each classifier test set, and Label represents the corresponding body shape label of the person to be tested.

[0099] The final result is obtained by multiplying the classification result of each classifier by the correlation coefficient of the corresponding classifier:

[0100] W final =α1C1+α2C2+α3C3+…

[0101] Here, C represents the output result of each classifier.

[0102] Step 5: Video action recognition and posture detection. Divide the final input video to be detected into frames to obtain a set of video frame images. Use the Openpose program to quickly detect human key points and obtain a set of key points for the entire video. Input the overall key point set into the action recognition backbone network for action recognition, and output the action category and probability. Input the feature map in the action recognition backbone network recognition process into the attention module to obtain the category and probability of the action. The final action category and probability are obtained by arithmetically averaging the outputs on both sides to complete fine-grained action recognition. It is expressed as follows:

[0103]

[0104]

[0105] Where P action represents the network output, P model represents the output probability of the action recognition backbone network, P attention represents the output probability of the attention network (i.e., the combined output of the sequential attention network and the image attention network), V L Represents the probability feature of the global branch output, V s Represents the probability characteristics of the local branch output, V channel represents the action probability output by the image attention network (i.e., the probability feature of channel attention), Represents the voting operation.

[0106] The key point set of the entire video is sampled at intervals to obtain a fixed-frame key point set (the fixed-frame key point set includes a key point coordinate set, a key point connection edge set, and a key point category code). The fixed-frame key point set is input into a rapid posture recognition module to determine whether the human body in the video contains any posture health problems and the specific category of posture problems, thereby achieving rapid posture detection.

[0107] In another specific embodiment, the present solution can also be implemented by a fine-grained action recognition and posture detection system based on an attention mechanism, the system comprising:

[0108] A video data preprocessing module preprocesses the video data to obtain preprocessed video data;

[0109] The action feature map calculation module intercepts the preprocessed video data by frame to obtain frame images; obtains the target human skeleton point data based on the frame images, and forms a set of temporal key points after merging them in frame sequence; based on the set of temporal key points, extracts spatial features and temporal features through the action recognition backbone network to obtain an action feature map;

[0110] The attention network module compresses the action feature map and inputs it into the image attention network to obtain the attention matrix, and outputs the action category and action probability; the frame image is processed by the temporal attention network to obtain the action category and action probability; the temporal attention network includes a local branch network and a global branch network; based on the action category and action probability obtained by the image attention network and the temporal attention network, the final action category and final action probability are obtained to complete fine-grained action recognition;

[0111] The posture detection module extracts a set of time-series key points to obtain a set of fixed-frame key points, which is input into a posture recognition network; the posture recognition network includes multiple classifiers, and a final classification result of the posture is obtained based on the correlation coefficients of the multiple classifiers.

[0112] Furthermore, the temporal key point set includes a key point coordinate set, a key point connection edge set, and a key point category code.

[0113] In another embodiment, the present solution can be implemented by means of a device, which may include a corresponding module for performing each or several steps in each of the above-mentioned embodiments. Therefore, each step or several steps of each of the above-mentioned embodiments can be performed by a corresponding module, and the electronic device may include one or more of these modules. The module may be one or more hardware modules specifically configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination. The device can be implemented using a bus architecture.

[0114] Any process or method description in the flowchart or otherwise described herein may be understood to represent a module, fragment or portion of a code including one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiment of the present solution includes other implementations, in which the functions may not be performed in the order shown or discussed, including performing the functions in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by a technician in the technical field to which the embodiments of the present solution belong. The processor performs the various methods and processes described above. Alternatively, in other embodiments, the processor may be configured to perform one of the above methods by any other appropriate means (e.g., by means of firmware).

[0115] The logic and / or steps represented in the flowchart or otherwise described herein may be embodied in any readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from and execute instructions on an instruction execution system, apparatus, or device), or in combination with such instruction execution systems, apparatuses, or devices.

[0116] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A fine-grained action recognition and posture detection method based on attention mechanism, characterized in that: The method comprises: S1. Acquire video data, preprocess the video data, and obtain preprocessed video data; S2, intercepting the pre-processed video data by frame to obtain a frame image; obtaining the target human skeleton point data based on the frame image, and merging them in frame sequence to form a time sequence key point set; Based on the set of temporal key points, the action recognition backbone network extracts spatial features and temporal features to obtain the action feature map, and outputs the action category and action probability; S3, compressing the action feature map and inputting it into the image attention network to obtain an attention matrix, and outputting the action category and action probability; The video data is processed by a temporal attention network to obtain action categories and action probabilities; the temporal attention network includes a local branch network and a global branch network, the short-term output sequence obtained after the video data is segmented is input into the local branch network to obtain the probability of the local branch output, and the video data is input into the global branch network as a long-term output sequence to obtain the probability of the global branch output; Based on the action categories and action probabilities obtained by the action recognition backbone network, as well as the action categories and action probabilities obtained by the image attention network and the temporal attention network, the final action categories and final action probabilities are obtained to complete fine-grained action recognition; S4. Extract the time sequence key point set in S2 to obtain a fixed frame key point set, and input the fixed frame key point set into a posture recognition network; the posture recognition network includes multiple classifiers, and obtains a final classification result of the posture based on the correlation coefficients of the multiple classifiers.

2. The method according to claim 1, characterized in that In S3, the image attention network solves the action probability in the following way: The action feature map is subjected to average pooling and maximum pooling operations to obtain two one-dimensional vectors as spatial dimension features; The spatial dimension features are processed by a multi-layer perceptron and a sigmoid function to obtain an attention matrix: M c (F)=σ(MLP(AvgPool(F))+MLP(maxPool(F))) Among them, σ represents the sigmoid function, F represents the action feature map, MLP represents the multi-layer perceptron, and M c represents the attention matrix; A linear layer is set after the sigmoid function layer to obtain the action probability.

3. The method according to claim 1, characterized in that In S3, the processing method of the local branch network is: The short-term output sequence is processed by the first temporal convolution to obtain the local field of view temporal features; After batch normalization of the local field of view temporal features, they are input into the second temporal convolution for processing and then output through the sigmoid function layer. The specific method is as follows: Among them, V s Represents the probability of the local branch output, Conv1D represents one-dimensional temporal convolution, BN represents batch normalization, σ represents the sigmoid function, represents a short-term output sequence; the short-term output sequence refers to a data sequence formed by storing the segmented video data.

4. The method according to claim 1, characterized in that: In S3, the processing method of the global branch network is: Through the two connected fully connected layers, the long-term output sequence is dynamically convolved, and then connected to the sigmoid function layer output: Among them, V L represents the probability of global branch output, FC represents the fully connected layer, W1 represents the dynamic convolution kernel weight of the first fully connected layer, and W2 represents the dynamic convolution kernel weight of the second fully connected layer. represents a long-term output sequence, σ represents a sigmoid function; and the video data is used as a long-term output sequence.

5. The method according to claim 1, characterized in that In S4, the correlation coefficient is calculated as follows: Based on the output results of each classifier, a result matrix is ​​constructed; Calculate the correlation coefficient using covariance and standard deviation: Among them, Cov represents covariance, σ represents variance, Model represents the set of output results of each classifier, and Label represents the corresponding body shape label of the person to be detected; The final classification results of body posture are: W final =α1C1+α2C2+α3C3+… Among them, C represents the output result of each classifier.

6. The method according to claim 1, characterized in that In S4, when the posture recognition network is trained, the training data set is constructed as follows: Extract the key point category code and key point coordinate set contained in the fixed frame key point set, and store each frame as text format data of [key point category, key point coordinate]; Use one-hot encoding to mark the abnormal posture in each frame, and construct a one-dimensional empty array as the label according to the number of abnormal posture categories; The labeled dataset is used as the training dataset.

7. The method according to claim 1, characterized in that When extracting temporal features, additional encoding is added to the key points of the same part to represent the same part, thereby increasing the dimension of the key point features; pooling operations are performed on the time dimension to reduce the dimension of the image frame.

Citation Information

Patent Citations

  • Depth video behavior identification method and system

    CN110059662A

  • Human body action recognition method fusing attention mechanism and space-time diagram convolutional neural network under security scene

    CN110119703A