A Human Action Recognition Method Based on Long Short-Term Memory Network and Transformer Algorithm

By using long and short-term memory networks and Transformer algorithms in the human body movement recognition method, the target position feature map is extracted and constructed, and the problem of low efficiency in the existing methods is solved, and more efficient human body movement recognition is achieved.

CN119360442BActive Publication Date: 2025-06-17BEIJING PACTERA JINXIN TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411405577.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-06-17
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

The existing human body movement recognition methods have a large amount of data in key point feature maps, resulting in low efficiency in processing of MLP networks.

Method used

Using a method based on long and short-term memory network and Transformer algorithm, the target position characteristics of the target dimensions in the key point heat map are extracted and the position feature map is constructed to reduce the amount of data and improve the recognition efficiency.

Benefits of technology

By reducing the amount of data in the position feature map, the data processing time is reduced and the efficiency of the human body movement recognition method is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360442B_ABST
    Figure CN119360442B_ABST
Patent Text Reader

Abstract

This application relates to a human action recognition method, device, and computer device based on long short-term memory networks and Transformer algorithms. The method includes: obtaining a key point heat map corresponding to image data; the key point heat map contains key point data of a target object; extracting target position features of a target dimension in the key point heat map according to a target recurrent neural network, and constructing a position feature map of the target object based on the target position features and the key point heat map; performing image recognition on the position feature map based on a sequence model and an activation function to obtain an action recognition result of the target object. Using this method can improve the efficiency of the human action recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly to a human action recognition method, device, computer device, computer-readable storage medium, and computer program product based on long short-term memory network and Transformer algorithm. Background Art

[0002] Understanding human behavior is one of the most important tasks in computer vision. By using the human action recognition method to recognize images, the action categories of people can be obtained, thereby providing a data basis for a wide range of applications such as subsequent human-computer interaction, robotics, and game control.

[0003] Currently, for the human action recognition method, based on the GCN network (Graph Convolution Network), the position features of each dimension of the key point data of the target object are extracted, and a key point feature map is constructed according to the position features of each dimension. Then, based on the MLP (Multilayer Perceptron) network, image recognition is performed on the key point feature map to obtain the action recognition result of the target object.

[0004] However, in the current human action recognition method, the GCN network will extract the position features of each dimension of the key point data, resulting in a large amount of data in the key point feature map, and further causing the MLP network to consume a large amount of time to process the key point feature map. Therefore, the efficiency of the current image recognition method is low. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a human action recognition method, device, computer device, computer-readable storage medium, and computer program product based on long short-term memory network and Transformer algorithm.

[0006] In a first aspect, the present application provides a human action recognition method based on long short-term memory network and Transformer algorithm, including:

[0007] Obtain a key point heat map corresponding to the image data; the key point heat map contains the key point data of the target object;

[0008] Extract the target position features of the target dimension from the key point heat map according to the target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map;

[0009] Perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object.

[0010] In one embodiment, extracting the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and constructing the position feature map of the target object based on the target position features and the key point heat map includes:

[0011] Input the key point heat map into the target recurrent neural network, and extract the position features of the full dimension of each key point data through the target recurrent neural network;

[0012] Based on the target recurrent neural network, screen the position features of the full dimension to obtain the target position features of the target dimension;

[0013] Construct a feature matrix map according to each target position feature, and splice the feature matrix map and the key point heat map to obtain the position feature map of the target object.

[0014] In one embodiment, performing image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object includes:

[0015] Divide the position feature map to obtain each initial position feature sub-map, and perform position encoding on each initial position feature sub-map to obtain the position feature sub-map;

[0016] Based on the position relationship between each position feature sub-map, determine the position relationship data;

[0017] Perform image recognition on each position feature sub-map according to the sequence model, the position relationship data, and the activation function to obtain an output result, and determine the action recognition result of the target object based on the output result.

[0018] In one embodiment, performing image recognition on each position feature sub-map according to the sequence model, the position relationship data, and the activation function to obtain an output result includes:

[0019] Determine each position feature sub-map as a patch embedding, and determine the target action category data and the position relationship data as class tokens; the target action category data is the action feature data of the action category to be classified;

[0020] Perform image recognition on the patch embedding and the class tokens according to the sequence model to obtain an initial output result;

[0021] Perform mapping processing on the initial output result according to the activation function to obtain the output result.

[0022] In one embodiment, the output result includes the probability of the target action category, and determining the action recognition result of the target object based on the output result includes:

[0023] Determine whether the probability of the target action category is greater than or equal to a preset probability threshold;

[0024] If the probability is greater than or equal to the probability threshold, determine that the action recognition result of the target object is the target action category;

[0025] If the probability is less than the probability threshold, determine that the action recognition result of the target object is a non-target action category.

[0026] In one embodiment, before extracting the target position features of the target dimension in the key point heat map according to the target recurrent neural network and constructing the position feature map of the target object based on the target position features and the key point heat map, the method further includes:

[0027] Obtain an image sample set; the image sample set includes an image training sample set and an image verification sample set;

[0028] Train the recurrent neural network based on the image training sample set to obtain a trained recurrent neural network, and verify the trained recurrent neural network based on the image verification sample set to obtain a verification result;

[0029] Determine whether the verification result meets a preset training completion condition;

[0030] When the verification result meets the training completion condition, determine the trained recurrent neural network as the target recurrent neural network.

[0031] In one embodiment, the verification result is a loss value. After determining whether the verification result meets a preset training completion condition, the method further includes:

[0032] When the verification result does not meet the training completion condition, determine the gradient value based on the loss value and the chain algorithm;

[0033] Update the parameters of the trained recurrent neural network according to the gradient descent method and the gradient value; the parameters are used to determine the target dimension;

[0034] Execute the step of training the recurrent neural network based on the image training sample set until the verification result meets the training completion condition.

[0035] In one embodiment, obtaining the key point heat map corresponding to the image data includes:

[0036] Obtain initial image data, and preprocess the initial image data according to a preset image format to obtain image data;

[0037] Extract key point data of a target object from the image data based on a human pose estimation model to obtain a key point heat map.

[0038] In a second aspect, the present application further provides a human action recognition device based on a long short-term memory network and a Transformer algorithm, including:

[0039] An acquisition module, configured to acquire a key point heat map corresponding to image data; each key point data of a target object is included in the key point heat map;

[0040] An extraction module, configured to extract target position features of a target dimension from the key point heat map according to a target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map;

[0041] An identification module, configured to perform image identification on the position feature map based on a sequence model and an activation function to obtain an action recognition result of the target object.

[0042] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0043] Acquire a key point heat map corresponding to image data; each key point data of a target object is included in the key point heat map;

[0044] Extract target position features of a target dimension from the key point heat map according to a target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map;

[0045] Perform image identification on the position feature map based on a sequence model and an activation function to obtain an action recognition result of the target object.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0047] Acquire a key point heat map corresponding to image data; each key point data of a target object is included in the key point heat map;

[0048] Extract the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and construct the position feature map of the target object based on the target position features and the key point heat map;

[0049] Perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object.

[0050] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0051] Obtain the key point heat map corresponding to the image data; each key point data of the target object is included in the key point heat map;

[0052] Extract the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and construct the position feature map of the target object based on the target position features and the key point heat map;

[0053] Perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object.

[0054] The above human action recognition method, device, computer device, computer-readable storage medium and computer program product based on the long short-term memory network and the Transformer algorithm obtain the key point heat map corresponding to the image data; each key point data of the target object is included in the key point heat map; extract the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and construct the position feature map of the target object based on the target position features and the key point heat map; perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object. By using this method, the target position features of the target dimension in the key point heat map are extracted through the recurrent neural network, and the position feature map is constructed based on the target position features and the key point heat map, reducing the data volume of the position feature map. Then, image recognition is performed on the position feature map with less data volume based on the sequence model and the activation function to obtain the action recognition result, reducing the data processing time and improving the efficiency of the human action recognition method. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.

[0056] Figure 1 Schematic flowchart of a human action recognition method based on long short - term memory network and Transformer algorithm in an embodiment;

[0057] Figure 2 Schematic flowchart of determining a position feature map in an embodiment;

[0058] Figure 3 Schematic flowchart of determining an action recognition result in an embodiment;

[0059] Figure 4 Schematic flowchart of determining an output result in an embodiment;

[0060] Figure 5 Schematic diagram of a model for determining an initial output result in an embodiment;

[0061] Figure 6 Schematic flowchart of determining a target recurrent neural network in an embodiment;

[0062] Figure 7 Schematic flowchart of updating parameters of a trained recurrent neural network in an embodiment;

[0063] Figure 8 Schematic flowchart of determining an action recognition result based on an output result in an embodiment;

[0064] Figure 9 Schematic flowchart of obtaining a key - point heat map in an embodiment;

[0065] Figure 10 Schematic flowchart of a human action recognition method based on long short - term memory network and Transformer algorithm in another embodiment;

[0066] Figure 11 Structure block diagram of a human action recognition device based on long short - term memory network and Transformer algorithm in an embodiment;

[0067] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0068] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0069] Understanding human behavior is one of the most important tasks in computer vision. By using human action recognition methods to identify images, the action categories of people can be obtained, thus providing a data basis for a wide range of applications such as subsequent human-computer interaction, robotics, and game control.

[0070] Current human action recognition methods extract the position features of each key point data of the target object in each dimension from the image data through a GCN network. For each key point data, the GCN network determines the position features of other key point data in each dimension according to the position features of the key point data in each dimension, determines the dependence relationship between the key point represented by the key point data and other key points, and thus constructs a skeleton feature map of the target object. The GCN network constructs a key point feature map based on the skeleton feature map of the target object and the position features of each key point data in each dimension. Then, the MLP network performs image recognition based on the key point feature map of the target object to obtain the action recognition result of the target object.

[0071] However, in current human action recognition methods, since the key point feature map contains the position features of each key point data in each dimension and the skeleton feature map of the target object. Therefore, the data volume of the key point feature map is large. Furthermore, the MLP network will consume a lot of time to process a large amount of data, and the efficiency is low. And, a lot of time will be consumed in the process of the GCN network constructing the key point feature map to determine the dependence relationship between each key point, further reducing the processing efficiency of the human action recognition method.

[0072] Based on the above traditional technologies, the embodiments of the present application provide a human action recognition method based on a long short-term memory network and a Transformer algorithm. The target position features of the target dimension in the key point heat map are extracted through a recurrent neural network, and a position feature map is constructed based on the target position features and the key point heat map, reducing the data volume of the position feature map. Then, image recognition is performed on the position feature map with less data volume based on a sequence model and an activation function to obtain the action recognition result, reducing the data processing time and improving the efficiency of the human action recognition method. And, the present application uses a human pose estimation model to extract the key point data of the target object from the image data according to fixed key points, avoiding determining the dependence relationship between key points based on the position features of the key point data, and can quickly obtain a key point human body map, further improving the efficiency of the human action recognition method.

[0073] In one embodiment, as Figure 1As shown, a human action recognition method based on a long short-term memory network and a Transformer algorithm is provided. In the embodiments of the present application, this method is described by taking its application to a computer device as an example. The embodiments of the present application do not limit the execution device of the human action recognition method based on the long short-term memory network and the Transformer algorithm, and include the following steps 102 to 106:

[0074] Step 102: Obtain a key point heat map corresponding to the image data.

[0075] Among them, each key point data of the target object is included in the key point heat map.

[0076] In implementation, the computer device obtains image data including the target object, and extracts each key point data of the target object in the image data based on the human pose estimation model to obtain the key point heat map.

[0077] Specifically, the computer device obtains the initial image data, preprocesses the initial image data to obtain the image data. Then, the computer device extracts the key point data of the target object in the image data according to the human pose estimation model to obtain the key point heat map.

[0078] In an optional embodiment, the computer device obtains the video data of the target object, transcodes the video data to obtain each image data. Then, for each image data, the computer device extracts the key point data of the target object in the image data according to the human pose estimation model to obtain the key point heat map.

[0079] Step 104: Extract the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map.

[0080] In implementation, the computer device obtains an image sample set, and determines the target recurrent neural network based on the image sample set and the recurrent neural network. Then, the computer device inputs the key point heat map into the target recurrent neural network, and extracts the target position features of the target dimension in the key point heat map through the target recurrent neural network. Then, the computer device constructs a position feature map of the target object according to the target position features and the key point heat map.

[0081] Specifically, the computer obtains an image sample set and trains a recurrent neural network based on the image sample set to obtain a target recurrent neural network that meets the training completion condition. Then, the computer device inputs the key point heat map into the target recurrent neural network, extracts the position features of each position through the target recurrent neural network, and filters the target position features of the target dimension from the position features. Then, the computer device constructs a position feature map of the target object based on the target position features and the key point heat map.

[0082] Step 106, perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object.

[0083] In implementation, the computer device preprocesses the position feature map to obtain each position feature sub-map. Then, the computer device performs image recognition on each position feature sub-map according to the sequence model and the activation function to obtain the image recognition result of the target object.

[0084] Specifically, the computer device performs division processing and position encoding processing on the position feature map to obtain each position feature sub-map. Then, the computer device determines the position relationship between each position feature sub-map, thereby generating position relationship data. The computer device performs image recognition on each position feature sub-map according to the position relationship data, the sequence model, and the activation function to obtain the output result of the target object, and determines the action recognition result of the target object based on the output result.

[0085] In the above human action recognition method based on the long short-term memory network and the Transformer algorithm, the target position features of the target dimension in the key point heat map are extracted through the recurrent neural network, and the position feature map is constructed based on the target position features and the key point heat map, reducing the data volume of the position feature map. Then, image recognition is performed on the position feature map with less data volume based on the sequence model and the activation function to obtain the action recognition result, reducing the data processing time and improving the efficiency of the human action recognition method.

[0086] In an exemplary embodiment, as Figure 2 shown, the specific processing process of step 104 includes steps 202 to 206. Among them:

[0087] Step 202, input the key point heat map into the target recurrent neural network, and extract the position features of the full dimension of each key point data through the target recurrent neural network.

[0088] Among them, the target recurrent neural network is the target LSTM network, and the key point heat map contains 17 key point data of the target object.

[0089] In implementation, the computer device inputs the key-point heat map into the target LSTM network, and extracts the position features of the full-dimensional data of each key point through the target LSTM network.

[0090] Specifically, the computer device inputs the key-point heat map into the target LSTM network, and extracts the position features of each of the 17 key-point data in 226 dimensions through the target LSTM network.

[0091] In an exemplary embodiment, the processing order of 17 key-point data is preset in the target LSTM network, and this processing order represents the order from top to bottom among the 17 key points when the target object is standing normally. The computer device inputs the key-point heat map into the target LSTM network, and extracts the position features of each of the 17 key-point data in 226 dimensions through the target LSTM network and the processing order. By extracting the position features of each key-point data through the processing order, serial modeling of the key-point data is achieved.

[0092] Step 204, based on the target recurrent neural network, screen the position features of the full dimension to obtain the target position features of the target dimension.

[0093] In implementation, for the position features of the full dimension corresponding to each key-point data, the computer device screens the position features of the full dimension through the target recurrent neural network to obtain the target position features of the target dimension.

[0094] Specifically, the parameters in the target recurrent neural network can determine the target position features of the target dimension required for each key-point data. Therefore, the target recurrent neural network in the computer device screens the position features of the full dimension corresponding to each key-point data to obtain the target position features of the target dimension.

[0095] In an alternative embodiment, the computer device contains the key-point heat maps corresponding to each image data. For each key-point heat map, the computer device inputs the key-point heat map into the target LSTM network, and extracts the position features of each of the 17 key-point data in 226 dimensions through the target LSTM network and the processing order. Then, the computer device screens the position features of 226 dimensions through the target LSTM network to obtain the target position features of the target dimension. Since the target recurrent neural network is the target LSTM network, and the target LSTM network processes multiple key-point heat maps corresponding to a video data, there will also be some temporal features in the position features. By extracting the target position features of the target dimension through the target LSTM network, the position of each human key point in each key-point heat map is converted into a richer temporal feature representation.

[0096] Step 206: Construct a feature matrix diagram based on each target position feature, and splice the feature matrix diagram and the key point heat map to obtain the position feature map of the target object.

[0097] In implementation, the computer device constructs a feature matrix diagram based on the target position features corresponding to each key point data. Then, the computer device splices the feature matrix diagram and the key point heat map together to obtain the position feature map of the target object.

[0098] Specifically, the computer device combines the target position features corresponding to each key point data to obtain a 17*64*48 feature matrix diagram. The computer device splices the feature matrix diagram and the key point heat map together according to the correspondence between the target dimension of the feature matrix diagram and the dimension of the key point heat map to obtain the position feature map of the target object.

[0099] In an exemplary embodiment, the computer device contains the target position features of each key point heat map. For each key point heat map, the computer device combines the target position features corresponding to the key point heat map to obtain a 17*64*48 feature matrix diagram. The computer device splices the feature matrix diagram and the key point heat map together according to the correspondence between the target dimension of the feature matrix diagram and the dimension of the key point heat map to obtain the position feature map of the target object.

[0100] In this embodiment, the key point data in the key point heat map is modeled by a target recurrent neural network to serialize each key point data and extract the dynamic information and long-term dependencies in the sequence. At the same time, the target recurrent neural network can also help the sequence model learn more robust feature representations and reduce errors caused by factors such as occlusion and perspective changes. In addition, serializing and modeling the human key point data by the target recurrent neural network enables the sequence model to extract additional temporal features from the two-dimensional pose data, which helps to enhance the feature representation of the key point heat map and thus improve the accuracy of action recognition.

[0101] In an exemplary embodiment, as Figure 3 shown, the specific processing process of step 106 includes steps 302 to 306. Among them:

[0102] Step 302: Divide the position feature map to obtain each initial position feature sub-map, and perform position encoding on each initial position feature sub-map to obtain the position feature sub-map.

[0103] In implementation, an initial position feature sub - map format and a position encoding algorithm are pre - set in the computer device. The computer device divides the position feature map according to the initial position feature sub - map format to obtain each initial position feature sub - map. Then, the computer device performs position encoding on each initial position feature sub - map according to the pre - set position encoding algorithm to obtain the position feature sub - map corresponding to the target object.

[0104] In an exemplary embodiment, there are multiple position feature maps in the computer device. The position encoding algorithm is an absolute position encoding algorithm. For each position feature map, the computer device divides the position feature map according to the initial position feature sub - map format to obtain each initial position feature sub - map corresponding to the position feature map. Then, the computer device performs position encoding processing on each initial position feature sub - map according to the absolute position encoding algorithm to obtain the position feature sub - map corresponding to the target object. Due to the use of the self - attention mechanism in the subsequent sequence model, the sequence model requires additional position encoding to understand the relative positions of words in the sequence. Therefore, performing position encoding on each initial position feature sub - map facilitates the subsequent sequence model to utilize the sequence information of the 17 key - point data in the position feature map. Moreover, by performing position encoding on the initial position feature sub - maps, the spatial position relationship between each position feature sub - map is retained.

[0105] Optionally, the initial position feature sub - map format is determined according to the size of the sequence model and the target image data. The initial position feature sub - map format can be but is not limited to being set to 13px * 13px. The position encoding algorithm can be an absolute position encoding algorithm, a relative position encoding algorithm, a common position encoding algorithm, etc. The embodiments of the present application do not limit the initial position feature sub - map format and the position encoding algorithm here.

[0106] Step 304, determine position relationship data based on the position relationships between the position feature sub - maps.

[0107] In implementation, for each position feature sub - map among the position feature sub - maps, the computer device determines the position relationship between this position feature sub - map and other position feature sub - maps among the position feature sub - maps, and generates position relationship data according to the position relationship.

[0108] In an exemplary embodiment, the position relationship represents the relative position and relative direction between two position feature sub - maps. The computer device contains the position feature sub - maps corresponding to multiple position feature maps. For each position feature sub - map corresponding to each position feature map, the computer device determines the position relationship between each position feature sub - map and other position feature sub - maps. This position relationship is the relative direction and relative position of this position feature sub - map and other position feature sub - maps in the position feature map. The computer device generates position relationship data according to the position relationships.

[0109] Step 306: Perform image recognition on each position feature sub - graph according to the sequence model, positional relationship data, and activation function to obtain an output result, and determine the action recognition result of the target object based on the output result.

[0110] In implementation, the computer device constructs class tokens and patch embeddings based on the positional relationship data and each position feature sub - graph. Then, the computer device processes the class tokens and patch embeddings according to the sequence model to obtain an initial output result, and performs data processing on the initial output result based on the activation function to obtain the output result. The computer device determines the action recognition result of the target object based on the output result and a preset probability threshold.

[0111] Specifically, the output result includes the probability of the target action category. The class token is the classtoken in the neural network, and the patch embedding is the Position Embedding of the neural network. The computer device determines each position feature sub - graph as the Position Embedding and constructs the class token based on the positional relationship data. Then, the computer device performs image recognition processing on the Position Embedding and the class token based on the sequence model to obtain an initial output result, and performs mapping processing on the initial output result according to the activation function to obtain the output result. Then, the computer device judges the magnitude relationship between the probability and the probability threshold to obtain a judgment result. The computer device determines the action recognition result of the target object according to the judgment result.

[0112] In this embodiment, by partitioning the position feature map and performing position encoding, each position feature sub - graph is obtained, and the positional relationship data is determined based on each position feature sub - graph, retaining the spatial positional relationship between each position feature sub - graph. Furthermore, based on the sequence model, the positional relationship data, and the position feature sub - graph, the action recognition result is determined, clarifying the action type of the target object in the image data.

[0113] In an exemplary embodiment, as Figure 4 shown, the specific processing process of performing image recognition on each position feature sub - graph according to the sequence model and activation function in step 306 to obtain the output result includes steps 402 to 406. Among them:

[0114] Step 402: Determine each position feature sub - graph as the patch embedding, and determine the target action category data and the positional relationship data as the class token.

[0115] Among them, the target action category data is the action feature data of the action category to be classified. The patch embedding is the Position Embedding of the deep - learning model. The class token is the class token in the deep - learning model.

[0116] In implementation, the computer device determines each position feature sub - graph as a patch embedding and obtains the target action category data. Then, the computer device concatenates the target action category data and the position relationship data to obtain a class token.

[0117] Specifically, the computer device determines each position feature sub - graph as Position Embedding. When the computer device needs to identify the target action category of the target object in the image data, the user determines the target action category to be identified and inputs the target action category data of this target action category into the computer device. The computer device obtains the target action category data. Then, the computer device concatenates the position relationship data and the target action category data to obtain a class token.

[0118] For example, the computer device determines each position feature sub - graph as Position Embedding. The user determines that the target action category to be identified for the target object is falling, and inputs the falling data of the fall into the computer device. The computer device obtains the falling data and concatenates the position relationship data and the falling data to obtain a class token.

[0119] In an exemplary embodiment, when the user needs to identify multiple target action categories of the target user, the user inputs the target action category data of multiple target action categories into the computer device. The computer device obtains each target action category. Then, for each target action category, the computer device concatenates the target action category and the temporal feature data together to obtain a class token.

[0120] Step 404, perform image recognition on the patch embedding and the class token according to the sequence model to obtain an initial output result.

[0121] In implementation, the computer device inputs the patch embedding and the class token into the sequence model, and performs image recognition processing on the patch embedding and the class token through the sequence model to obtain an initial output result.

[0122] Specifically, the sequence model is a Transformer model (a converter model, a deep neural network model based on the self - attention mechanism). The computer device inputs the patch embedding and the class token into the Transformer model. The Transformer model extracts the position and the connection relationship between each key point of the target object from the patch embedding and the class token, and performs image recognition based on the position and the connection relationship to obtain an initial output result.

[0123] In an exemplary embodiment, a sequence model for recognizing multiple target action categories is set in a computer device, and the computer device includes class tokens for multiple target action categories. For each class token of a target action category, the computer device determines the sequence model corresponding to the class token according to the target action category of the class token. Then, the computer device inputs the class token and the patch embedding into the sequence model corresponding to the class token, and performs image recognition of the target object on the class token and the patch embedding through the sequence model to obtain an initial output classification result. This initial output classification result indicates whether the action performed by the target object is the target action category.

[0124] For example, a sequence model for recognizing falling, normal walking, and riding is set in a computer device, and the computer device includes class tokens for falling, normal walking, and riding. The computer device determines that the sequence model corresponding to the class token of falling is the sequence model for recognizing falling, and inputs the class token of falling and the patch embedding into the sequence model for recognizing falling. Image recognition is performed on the class token of falling and the patch embedding through the sequence model for recognizing falling to obtain an initial output result. This initial output result indicates whether the action performed by the target object in the image data is falling.

[0125] The computer device determines that the sequence model corresponding to the class token of normal walking is the sequence model for recognizing normal walking, and inputs the class token of normal walking and the patch embedding into the sequence model for recognizing normal walking. Image recognition is performed on the class token of normal walking and the patch embedding through the sequence model for recognizing normal walking to obtain an initial output result. This initial output result indicates whether the action performed by the target object in the image data is normal walking.

[0126] The computer device determines that the sequence model corresponding to the class token of riding is the sequence model for recognizing riding, and inputs the class token of riding and the patch embedding into the sequence model for recognizing riding. Image recognition is performed on the class token of riding and the patch embedding through the sequence model for recognizing riding to obtain an initial output result. This initial output result indicates whether the action performed by the target object in the image data is riding.

[0127] Step 406: Perform a mapping process on the initial output result according to an activation function to obtain an output result.

[0128] In implementation, an activation function is set in the computer device. The computer device performs a mapping process on the initial output result according to the activation function to obtain an output result.

[0129] Specifically, the mapping function is the Softmax (flexible maximum) layer of a neural network. The computer device inputs the initial output result into the Softmax layer, and performs a mapping process on the initial output result through the Softmax layer to obtain an output result.

[0130] For example, the initial output result shows that the ratio of the action category to the non-action category is 2:8. The computer device inputs the initial output result into the Softmax layer, and the Softmax layer performs mapping processing on the initial output result, making the sum of the ratios of the action category to the non-action category equal to 1, and obtaining an output result where the ratio of the action category to the non-action category is 0.2:0.8.

[0131] In an optional embodiment, a normalization algorithm and a mapping function are set in the computer device. The normalization algorithm is the Normalization layer of the neural network, and the mapping function is the Softmax layer of the neural network. The computer device inputs the initial output result into the Normalization layer, and the Normalization layer performs normalization processing on the initial output result to obtain the normalized initial output result. Then, the computer device inputs the normalized initial output result into the Softmax layer, and the Softmax layer performs mapping processing on the normalized initial output result to obtain the output result.

[0132] In an exemplary embodiment, Figure 5 is a schematic diagram of the model for determining the initial output result. As Figure 5 shown, the computer device splices the feature matrix map (LSTM Features) and the keypoint heat map (Keypoints feature) to obtain the position feature map of the target object. Then, the computer device performs partitioning and position encoding processing on the position feature map to obtain each position feature sub-map. The computer device constructs a class token (Class Toeken) based on the target action category data and the position relationship data between each position feature sub-map, and determines each position feature sub-map as a patch embedding. The computer device inputs the class token and the patch embedding into the Transformer model to obtain the initial output result, and determines the action recognition result (Class) of the target object based on the initial output.

[0133] In this embodiment, by inputting the patch embedding and the class token into the sequence model, the sequence model can understand the dependency relationship between keypoints under different human actions, determine the initial output result based on the dependency relationship, and then determine the output result based on the initial output result, which is convenient for subsequent determination of the action recognition result of the target object.

[0134] In an exemplary embodiment, as Figure 6 shown, the output result includes the probability of the target action category. The specific processing process of determining the action recognition result of the target object based on the output result in step 306 includes steps 602 to 606. Among them:

[0135] Step 602, determine whether the probability of the target action category is greater than or equal to a preset probability threshold.

[0136] In implementation, a probability threshold is preset in the computer device. The computer device determines whether the probability of the action category is greater than or equal to the preset probability threshold.

[0137] In an exemplary embodiment, the action category is falling. The computer device determines whether the probability of falling is greater than or equal to the preset probability threshold. This probability is the confidence level of whether the target object has performed an action of the target action category. The greater the probability, the higher the confidence level that the target object has performed an action of the target action category; the smaller the probability, the lower the confidence level that the target object has performed an action of the target action category.

[0138] In an alternative embodiment, for each action category, a probability threshold corresponding to the action category can be set, or a single probability threshold can be set uniformly for all action categories. If a probability threshold is set for each action category, the computer device determines the probability threshold corresponding to the action category and determines whether the probability of the action category is greater than or equal to the probability threshold.

[0139] Optionally, the probability threshold can be, but is not limited to, set to 90%, which is determined according to the image recognition requirements. The embodiments of the present application do not limit the probability threshold here.

[0140] Step 604, if the probability is greater than or equal to the probability threshold, determine that the action recognition result of the target object is the target action category.

[0141] In implementation, if the probability is greater than or equal to the probability threshold, it indicates that in the image data, the probability that the target object has performed an action of the target action category is relatively high. Therefore, the computer device determines the action recognition result of the target object as the target action category.

[0142] In an exemplary embodiment, the action category is falling. If the probability is greater than or equal to the probability threshold, it indicates that in the image data, the probability that the target object has fallen is relatively high. Therefore, the computer device determines the action recognition result of the target object as wrestling.

[0143] In an alternative embodiment, the computer device includes probabilities of multiple target action categories. The computer device determines the target action category with the highest probability as the action recognition result of the target object. For example, the computer device includes probabilities of falling, normal walking, and cycling. The probability of falling is 95%, the probability of normal walking is 90%, and the probability of cycling is 95%. The computer device determines that the target action category with the highest probability is cycling and determines the action recognition result of the target object as cycling.

[0144] Step 606, if the probability is less than the probability threshold, determine that the action recognition result of the target object is a non-target action category.

[0145] In implementation, if the probability is less than the probability threshold, it indicates that in the image data, the probability that the target object performs the action of the target action category is small. Therefore, the computer device determines the action recognition result of the target object as a non-target action category.

[0146] In an exemplary embodiment, the action category is falling. If the probability is less than the probability threshold, it indicates that in the image data, the probability that the target object falls is small. Therefore, the computer device determines the action recognition result of the target object as not wrestling.

[0147] In this embodiment, the action recognition result of the target object is determined based on the magnitude relationship between the output result and the probability threshold, clarifying the target action category of the action performed by the target object in the image data.

[0148] In an exemplary embodiment, before extracting the target position features according to the target recurrent neural network, it is necessary to first train the recurrent neural network to obtain the target neural network. As Figure 7 shown, before step 104 is executed, the specific processing procedure of the human action recognition method based on the long short-term memory network and the Transformer algorithm further includes steps 702 to 708. Among them:

[0149] Step 702, obtain an image sample set.

[0150] Among them, the image sample set includes an image training sample set and an image verification sample set.

[0151] In implementation, the computer device obtains an initial image sample set, annotates each initial sample image in the initial image sample set to obtain sample images. The computer device constructs an image sample set according to the sample images. Then, the computer device divides the image sample set according to a preset division ratio to obtain an image training sample set and an image verification sample set.

[0152] Specifically, the computer device obtains an initial image sample set, and the sample images in the initial image sample set include first initial sample images representing the target action category and second initial sample images representing non-target action categories. The target category is the action category to be recognized. The computer device sets a first label for the first initial sample images to obtain first sample images, and sets a second label for the second initial sample images to obtain second sample images. Then, the computer device constructs an image sample set according to the first sample images and the second sample images. Then, the computer device divides the image sample set according to a preset division ratio to obtain an image training sample set and an image verification sample set.

[0153] In an exemplary embodiment, the target action category is falling. The computer device obtains an initial image sample set, and the sample images in the initial image sample set include first initial sample images representing falling and second initial sample images representing non-falling (other action categories). The computer device sets a first label 1 for the first initial sample images to obtain first sample images, and sets a second label 0 for the second initial sample images to obtain second sample images. Then, the computer device constructs an image sample set based on each of the first sample images and each of the second sample images. Then, the computer device divides the image sample set according to the segmentation ratio of image training sample set: image validation sample set = 8:2 to obtain an image training sample set and an image validation sample set.

[0154] In an alternative embodiment, the recurrent neural network model needs to identify multiple target action categories, including falling, normal walking, and cycling. The computer device obtains each initial image sample set. The sample images in the initial image sample set include first initial image samples representing falling, second initial image samples representing normal walking, third initial image samples representing cycling, and fourth initial image samples representing other action categories (non-falling, normal walking, and cycling). Then, the computer device sets a first label 1 for the first initial sample images to obtain first sample images, and sets a second label 0 for the second initial sample images, third initial sample images, and fourth initial sample images to obtain second sample images, third sample images, and fourth sample images. The computer device constructs a falling image sample set based on each of the first sample images, each of the second sample images, each of the third sample images, and the fourth sample images. Then, the computer device divides the falling image sample set according to the segmentation ratio of image training sample set: image validation sample set = 8:2 to obtain a falling image training sample set and a falling image validation sample set.

[0155] The computer device sets a first label 1 for the second initial sample images to obtain second sample images, and sets a second label 0 for the first initial sample images, third initial sample images, and fourth initial sample images to obtain first sample images, third sample images, and fourth sample images. The computer device constructs a normal walking image sample set based on each of the first sample images, each of the second sample images, each of the third sample images, and the fourth sample images. Then, the computer device divides the normal walking image sample set according to the segmentation ratio of image training sample set: image validation sample set = 8:2 to obtain a normal walking image training sample set and a normal walking image validation sample set.

[0156] The computer device sets the first label 1 for the third initial sample image to obtain the third sample image, and sets the second label 0 for the first initial sample image, the second initial sample image, and the fourth initial sample image to obtain the first sample image, the second sample image, and the fourth sample image. The computer device constructs a cycling image sample set based on each first sample image, each second sample image, each third sample image, and the fourth sample image. Then, the computer device divides the cycling image sample set according to the segmentation ratio of image training sample set: image verification sample set = 8:2 to obtain a cycling image training sample set and a cycling image verification sample set.

[0157] Optionally, the first label can be set to 1, and the second label can be set to 0. The embodiments of the present application do not limit the first label and the second label.

[0158] Optionally, the segmentation ratio can be but is not limited to image training sample set: image verification sample set = 8:2, which is determined according to the training requirements of the recurrent neural network. The embodiments of the present application do not limit the segmentation ratio.

[0159] Step 704, train the recurrent neural network based on the image training sample set to obtain a trained recurrent neural network, and verify the trained recurrent neural network based on the image verification sample set to obtain a verification result.

[0160] In implementation, the computer device is pre-set with a recurrent neural network. The computer device inputs the sample images in the image training sample set into the recurrent neural network, and performs image recognition on the sample images through the recurrent neural network to obtain a trained recurrent neural network. Then, the computer device inputs the sample set in the image verification sample set into the trained recurrent neural network to obtain an output result. Then, the computer device performs data operations on the output result and the labels of each sample image in the image verification sample set according to the verification algorithm to obtain a verification result.

[0161] Specifically, the recurrent neural network is an LSTM (Long Short-Term Memory) network. The verification algorithm is a loss function algorithm. The computer device inputs the sample images in the image training sample set into the LSTM network, and performs image recognition on the sample images through the LSTM network to obtain a trained LSTM network. Then, the computer device inputs the sample set in the image verification sample set into the trained LSTM network to obtain an output result. This output result is the position feature of the target object in the sample image. The computer device calculates each output result and the labels of each sample image in the image verification sample set according to the loss function algorithm to obtain the loss value of the trained LSTM network.

[0162] In an exemplary embodiment, the image sample set is a wrestling image sample set. The wrestling image sample set includes a wrestling image training sample set and a wrestling image verification sample set. The computer device inputs the sample images in the wrestling image training sample set into the LSTM network, performs image recognition on the sample images through the LSTM network, and obtains the trained LSTM network. Then, the computer device inputs the sample set in the wrestling image verification sample set into the trained LSTM network to obtain the output results. The computer device calculates each output result and the labels of each sample image in the wrestling image verification sample set according to the loss function algorithm to obtain the loss value of the trained LSTM network.

[0163] Step 706, determine whether the verification result meets the preset training completion condition.

[0164] In implementation, the computer device is pre-set with the training completion condition of the recurrent neural network. The computer device determines whether the verification result meets the training completion condition of the recurrent neural network. If the verification result meets the training completion condition, the computer device executes the following step 708. If the verification result does not meet the training completion condition, the computer device executes the following step 802.

[0165] Specifically, the verification result is the loss value, and the training completion condition is the loss threshold. The computer device is pre-set with the loss threshold of the recurrent neural network. The computer device determines whether the loss value is less than or equal to the loss threshold. If the loss value is less than or equal to the loss threshold, the computer device determines that the verification result meets the preset training completion condition. If the loss value is greater than the loss threshold, the computer device determines that the verification result does not meet the training completion condition.

[0166] In an exemplary embodiment, the loss value is 0.2 and the loss threshold is 0.1. The computer device determines that the loss value is greater than the loss threshold and determines that the verification result does not meet the training completion condition.

[0167] Optionally, the loss threshold can be but is not limited to being set to 0.1, which is determined according to the training requirements of the recurrent neural network. The embodiments of the present application do not limit the loss threshold.

[0168] Step 708, in the case where the verification result meets the training completion condition, determine the trained recurrent neural network as the target recurrent neural network.

[0169] In implementation, in the case where the verification result meets the training completion condition, the computer device takes the trained recurrent neural network as the target recurrent neural network and uses the target recurrent neural network in the subsequent execution process of the human action recognition method. The target recurrent neural network can extract the position features of the target action category.

[0170] Specifically, the recurrent neural network is an LSTM network. When the verification result meets the training completion condition, the computer device uses the trained LSTM network as the target LSTM network.

[0171] In an exemplary embodiment, the computer device trains the LSTM network based on a wrestling image sample set. When the verification result meets the training completion condition, the computer device determines the trained LSTM network as the target LSTM network and determines that the target LSTM network can extract the position features of the target object during wrestling.

[0172] In an exemplary embodiment, the computer device contains a wrestling image sample set, a normal walking image sample set, and a cycling image sample set. The computer device trains a recurrent neural network based on the wrestling image sample set to obtain a first target recurrent neural network that meets the training completion condition. At this time, the first target recurrent neural network can accurately extract the position features of the target object during wrestling in the image data. Then, the computer device trains the first target recurrent neural network based on the normal walking image sample set to obtain a second target recurrent neural network that meets the training completion condition. At this time, the second target recurrent neural network can accurately extract the position features of the target object during normal walking and wrestling in the image data. Then, the computer device trains the second target recurrent neural network based on the cycling image sample set to obtain a third target recurrent neural network that meets the training completion condition. At this time, the third target recurrent neural network can accurately extract the position features of the target object during cycling, normal walking, and wrestling in the image data.

[0173] In this embodiment, the recurrent neural network is trained through the image sample set to obtain a target recurrent neural network that meets the training completion condition, improving the accuracy of the target recurrent neural network. Furthermore, based on the target recurrent neural network to extract the target position features, the accuracy of the target position features is improved, and the accuracy of the human action recognition method is improved.

[0174] In an exemplary embodiment, when the verification result does not meet the training completion condition, the parameters of the recurrent neural network need to be updated. The verification result is a loss value. As Figure 8 shown, before step 706 is executed, the specific processing procedure of the human action recognition method based on the long short-term memory network and the Transformer algorithm further includes steps 802 to 806. Among them:

[0175] Step 802, when the verification result does not meet the training completion condition, determine the gradient value based on the loss value and the chain algorithm.

[0176] In implementation, a chain algorithm is preset in the computer device. When the verification result does not meet the training completion condition, the computer device performs data operations on the loss value according to the preset chain algorithm to obtain a gradient value.

[0177] In an exemplary embodiment, in a composite function, the total derivative is equal to the product of the partial derivatives of each function. In a recurrent neural network, the computer device starts from the output layer of the recurrent neural network and calculates the gradient of the loss value with respect to the parameters of each layer layer by layer. Specifically, the recurrent neural network contains an automatic differentiation tool. The computer device automatically processes the loss value through the automatic differentiation tool to obtain the gradient of the recurrent neural network.

[0178] Optionally, the loss function can be but is not limited to mean square error or cross-entropy loss, etc. The present application does not limit the loss function.

[0179] Step 804, update the parameters of the trained recurrent neural network according to the gradient descent method and the gradient value.

[0180] Among them, the parameters are used to determine the target dimension.

[0181] In implementation, the computer device processes the gradient value according to the gradient descent method to update the parameters of the trained recurrent neural network.

[0182] Specifically, the basic formula for parameter update is shown in the following formula (1):

[0183] (1)

[0184] Among them, in the above formula (1), is the model parameter, is the loss function, is the learning rate, is the gradient of the loss function with respect to the parameter.

[0185] During the update process of the parameters of the recurrent neural network, the recurrent neural network will automatically learn the target position features of the target dimension required, so as to facilitate the subsequent screening of the target position features of the target dimension.

[0186] Optionally, the gradient descent method and other optimization algorithms (such as Adam (Adaptive Moment Estimation)) can be but are not limited to be used to update the parameters of the recurrent neural network. The present application does not limit the optimization algorithm.

[0187] Step 806, execute the step of training the recurrent neural network based on the image training sample set until the verification result meets the training completion condition.

[0188] In implementation, the computer device executes the above steps 704 until the verification result of the recurrent neural network meets the training completion condition, and the computer device determines the trained recurrent neural network as the target recurrent neural network. The specific processing procedure of step 704 has been elaborated in detail in the above embodiments, and will not be repeated in this embodiment of the present application.

[0189] In this embodiment, in the case where the verification result does not meet the training completion condition, the parameters of the recurrent neural network are updated based on the loss value, improving the accuracy of the recurrent neural network, and then the target position features of the target dimension are determined based on the parameters of the recurrent neural network, improving the accuracy of the target position features.

[0190] In an exemplary embodiment, as Figure 9 shown, the specific processing procedure of step 102 includes steps 902 to 904. Among them:

[0191] Step 902, obtain initial image data, and preprocess the initial image data according to a preset image format to obtain image data.

[0192] In implementation, the image format is preset in the computer device. The computer device obtains the initial image data. Then, the computer device adjusts the initial image data to image data in the image format.

[0193] Specifically, the computer device is network-connected to an image acquisition device. The image acquisition device acquires the initial image data of the target object and transmits the initial image data to the computer device through the network connection. The computer device obtains the initial image data. Then, the computer device adjusts the format of the initial image data according to the preset image format to obtain image data.

[0194] In an exemplary embodiment, the image acquisition device is a monitoring device. The computer device needs to identify the action category of the target object to determine whether the behavior of the target object is abnormal. The monitoring device acquires the video data of the target object and transmits the video data to the computer device in the form of a video stream. The computer device receives the video data and performs transcoding processing on the video data to obtain the initial image data of each frame. Then, for each frame of the initial image data, the computer device preprocesses the format of the initial image data according to the preset image format to obtain image data, so that the length and width of the image data meet the processing requirements of the human pose estimation model. This image data is an RGB (Red, Green, Blue) image.

[0195] Optionally, the image format meets the input requirements of the human pose estimation model, and can be but is not limited to 256px (pixel) * 192px. The embodiments of the present application do not limit the image format.

[0196] Optionally, the target object is generally a person, including men, women, and children, and the image data may but is not limited to include 1 target object. The embodiments of the present application do not limit the target object.

[0197] Step 904, based on the human pose estimation model, extract the key point data of the target object in the image data to obtain a key point heat map.

[0198] Among them, the key point data is the position data and confidence data of the key points. In human pose estimation, the key points refer to specific parts of the human body used to describe the pose, such as the head, shoulders, elbows, wrists, hips, knees, ankles, etc.

[0199] In implementation, a human pose estimation model is pre-set in the computer device. The computer device inputs the image data into the human pose estimation model, and extracts the key point data of the target image in the image data through the human pose estimation model to obtain a key point heat map.

[0200] Specifically, the human pose estimation model is the ViT Pose (ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation, a simple vision Transformer baseline for human pose estimation, and Transformer is a neural network structure) model. The computer device inputs the image data into the ViT Pose model, and extracts the key point data of the target image in the image data through the human pose estimation model to obtain a key point heat map.

[0201] In an exemplary embodiment, there are multiple frames of image data in the computer device. For each image data, the computer device inputs the image data into the ViT Pose model, performs image recognition on the target object in the image data through the ViT Pose model, and extracts 17 key point data of the target object, so as to construct a key point heat map corresponding to the target object based on the 17 key point data. Among them, the ViT Pose model is a pre-trained ViT Pose model, and this ViT Pose model contains a multi-layer Transformer structure. The pre-trained ViT Pose model with a multi-layer Transformer structure accurately and efficiently predicts the human key point heat map of the target image data.

[0202] In this embodiment, the initial image data is pre-processed, and each key point data of the target object in the pre-processed image data is extracted according to the human pose estimation model to obtain a human key point heat map, which clarifies the relationship between the joints of the target object and facilitates subsequent data processing of the human key point heat map.

[0203] In an exemplary embodiment, a human action recognition method based on a long short-term memory network and a Transformer algorithm is provided, as Figure 10 shown. Figure 10 FIG. is a schematic diagram of the processing flow of a human action recognition method based on a long short-term memory network and a Transformer algorithm in another embodiment. The human action recognition method based on a long short-term memory network and a Transformer algorithm includes:

[0204] Step 1001, obtain initial image data, and preprocess the initial image data to obtain image data. The image data is RGB image data.

[0205] Step 1002, extract key point data of the target object in the image data based on the ViT Pose model to obtain a key point heat map.

[0206] Step 1003, extract target position features of the target dimension in the key point heat map according to the target LSTM model, and splice the key point heat map and each target position feature to obtain a position feature map.

[0207] Step 1004, perform partitioning and position encoding processing on the position feature map to obtain each position feature sub-map, and determine each position feature sub-map as a patch embedding.

[0208] Step 1005, determine position relationship data based on each position feature sub-map, and construct class tokens based on the position relationship data and target action category data.

[0209] Step 1006, perform image recognition on the class tokens and patch embeddings based on a sequence model and an activation function to obtain an output result.

[0210] Step 1007, determine the action recognition result of the target object based on the output result.

[0211] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0212] Based on the same inventive concept, an embodiment of the present application further provides a human action recognition device based on a long short-term memory network and a Transformer algorithm for implementing the above-mentioned human action recognition method based on a long short-term memory network and a Transformer algorithm. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the human action recognition device based on a long short-term memory network and a Transformer algorithm provided below can refer to the limitations on the human action recognition method based on a long short-term memory network and a Transformer algorithm in the above text, and will not be repeated here.

[0213] In an exemplary embodiment, as Figure 11 shown, a human action recognition device 1100 based on a long short-term memory network and a Transformer algorithm is provided, including: an acquisition module 1101, an extraction module 1102, and a recognition module 1103, where:

[0214] The acquisition module 1101 is used to acquire a key point heat map corresponding to image data; each key point data of the target object is included in the key point heat map.

[0215] The extraction module 1102 is used to extract the target position features of the target dimension in the key point heat map according to the target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map.

[0216] The recognition module 1103 is used to perform image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object.

[0217] In an exemplary embodiment, the extraction module 1102 includes:

[0218] The first extraction sub-module is used to input the key point heat map into the target recurrent neural network, and extract the position features of the full dimension of each key point data through the target recurrent neural network.

[0219] The first screening sub-module is used to screen the position features of the full dimension based on the target recurrent neural network to obtain the target position features of the target dimension.

[0220] The first construction sub-module is used to construct a feature matrix map according to each target position feature, and splice the feature matrix map and the key point heat map to obtain the position feature map of the target object.

[0221] In an exemplary embodiment, the recognition module 1103 includes:

[0222] The first sub-module for partitioning is used to partition the position feature map to obtain each initial position feature sub-map, and perform position encoding on each initial position feature sub-map to obtain the position feature sub-map;

[0223] The first determination sub-module is used to determine position relationship data based on the position relationships between the position feature sub-maps;

[0224] The first recognition sub-module is used to perform image recognition on each position feature sub-map according to the sequence model, the position relationship data, and the activation function to obtain an output result, and determine the action recognition result of the target object based on the output result.

[0225] In an exemplary embodiment, the first recognition sub-module includes a second recognition sub-module and a second determination sub-module. Among them, the second recognition sub-module includes:

[0226] The third determination sub-module is used to determine each position feature sub-map as a patch embedding, and determine the target action category data and the position relationship data as class tokens; the target action category data is the action feature data of the action category to be classified.

[0227] The third recognition sub-module is used to perform image recognition on the patch embedding and the class tokens according to the sequence model to obtain an initial output result.

[0228] The first processing sub-module is used to perform mapping processing on the initial output result according to the activation function to obtain the output result.

[0229] In an exemplary embodiment, the output result includes the probability of the target action category. The first recognition sub-module includes a second recognition sub-module and a second determination sub-module. Among them, the second determination sub-module includes:

[0230] The first judgment sub-module is used to judge whether the probability of the target action category is greater than or equal to a preset probability threshold.

[0231] The fourth determination sub-module is used to determine that the action recognition result of the target object is the target action category if the probability is greater than or equal to the probability threshold.

[0232] The fifth determination sub-module is used to determine that the action recognition result of the target object is a non-target action category if the probability is less than the probability threshold.

[0233] In an exemplary embodiment, the human action recognition device 1100 based on the long short-term memory network and the Transformer algorithm further includes:

[0234] The first acquisition module is used to acquire an image sample set; the image sample set includes an image training sample set and an image verification sample set.

[0235] A training module, configured to train a recurrent neural network based on an image training sample set to obtain a trained recurrent neural network, and verify the trained recurrent neural network based on an image verification sample set to obtain a verification result.

[0236] A first judgment module, configured to judge whether the verification result meets a preset training completion condition.

[0237] A first determination module, configured to determine the trained recurrent neural network as a target recurrent neural network when the verification result meets the training completion condition.

[0238] In an exemplary embodiment, the verification result is a loss value, and the human action recognition device 1100 based on the long short-term memory network and the Transformer algorithm further includes:

[0239] A second determination module, configured to determine a gradient value based on the loss value and a chain algorithm when the verification result does not meet the training completion condition.

[0240] An update module, configured to update the parameters of the trained recurrent neural network according to the gradient descent method and the gradient value; the parameters are used to determine a target dimension.

[0241] An execution module, configured to execute the step of training the recurrent neural network based on the image training sample set until the verification result meets the training completion condition.

[0242] In an exemplary embodiment, the acquisition module 1101 includes:

[0243] A second acquisition sub-module, configured to acquire initial image data and preprocess the initial image data according to a preset image format to obtain image data.

[0244] A first extraction sub-module, configured to extract key point data of a target object in the image data based on a human pose estimation model to obtain a key point heat map.

[0245] Each module in the above-mentioned human action recognition device based on the long short-term memory network and the Transformer algorithm can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0246] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a human action recognition method based on a long short-term memory network and a Transformer algorithm. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0247] Those skilled in the art can understand that Figure 12 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0248] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0249] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0250] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0251] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0252] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0253] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A human action recognition method based on long short-term memory network and Transformer algorithm, characterized in that: The method comprises: Obtaining a key point heat map corresponding to the image data; the key point heat map includes key point data of the target object; Extracting target position features of the target dimension in the key point heat map according to the target recurrent neural network, and constructing a position feature map of the target object based on the target position features and the key point heat map; Image recognition is performed on the position feature map based on a sequence model and an activation function to obtain an action recognition result of the target object.

2. The method according to claim 1, characterized in that: The step of extracting target position features of the target dimension in the key point heat map according to the target recurrent neural network, and constructing a position feature map of the target object based on the target position features and the key point heat map, includes: Inputting the key point heat map into the target recurrent neural network, and extracting the full-dimensional position features of each key point data through the target recurrent neural network; Based on the target recurrent neural network, the position features of the full dimension are screened to obtain the target position features of the target dimension; A feature matrix diagram is constructed according to each of the target position features, and the feature matrix diagram and the key point heat map are spliced ​​to obtain a position feature map of the target object.

3. The method according to claim 1, characterized in that The performing image recognition on the position feature map based on the sequence model and the activation function to obtain the action recognition result of the target object includes: Dividing the position feature graph to obtain initial position feature subgraphs, and position encoding each of the initial position feature subgraphs to obtain position feature subgraphs; Determining position relationship data based on the position relationship between each of the position feature subgraphs; Image recognition is performed on each of the position feature subgraphs according to the sequence model, the position relationship data and the activation function to obtain an output result, and an action recognition result of the target object is determined based on the output result.

4. The method according to claim 3, characterized in that The performing image recognition on each of the position feature subgraphs according to the sequence model, the position relationship data and the activation function to obtain an output result includes: Determine each of the position feature subgraphs as a patch embedding, and determine the target action category data and the position relationship data as a class token; the target action category data is the action feature data of the action category to be classified; Performing image recognition on the patch embedding and the class token according to the sequence model to obtain an initial output result; The initial output result is mapped according to the activation function to obtain an output result.

5. The method according to claim 3, characterized in that: The output result includes the probability of the target action category, and the step of determining the action recognition result of the target object based on the output result includes: Determining whether the probability of the target action category is greater than or equal to a preset probability threshold; If the probability is greater than or equal to the probability threshold, determining that the action recognition result of the target object is a target action category; If the probability is less than the probability threshold, it is determined that the action recognition result of the target object is a non-target action category.

6. The method according to claim 1, characterized in that Before extracting the target position features of the target dimension in the key point heat map according to the target recurrent neural network and constructing the position feature map of the target object based on the target position features and the key point heat map, the method further includes: Acquire an image sample set; the image sample set includes an image training sample set and an image verification sample set; Training the recurrent neural network based on the image training sample set to obtain a trained recurrent neural network, and verifying the trained recurrent neural network based on the image verification sample set to obtain a verification result; Determining whether the verification result meets a preset training completion condition; When the verification result satisfies the training completion condition, the trained recurrent neural network is determined to be the target recurrent neural network.

7. The method according to claim 6, characterized in that The verification result is a loss value. After determining whether the verification result meets a preset training completion condition, the method further includes: When the verification result does not satisfy the training completion condition, determining a gradient value based on the loss value and a chain algorithm; Updating parameters of the trained recurrent neural network according to the gradient descent method and the gradient value; the parameters are used to determine the target dimension; The step of training the recurrent neural network based on the image training sample set is performed until the verification result satisfies the training completion condition.

8. The method according to claim 1, characterized in that The step of obtaining a key point heat map corresponding to the image data includes: Acquire initial image data, and preprocess the initial image data according to a preset image format to obtain image data; Based on the human body posture estimation model, key point data of the target object in the image data is extracted to obtain a key point heat map.

9. A human action recognition device based on long short-term memory network and Transformer algorithm, characterized in that: The device comprises: An acquisition module is used to acquire a key point heat map corresponding to the image data; the key point heat map contains key point data of the target object; An extraction module, configured to extract target position features of a target dimension in the key point heat map according to a target recurrent neural network, and construct a position feature map of the target object based on the target position features and the key point heat map; The recognition module is used to perform image recognition on the position feature map based on a sequence model and an activation function to obtain an action recognition result of the target object.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method for recognizing human body behaviors by utilizing posture mask

    CN111274921A

  • Action recognition method based on deep joint and manual apparent feature fusion

    CN113688672A