Boundary sensitive network-based human body action recognition method and device

By extracting frames on video data and constructing timing position feature data, and classifying the video model, the problem of low human body movement recognition efficiency in the prior art is solved, and more efficient action recognition is achieved.

CN120088842APending Publication Date: 2025-06-03BEIJING PACTERA JINXIN TECH LTD
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202411404633.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing human body movement recognition methods are less efficient and take a lot of time due to processing a large amount of image data.

Method used

The human body movement recognition method based on the boundary sensitive network is adopted, and the target image data is obtained by extracting frames on the video data, and the timing position feature data is constructed based on the human body pose estimation model and action category data, and the sequence model is used for classification processing to determine the target action category.

Benefits of technology

By reducing the amount of image data to be processed, the efficiency of the human body motion recognition method is improved and the processing time is shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088842A_ABST
    Figure CN120088842A_ABST
Patent Text Reader

Abstract

The invention relates to a human body action recognition method and device based on a boundary sensitive network and computer equipment. The method comprises the following steps: acquiring video data, and performing frame extraction processing on the video data to obtain target image data; the target image data is image data when the target object acts; constructing time sequence position feature data of the target object based on the human body posture estimation model, the target image data and the action category data; performing classification processing on the time sequence position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data; and determining a target action category of the target object according to each initial classification result. By adopting the method, the efficiency of the human body action recognition method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a human action recognition method, device, computer device, computer-readable storage medium, and computer program product based on a boundary-sensitive network. Background Art

[0002] In multiple applications such as human-computer interaction, robot data, and game control, it is necessary to determine the action category of a person. By using the human action recognition method in the field of computer vision to identify and process the video, the action category of the person can be obtained, providing a data basis for multiple applications.

[0003] In the current human action recognition method, the video data contains continuous frame image data. Feature extraction is performed on each frame of image data through a three-dimensional convolutional network to obtain feature data, and the feature data is spliced into a feature data set. The three-dimensional convolutional network performs prediction processing on the feature data set to obtain the classification result of the action category.

[0004] However, in the current human action recognition method, since the video data contains a large amount of image data, it takes a lot of time to perform feature extraction and prediction processing on a large amount of image data through a three-dimensional convolutional network. Therefore, the efficiency of the current human action recognition method is low. Summary of the Invention

[0005] Based on this, it is necessary to provide a human action recognition method, device, computer device, computer-readable storage medium, and computer program product based on a boundary-sensitive network for the above technical problems.

[0006] In a first aspect, this application provides a human action recognition method based on a boundary-sensitive network, including:

[0007] Obtain video data, and perform frame extraction processing on the video data to obtain target image data; the target image data is the image data of the target object performing an action;

[0008] Construct the temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data;

[0009] Perform classification processing on the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data;

[0010] Determine the target action category of the target object according to each initial classification result.

[0011] In one of the embodiments, the performing frame extraction processing on the video data to obtain target image data includes:

[0012] Perform image recognition on the video data to obtain the action start sequence, action end sequence, and action probability sequence of the target object; the video data includes image data of each frame.

[0013] Based on the action start sequence and the action end sequence, screen each candidate target image data in each frame of the image data.

[0014] Based on the action probability sequence and a preset action probability threshold, screen the target image data from each of the candidate target image data.

[0015] In one embodiment, constructing the temporal position feature data of the target object based on the human pose estimation model, the target image data, and the action category data includes:

[0016] Based on the human pose estimation model and the target image data, determine each key point heat sub-map corresponding to the target object.

[0017] Construct class tokens according to each key point heat sub-map, the target image data, and the action category data.

[0018] Concatenate the class tokens and each key point heat sub-map to obtain the temporal position feature data corresponding to the target object.

[0019] In one embodiment, based on the human pose estimation model and the target image data, determining each key point heat sub-map corresponding to the target object includes:

[0020] Extract the key point data of the target object in the target image data based on the human pose estimation model to obtain the key point heat map corresponding to the target object.

[0021] Divide the key point heat map to obtain each initial key point heat sub-map.

[0022] Perform position encoding on each of the initial key point heat sub-maps to obtain each key point heat sub-map corresponding to the target object.

[0023] In one embodiment, constructing class tokens according to each key point heat sub-map, the target image data, and the action category data includes:

[0024] Determine the position relationship data according to the position relationship between each key point heat sub-map.

[0025] Determine the position of the target image data in the video data to obtain the temporal position data of the target image data.

[0026] Construct class tokens based on the position relationship data, the temporal position data, and the action category data; the action category data is the action feature data of the action categories to be recognized.

[0027] In one embodiment, the classifying the temporal position feature data according to the sequence model to obtain the initial classification result corresponding to the target image data includes:

[0028] Classify the temporal position feature data through a sequence model to obtain an output result; the output result includes the probability of the action type.

[0029] Judge whether the probability of the action category is greater than or equal to a preset probability threshold to obtain a judgment result.

[0030] Determine the initial classification result corresponding to the target image data based on the judgment result.

[0031] In one embodiment, the determining the target action category of the target object according to each of the initial classification results includes:

[0032] Filter each of the initial classification results to obtain each initial target classification result; the initial target classification result represents the action category.

[0033] Determine the proportion of each action category according to each of the initial target classification results.

[0034] Determine the action category with the highest proportion as the target action category.

[0035] In a second aspect, the present application further provides a human action recognition device based on a boundary-sensitive network, including:

[0036] A first processing module, configured to obtain video data and perform frame extraction processing on the video data to obtain target image data; the target image data is the image data of a target object performing an action.

[0037] A construction module, configured to construct the temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data.

[0038] A second processing module, configured to classify the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data.

[0039] A determination module, configured to determine the target action category of the target object according to each of the initial classification results.

[0040] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0041] Obtain video data, and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of a target object performing an action;

[0042] Construct the temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data;

[0043] Perform classification processing on the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data;

[0044] Determine the target action category of the target object according to each of the initial classification results.

[0045] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0046] Obtain video data, and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of a target object performing an action;

[0047] Construct the temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data;

[0048] Perform classification processing on the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data;

[0049] Determine the target action category of the target object according to each of the initial classification results.

[0050] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0051] Obtain video data, and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of a target object performing an action;

[0052] Construct the temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data;

[0053] Perform classification processing on the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data;

[0054] Determine the target action category of the target object according to each of the initial classification results.

[0055] The above-mentioned human action recognition method, device, computer device, computer-readable storage medium and computer program product based on the boundary-sensitive network obtain video data, perform frame extraction processing on the video data to obtain target image data; the target image data is the image data of the target object performing an action; construct the temporal position feature data of the target object based on the human pose estimation model, the target image data and the action category data; perform classification processing on the temporal position feature data according to the sequence model to obtain the initial classification result corresponding to the target image data; determine the target action category of the target object according to each of the initial classification results. By using this method, through frame extraction processing on the video data, the target image data of the target object performing an action is obtained, and the temporal position feature data of the target object is constructed according to the human pose estimation model and the action category data, obtaining the temporal features and position features of the target object, and reducing the data volume of the target image data to be processed. Then, through the sequence model, classification processing is performed on the temporal position feature data to obtain the initial classification result, and the target action category of the target object is determined according to each initial classification result, improving the efficiency of the human action recognition method. Brief Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of a human action recognition method based on a boundary-sensitive network in an embodiment;

[0058] Figure 2 It is a schematic flowchart of determining target image data in an embodiment;

[0059] Figure 3 It is a schematic flowchart of constructing temporal position feature data in an embodiment;

[0060] Figure 4 It is a schematic flowchart of determining each key point heat sub-map in an embodiment;

[0061] Figure 5 It is a schematic flowchart of constructing class tokens in an embodiment;

[0062] Figure 6Schematic flowchart of determining an initial classification result in an embodiment;

[0063] Figure 7 Schematic flowchart of determining a target action category in an embodiment;

[0064] Figure 8 Schematic flowchart of determining target image data in an embodiment;

[0065] Figure 9 Structural block diagram of a human action recognition device based on a boundary-sensitive network in an embodiment;

[0066] Figure 10 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0067] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0068] In multiple applications such as human-computer interaction, robot data, and game control, it is necessary to determine the action category of a person. The video data is recognized and processed by a human action recognition method in the field of computer vision to obtain the action category of the person, providing a data basis for multiple applications.

[0069] In the traditional technology, a computer device acquires video data. The video data is constructed from a series of consecutive frame image data. The computer device inputs all the image data in the video data into a three-dimensional convolutional network, models the temporal information of the video data through the three-dimensional convolutional network to obtain a time series dimension, and extracts features from each frame of image data in the video data based on the time series dimension to obtain respective feature data. Then, the computer device splices the respective feature data into a feature data set through the three-dimensional convolutional network and performs prediction processing on the feature data set to obtain a classification result of the action category.

[0070] However, in the current human action recognition method, since the video data contains a large amount of image data, it takes a lot of time to extract features and perform prediction processing on all the image data in the video data through the three-dimensional convolutional network. Therefore, the efficiency of the current human action recognition method is low.

[0071] Based on the above traditional technologies, an embodiment of the present application provides a human action recognition based on a boundary-sensitive network. By performing frame extraction on video data, target image data of the target object performing an action is obtained, reducing the amount of data that the human pose estimation model and the sequence model need to process. Then, based on the human pose estimation model and the sequence model, an initial classification result of the target image data is determined, and based on each initial classification result, the target action category of the target object is determined, improving the efficiency of the human action recognition method.

[0072] In one embodiment, as Figure 1 shown, a human action recognition method based on a boundary-sensitive network is provided. In the embodiment of the present application, taking the application of this method to a computer device as an example for illustration, the embodiment of the present application does not limit the execution device of the human action recognition method based on the boundary-sensitive network, including the following steps 102 to step 108:

[0073] Step 102, obtain video data, and perform frame extraction on the video data to obtain target image data.

[0074] Among them, the target image data is the image data of the target object performing an action. The video data contains consecutive frame image data.

[0075] In implementation, the computer device obtains the video data that needs to be recognized. Then, the computer device performs frame extraction on each frame of image data to obtain each target image data.

[0076] Specifically, the computer device is connected to the image acquisition device through a network. The image acquisition device acquires video data and sends the video data to the computer device through the network. The computer device receives the video data. Then, the computer device performs image recognition on each frame of image data in the video data to obtain the action start sequence, action end sequence, and action probability sequence of the target object in the image data. Then, the computer device determines the image data of the target object performing an action in each frame of image data based on the action start sequence, action end sequence, and action probability sequence of the target object, and determines the image data of the target object performing an action as the target image data.

[0077] In an exemplary embodiment, the computer device needs to identify the action category of the target object in the video data, thereby providing a data basis for determining whether the behavior of the target object is abnormal. The monitoring device collects the video data of the target object and sends the video data to the computer device, and the computer device receives the video data sent by the monitoring device. Then, the computer device performs frame extraction processing on the video data based on the BSN network (Boundary Sensitive Network), and obtains the target image data. Specifically, the computer device inputs the video data into the BSN network, and performs image recognition on each frame of image data in the video data through the BSN network, and obtains the action start sequence, action end sequence, and action probability sequence of the target object in the image data. Then, the BSN network in the computer device determines the image data in which the target object performs an action according to the action start sequence, action end sequence, and action probability sequence of the target object in each frame of image data, and determines the image data in which the target object performs an action as the target image data.

[0078] Optionally, the computer device may, but is not limited to, obtain video data from an image acquisition device or a computer-readable storage medium, and the computer device may, but is not limited to, obtain video data through networks, Bluetooth, etc. The embodiments of the present application do not limit the source and path of the video data.

[0079] Optionally, the target object is generally a person, and the person here may be a child, a man, a woman, etc. The number of target objects may, but is not limited to, be 1. The embodiments of the present application do not limit the type of the target object and the number of target objects.

[0080] Step 104, construct the temporal and spatial position feature data of the target object based on the human body pose estimation model, the target image data, and the action category data.

[0081] Among them, the action category data is the action feature data of the action category to be recognized.

[0082] In implementation, the computer device determines the heat submaps of each key point of the target object according to the human body pose estimation model and the target image data, and constructs class tokens according to the heat submaps of each key point, the target image data, and the action category data. The class tokens here are class tokens in the deep learning model. Then, the computer device concatenates the class tokens and the heat submaps of each key point together to obtain the temporal and spatial position feature data of the target object. The temporal and spatial position feature data contains the position feature and temporal feature of the target object in the video data.

[0083] Specifically, the computer device extracts the key point data of the target object in the target image data model based on the human pose estimation model to obtain heat submaps of each key point. Then, the computer device constructs class tokens based on the positional relationships between the heat submaps of each key point, the target image data, and the action category data. The computer device splices the class tokens and the heat submaps of each key point to obtain the temporal position feature data of the target object.

[0084] Step 106: Classify the temporal position feature data according to the sequence model to obtain the initial classification result corresponding to the target image data.

[0085] In implementation, the computer device performs image recognition on the temporal position feature data through the sequence model to obtain the initial output result. Then, the computer device preprocesses the initial output result to obtain the initial classification result corresponding to the image data.

[0086] Specifically, the computer device performs image recognition processing of the target object on the temporal position feature data according to the sequence model to obtain the initial output result. Then, the computer device preprocesses the initial output result according to the normalization algorithm and the activation function to obtain the output result. Then, the computer device determines the initial classification result corresponding to the target image data according to the output result.

[0087] In an exemplary embodiment, multiple trained sequence models are set in the computer device, and each sequence model recognizes one action category. The computer device performs image recognition processing on the temporal position feature data based on the multiple sequence models simultaneously to obtain multiple initial output results. Then, the computer device preprocesses each initial output result according to the normalization algorithm and the activation function to obtain the output result. Then, the computer device determines the initial classification result corresponding to the target image data according to each output result.

[0088] Step 108: Determine the target action category of the target object according to each initial classification result.

[0089] In implementation, the computer device screens the initial target classification results from each initial classification result. Then, the computer device determines the target action category of the target object according to each initial target classification result.

[0090] Specifically, the initial classification results include action categories and non-action categories. The computer device screens the initial target classification results that are action categories from each initial classification result. Then, the computer device determines the action category with the highest proportion as the target action category according to each initial target classification result.

[0091] In the above-mentioned human action recognition method based on the boundary-sensitive network, by performing frame extraction on video data, target image data of the target object performing an action is obtained, and based on the human pose estimation model and action category data, temporal position feature data of the target object is constructed, obtaining the temporal features and position features of the target object, and reducing the amount of target image data to be processed. Then, through a sequence model, the temporal position feature data is classified to obtain an initial classification result, and based on each initial classification result, the target action category of the target object is determined, improving the efficiency of the human action recognition method based on the boundary-sensitive network.

[0092] In an exemplary embodiment, as Figure 2 shown, the specific processing process of performing frame extraction on video data in step 102 to obtain target image data includes steps 202 to 206. Among them:

[0093] Step 202, perform image recognition on the video data to obtain an action start sequence, an action end sequence, and an action probability sequence of the target object.

[0094] Among them, the video data contains each frame of image data.

[0095] In implementation, the computer device performs image recognition on the video data to obtain an action start sequence, an action end sequence, and an action probability sequence of the target object. The action start sequence is the probability that an action starts for each frame of image data in the video data, that is, the action start probability. The action end sequence is the probability that the action ends in each frame of image data in the video data. The action probability sequence is the probability that an action exists in each frame of image data in the video data, that is, the action end probability.

[0096] Specifically, the BSN network includes a temporal evaluation module, a nomination generation module, and a nomination evaluation module. The temporal evaluation module includes three layers of temporal convolutional layers. The computer device inputs the video data into the BSN network, and through the temporal evaluation module in the BSN network, an image feature sequence (temporal action region) in the video data is extracted. The image feature sequence includes an action start sequence, an action end sequence, and an action probability sequence.

[0097] In an exemplary embodiment, the computer device inputs the video data into the BSN network. The three layers of temporal convolutional layers in the BSN network simultaneously model the probability that an action starts, the probability that an action ends, and the probability that an action is being performed on each frame of image data in the video data, generating an action start probability sequence, an action end probability sequence, and an action probability sequence, providing a data basis for the subsequent nomination generation module.

[0098] Step 204, based on the action start sequence and the action end sequence, screen each candidate target image data in each frame of image data.

[0099] In implementation, the computer device determines a target action start probability and a target action end probability in an action start sequence and an action end sequence according to a preset node action probability threshold. Then, the computer device filters each candidate target image data from the frame image data according to the target action start probability and the target action end probability.

[0100] Specifically, a nomination generation module is included in the BSN network. A node action probability threshold is preset in the computer device. The computer device determines a target action start probability in the action start sequence whose action start probability is greater than the preset node action probability threshold based on the nomination generation model. At the same time, the computer device determines a target action end probability in the action end sequence whose action end probability is greater than the preset node action probability threshold based on the nomination generation module. The computer device determines first image data corresponding to the target action start probability and second image data corresponding to the target action end probability in the video data. Then, the computer device uses the image data (including the first image data and the second image data) between the first image data and the second image data in the video data as candidate target image data.

[0101] In an exemplary embodiment, the computer device determines a target action start probability in the action start sequence whose action start probability is greater than the preset node action probability threshold based on the nomination generation model. At the same time, the computer device determines a target action end probability in the action end sequence whose action end probability is greater than the preset node action probability threshold based on the nomination generation module. The nomination generation model in the computer device combines the target action start probability and the target action end probability pairwise to obtain a probability combination. Then, the nomination generation module in the computer device determines each candidate target image data from the video data based on the probability combination and generates a corresponding feature description for the candidate target image data, that is, Boundary-Sensitive Proposal (BSP) feature.

[0102] Optionally, the node action probability threshold is determined according to the screening requirement, and can be, but is not limited to, 90% (percent sign), etc. The embodiments of the present application do not limit the node action probability threshold here.

[0103] Step 206, filter target image data from each candidate target image data based on the action probability sequence and a preset action probability threshold.

[0104] In implementation, an action probability threshold is preset in the computer device. The computer device determines a target action probability in the action probability sequence based on the action probability threshold. Then, the computer device determines the target image data from each candidate target image data according to the target action probability.

[0105] Specifically, an action probability threshold is preset in the computer device. The computer device determines, through the BSN network, a target action probability in the action probability sequence that is greater than or equal to the action probability threshold. Then, the computer device determines, through the BSN network, the target image data corresponding to the target action probability from each candidate target image data.

[0106] In an exemplary embodiment, the computer device inputs video data into the BSN network, performs image recognition on the video sequence in the video data through the BSN network, and simultaneously models the probability of the start of an action, the probability of the end of an action, and the probability of the action in progress at each position in the video sequence, thereby generating an action start probability sequence, an action end sequence, and an action probability sequence. Then, the BSN network in the computer device constructs and generates a temporal action region based on the action start probability sequence, the action end sequence, and the action probability sequence. The BSN network in the computer device determines a target time period in the time period corresponding to the video data according to a preset rule. The rule is as follows: (1) The action probability in the action probability sequence corresponding to the target time period is greater than the action probability threshold; (2) The action start probability at the start time of the target time period is higher than the probabilities of the previous and next moments (i.e., a probability peak); (3) The action end probability at the end time of the target time period is higher than the probabilities of the previous and next moments (i.e., a probability peak). The BSN network in the computer device determines the target video data corresponding to the target time period in the video data and crops the target video data into individual RGB (Red, Green, Blue) frames, thereby obtaining the target video data. The target video data is the region where the target object is most likely to perform an action.

[0107] Optionally, the action probability threshold is determined according to the screening requirements of the target image data. The action probability threshold can be, but is not limited to, 85% or the like. The embodiments of the present application do not limit the action probability threshold.

[0108] In this embodiment, through image recognition of video data, an action start sequence, an action end sequence, and an action probability sequence of the target object are obtained, and each image data in the video data is screened based on the action start sequence, the action end sequence, and the action probability sequence of the target object, so that the human pose estimation model can directly process the target image data where an action occurs, avoiding the human pose estimation model from processing all the image data in the video data. The amount of data of the target image data that the human pose estimation model needs to process is reduced, and the efficiency of the image recognition method is improved.

[0109] In an exemplary embodiment, as Figure 3 shown, the specific processing process of step 104 includes steps 302 to 306. Among them:

[0110] Step 302: Based on the human pose estimation model and the target image data, determine the heat submaps corresponding to each key point of the target object.

[0111] In implementation, the computer device extracts features from the target image data according to the human pose estimation model to obtain the heat map of key points. Then, the computer device divides the heat map of key points and processes the initial heat submaps after division to obtain the heat submaps corresponding to each key point.

[0112] Specifically, the computer device extracts the key point data of the target object in the target image data through the human pose estimation model to obtain the heat map of key points. Then, the computer device divides the heat map of key points to obtain each initial heat submap. The computer device performs position encoding processing on each initial heat submap to obtain the heat submap of key points.

[0113] In an exemplary embodiment, the human pose estimation model is the ViT Pose (ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation, and Transformer is a neural network structure) model. The computer device extracts the key point data of the target object in the target image data according to the ViT Pose model. Then, the computer device divides the heat map of key points into each initial heat submap and performs position encoding on each initial heat submap to obtain the heat submap of key points.

[0114] Step 304: Construct class tokens according to the heat submaps corresponding to each key point, the target image data, and the action category data.

[0115] In implementation, the computer device determines the position relationship data based on the position relationship between each heat submap of key points. Then, the computer device constructs class tokens based on the position of the target image data in the video data, the position relationship data, and the action category data. The class tokens contain the temporal features and position features of the target object.

[0116] Specifically, the computer device determines the position relationship between each heat submap of key points and other heat submaps of key points and generates position relationship data based on this position relationship. Then, the computer device generates the temporal position data of the target image data based on the position of the target image data in the video data. The computer device constructs class tokens according to the heat submaps corresponding to each key point, the target image data, and the action category data, so as to obtain the global information of each heat submap of key points.

[0117] Step 306: Concatenate the class token and the heatmaps of each key point to obtain the temporal position feature data corresponding to the target object.

[0118] In implementation, the computer device concatenates the class token and the heatmaps of each key point together to obtain the temporal feature data corresponding to the target object.

[0119] In an exemplary embodiment, the class token is the class token in the deep learning model, which represents the global information of the input vector. Therefore, the computer device concatenates the class token to the first dimension of the heatmaps of each key point to obtain the temporal feature data of the target object.

[0120] In this embodiment, the human pose estimation model is used to extract the feature of the target image data to obtain the heatmaps of each key point of the target object, thereby clarifying the position feature of the target object. Then, by constructing a class token and concatenating the class token and the heatmaps of each key point, the temporal position feature data corresponding to the target object is obtained, clarifying the position feature and the temporal feature of the target object, adding the temporal feature, and further determining the target action category based on the temporal position feature, thereby improving the accuracy of the target action category.

[0121] In an exemplary embodiment, as Figure 4 shown, the specific processing procedure of step 302 includes steps 402 to 406. Among them:

[0122] Step 402: Based on the human pose estimation model, extract the key point data of the target object in the target image data to obtain the heatmap corresponding to the target object.

[0123] Among them, the human pose estimation model is the ViT Pose model. The key point data is the position data and confidence data of the key points. In human pose estimation, the key points refer to specific parts of the human body used to describe the pose, such as the head, shoulders, elbows, wrists, hips, knees, ankles, etc.

[0124] In implementation, the computer device inputs the target image data into the ViT Pose model. The ViT Pose model performs image recognition on the target object in the target image data and extracts the key point data of the target object to obtain the heatmap. The heatmap represents the pose information of the human body.

[0125] Specifically, the computer device inputs the target image data into the ViT Pose model. The ViT Pose model performs image recognition on the target object in the target image data and extracts 17 key point data of the target object, thereby constructing a key point heat map corresponding to the target object based on the 17 key point data. Among them, the ViT Pose model is a pre-trained ViT Pose model, and this ViT Pose model contains multiple Transformer structures. The pre-trained ViT Pose model with multiple Transformer structures accurately and efficiently predicts the human key point heat map of the target image data.

[0126] In an exemplary embodiment, there are multiple frames of target image data in the computer device. For each target image data, the computer device inputs the target image data into the ViT Pose model. The ViT Pose model performs image recognition on the target object in the target image data and extracts 17 key point data of the target object, thereby constructing a key point heat map corresponding to the target object based on the 17 key point data.

[0127] Step 404: Divide the key point heat map to obtain each initial key point heat sub-map.

[0128] In implementation, a heat sub-map format is pre-set in the computer device. The computer device divides the key point heat map according to the heat sub-map format to obtain each initial key point heat sub-map.

[0129] In an exemplary embodiment, there are multiple key point heat maps in the computer device. For each key point heat map, the computer device divides the key point heat map according to the heat sub-map format to obtain each initial key point heat sub-map corresponding to the key point heat map.

[0130] Optionally, the heat sub-map format is determined according to the sequence model and the size of the target image data. The heat sub-map format can be but is not limited to being set to 16px (pixel) * 16px. The embodiments of the present application do not limit the heat sub-map format here.

[0131] Step 406: Perform position encoding on each initial key point heat sub-map to obtain each key point heat map corresponding to the target object.

[0132] In implementation, a position encoding algorithm is pre-set in the computer device. The computer device performs position encoding on each initial key point heat sub-map according to the pre-set position encoding algorithm to obtain the key point heat map corresponding to the target object.

[0133] In an exemplary embodiment, the position encoding algorithm is an absolute position encoding algorithm. The computer device performs position encoding processing on each initial key-point heat map sub-map corresponding to each key-point heat map according to the absolute position encoding algorithm to obtain the key-point heat map sub-maps corresponding to the target object. Through position encoding, each key-point heat map sub-map is added with learnable position embeddings, facilitating the subsequent sequence model to utilize the sequence information of the 17 key-point data in the key-point heat map.

[0134] Optionally, the position encoding algorithm may be an absolute position encoding algorithm or a relative position encoding algorithm, etc. The embodiments of the present application do not limit the position encoding algorithm.

[0135] In this embodiment, the key-point data of the target object in the target image data is extracted through a human pose estimation model to obtain a key-point heat map, and the key-point heat map is divided and position-encoded to obtain each key-point heat map sub-map, clarifying the position relationship between each key-point, facilitating subsequent sequence models for image recognition.

[0136] In an exemplary embodiment, as Figure 5 shown, the specific processing process of step 304 includes steps 502 to 506. Among them:

[0137] Step 502, determine position relationship data according to the position relationship between each key-point heat map sub-map.

[0138] In implementation, the computer device determines the position relationship between each key-point heat map sub-map in each key-point heat map sub-map and other key-point heat map sub-maps in each key-point heat map sub-map, and generates position relationship data according to the position relationship.

[0139] In an exemplary embodiment, the position relationship represents the relative position and relative direction between two key-point heat map sub-maps. The computer device includes key-point heat map sub-maps corresponding to multiple key-point heat maps. The computer device determines the position relationship between each key-point heat map sub-map corresponding to each key-point heat map and other key-point heat map sub-maps. This position relationship is the relative direction and relative position of this key-point heat map sub-map and other key-point heat map sub-maps in the key-point heat map. The computer device generates position relationship data according to each position relationship.

[0140] Step 504, determine the position of the target image data in the video data to obtain the timing position data of the target image data.

[0141] In implementation, the computer device determines the timing position of the target image data in the video data to obtain the timing position data of the target image data.

[0142] Specifically, the computer device determines the time of the target image data in the video data, thereby generating the timing position data of the target image data. This time is the number of minutes and seconds of the target image data in the video data, or the number of frames of the target image data in the target video data. For example, the target image data is at the 2nd minute and 30th second in the video data, or the image data is at the 240th frame in the video data.

[0143] Step 506, construct class tokens based on the position relationship data, timing position data, and action category data.

[0144] Among them, the action category data is the action feature data of the action category to be recognized.

[0145] In implementation, the computer device obtains the action category data. Then, the computer device splices the position relationship data, timing position data, and action category data to obtain class tokens.

[0146] Specifically, when the computer device needs to recognize the target action category of the target object in the video data, the user determines the recognized action category and inputs the action category data of this action category into the computer device. The computer device obtains the action category data. Then, the computer device splices the position relationship data, timing position data, and action category data to obtain class tokens.

[0147] For example, the user determines that the action category to be recognized by the target object is falling, and inputs the falling data of the fall into the computer device. The computer device obtains the falling data and splices the position relationship data, timing position data, and falling data to obtain class tokens.

[0148] In an exemplary embodiment, when the user needs to recognize multiple action categories of the target user, the user inputs the action category data of multiple action categories into the computer device. The computer device obtains each action category data. Then, for each action category data, the computer device splices the action category data, position relationship data, and timing feature data together to obtain class tokens. Then, the computer device splices each class token and each key point heat sub-map together to obtain the timing position feature data of each action category. The computer device determines the sequence model corresponding to each timing position feature data according to the action category, and determines the initial classification result of the target image data corresponding to this action category based on this sequence model and the timing position feature data.

[0149] In this embodiment, class tokens are constructed from positional relationship data, temporal position data, and action category data, clarifying the global information of each key-point heat sub-map. Moreover, by adding temporal position information to the class tokens, the global information is enriched. By constructing temporal position feature data based on the class tokens, the accuracy of the temporal position feature data is improved, thereby enhancing the accuracy of the human action recognition method based on the boundary-sensitive network.

[0150] In an exemplary embodiment, as Figure 6 shown, the specific processing procedure of step 106 includes steps 602 to 606. Among them:

[0151] Step 602: Classify the temporal position feature data through a sequence model to obtain an output result.

[0152] The output result includes the probability of the action type.

[0153] In implementation, the computer device inputs the temporal position feature data into the sequence model, and the sequence model performs image recognition of the target object on the temporal position feature data to obtain an initial output classification result. Then, the computer device performs normalization processing on the initial output result according to the normalization algorithm, and performs mapping processing on the normalized initial output result according to the activation function to obtain the output result.

[0154] Specifically, the sequence model is a Transformer model (a converter model, a deep neural network model based on the self-attention mechanism), and this Transformer model includes six layers of Transformer structures. The normalization algorithm is the Normalization layer of the neural network, and the mapping function is the Softmax layer of the neural network. The computer device inputs the temporal position feature data into the Transformer model, and the Transformer model extracts the positional and inter-connection relationships between each key point of the target object from the temporal position feature data, and performs image recognition based on the position and connection relationships to obtain the initial output result. The six-layer Transformer structure (Transformer self-attention mechanism) here can better extract the inter-correlation relationships between the 17 key points of the target object and achieve a more accurate recognition effect. The computer device inputs the initial output result into the Normalization layer, and the Normalization layer performs normalization processing on the initial output result to obtain the normalized initial output result. Then, the computer device inputs the normalized initial output result into the Softmax layer, and the Softmax layer performs mapping processing on the normalized initial output result to obtain the output result.

[0155] For example, the initial output result shows that the ratio of the action category to the non-action category is 200:800. The computer device inputs the initial output result into the Normalization layer, and the Normalization layer normalizes the initial output result to obtain the normalized initial output result where the ratio of the action category to the non-action category is 2:8. Then, the computer device inputs the normalized initial output result into the Softmax layer, and the Softmax layer maps the normalized initial output result so that the sum of the ratios of the action category to the non-action category is 1, obtaining an output result where the ratio of the action category to the non-action category is 0.2:0.8.

[0156] In an exemplary embodiment, a sequence model for recognizing multiple action categories is set in the computer device, and the computer device includes temporal position feature data of multiple action categories. For each piece of temporal position feature data, the computer device determines the sequence model corresponding to the temporal position feature data according to the action category of the temporal position feature data. Then, the computer device inputs the temporal position feature data into the sequence model corresponding to the temporal position feature data, and the sequence model performs image recognition of the target object on the temporal position feature data to obtain an initial output classification result. The initial output classification result indicates whether the action performed by the target object is of this action category.

[0157] For example, a sequence model for recognizing falling and normal walking is set in the computer device, and the computer device includes temporal position feature data of falling and normal walking. The computer device determines that the sequence model corresponding to the temporal position feature data of falling is the sequence model for recognizing falling, and inputs the temporal position feature data of falling into the sequence model for recognizing falling. The sequence model for recognizing falling performs image recognition on the temporal position feature data to obtain an initial output result. The initial output result indicates whether the action performed by the target object in the video data is falling. The computer device determines that the sequence model corresponding to the temporal position feature data of normal walking is the sequence model for recognizing normal walking, and inputs the temporal position feature data of normal walking into the sequence model for recognizing normal walking. The sequence model for recognizing normal walking performs image recognition on the temporal position feature data to obtain an initial output result.

[0158] Step 604, determine whether the probability of the action category is greater than or equal to a preset probability threshold to obtain a judgment result.

[0159] During implementation, a probability threshold is preset in the computer device. The computer device determines whether the probability of the action category is greater than or equal to the preset probability threshold to obtain a judgment result. If the probability of the action category is greater than or equal to the probability threshold, the computer device determines that the judgment result is the first judgment result. If the probability of the action category is less than the probability threshold, the computer device determines that the judgment result is the second judgment result.

[0160] In an exemplary embodiment, the action category is a fall. The computer device determines whether the probability of falling is greater than or equal to the preset probability threshold. If the probability of falling is greater than or equal to the probability threshold, the computer device determines that the judgment result is the first judgment result. If the probability of falling is less than the probability threshold, the computer device determines that the judgment result is the second judgment result.

[0161] In an alternative embodiment, for each action category, a probability threshold corresponding to the action category can be set, or a single probability threshold can be uniformly set for all action categories. If a probability threshold is set for each action category, the computer device determines the probability threshold corresponding to the action category and determines whether the probability of the action category is greater than or equal to the probability threshold.

[0162] Optionally, the probability threshold can be, but is not limited to, set to 80%, which is determined according to the image recognition requirements. The embodiments of the present application do not limit the probability threshold here.

[0163] Step 606, determine the initial classification result corresponding to the target image data based on the judgment result.

[0164] During implementation, if the judgment result is the first judgment result, it indicates that in the target image data, the probability that the target object performs the action of this action category is relatively high. Therefore, the computer device determines the initial classification result corresponding to the target image data as the action category. If the judgment result is the second judgment result, it indicates that in the target image data, the probability that the target object performs the action of this action category is relatively low. Therefore, the computer device determines the initial classification result corresponding to the target image data as not this action category, that is, the target object does not perform the action of this action category in the target image data.

[0165] In an exemplary embodiment, the action category is a fall. If the judgment result is the first judgment result, it indicates that in the target image data, the probability that the target object falls is relatively high. Therefore, the computer device determines the initial classification result corresponding to the target image data as a fall. If the judgment result is the second judgment result, it indicates that in the target image data, the probability that the target object falls is relatively low. Therefore, the computer device determines the initial classification result corresponding to the target image data as not a fall.

[0166] In this embodiment, the sequence model is used to classify the temporal position feature data to obtain an output result, and the initial classification result of the target image data is determined according to the output result, initially clarifying the action of the target object in the target image data, which is convenient for subsequent determination of the target action type.

[0167] In an exemplary embodiment, as Figure 7 shown, the specific processing process of step 108 includes steps 702 to 706. Among them:

[0168] Step 702, screen each initial classification result to obtain each initial target classification result.

[0169] Among them, the initial target classification result represents the action category.

[0170] In implementation, the computer device determines the initial classification result representing the action category among each initial classification result as the initial target classification result.

[0171] Specifically, for each initial classification result in each initial classification result, the computer device determines whether the initial classification result represents the action category. If the initial classification result represents the action category, the computer device determines the initial classification result as the initial target classification result.

[0172] Step 704, determine the proportion of each action category according to each initial target classification result.

[0173] In implementation, the computer device determines the action category represented by each initial target classification result. Then, the computer device determines the proportion of each action category according to each action category.

[0174] In an exemplary embodiment, there are 5 initial target classification results in the computer device, namely the initial target classification result A, the initial target classification result B, the initial target classification result C, the initial target classification result D, and the initial target classification result E. The action represented by the initial target classification result A is falling; the action represented by the initial target classification result B is falling; the action represented by the initial target classification result C is falling; the action represented by the initial target classification result D is falling; the action represented by the initial target classification result E is normal walking. The computer device determines that the proportion of normal walking is 20%, and determines that the proportion of falling is 80%.

[0175] Step 706, determine the action category with the highest proportion as the target action category.

[0176] In implementation, the computer device determines the action category with the highest proportion among each proportion. Then, the computer device takes this action category as the target action category.

[0177] In an exemplary embodiment, the proportion of normal walking is 20%, and the proportion of falling is 80%. The computer device determines that the action category with the highest proportion is falling. Then, the computer device determines falling as the target action category corresponding to the target object.

[0178] In an exemplary embodiment, the computer device votes on the initial classification results of each target image data to obtain the proportion of each action category. Then, the computer device determines the action category with the highest proportion as the target action category. This target action category represents the action category that occurs to the target object in the video data and is also the action classification of the video data.

[0179] In an alternative embodiment, the computer device determines whether the proportion of each action category exceeds a preset action category proportion threshold. If the proportion of an action category exceeds the action category proportion threshold, the computer device determines this action category as the initial target action category. If the number of initial target action categories is 1, the computer device determines this initial target action category as the target action category. If the number of initial target action categories is greater than 1 or 0, the computer device determines the action category with the highest proportion as the target action category.

[0180] Optionally, the action category proportion threshold can be, but is not limited to, set to 50%, which is determined according to the image recognition requirements. The embodiments of the present application do not limit the action category proportion threshold.

[0181] In this embodiment, the target action type of the target object is determined by the proportion of the action types represented by each initial classification result, clarifying the action performed by the target object in the video data. Moreover, determining the target action type based on the initial classification results corresponding to consecutive target image data is more accurate than determining the target action type based on single frames or discontinuous frames, improving the accuracy of the human action recognition method based on the boundary-sensitive network.

[0182] In an exemplary embodiment, a human action recognition method based on a boundary-sensitive network is provided, as Figure 8 shown. Figure 8 is a schematic diagram of the processing flow of the human action recognition method based on the boundary-sensitive network in another embodiment. The human action recognition method based on the boundary-sensitive network includes:

[0183] Step 801, perform frame extraction processing on the video data through the BSN network to obtain target image data.

[0184] Step 802, perform feature extraction on the target image data based on the ViT Pose model to obtain a key point heatmap (Heatmap).

[0185] Step 803: Divide the heatmap of key points and perform position encoding on the divided heatmap of key points to obtain sub-heatmaps of each key point.

[0186] Step 804: Construct class tokens, and splice the class tokens and the sub-heatmaps of each key point to obtain temporal position feature data.

[0187] Step 805: Perform image recognition on the temporal position feature data based on the Transformer model to obtain an initial output result.

[0188] Step 806: Process the initial output result based on the Normalization layer and the Softmax layer to obtain an output result, and determine an initial classification result (class) based on the output result.

[0189] Step 806: Perform a voting process (Vote) on each initial classification result to obtain the target action class (final class) of the target object.

[0190] The human action recognition method based on the boundary-sensitive network effectively uses multi-frame target image data for human action recognition and detection, and achieves high precision and recall. Compared with other methods that use multiple frames and may only intercept incomplete consecutive frames for judgment, resulting in a loss of precision in human action recognition, the human action recognition method based on the boundary-sensitive network can use all video frames containing the specified action class for action recognition, improving the accuracy and efficiency of the human action recognition method based on the boundary-sensitive network.

[0191] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0192] Based on the same inventive concept, an embodiment of the present application further provides a human action recognition device based on a boundary-sensitive network for implementing the above-mentioned human action recognition method based on a boundary-sensitive network. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following image recognition device embodiments can refer to the limitations on the human action recognition method based on a boundary-sensitive network in the above text, and will not be repeated here.

[0193] In an exemplary embodiment, as Figure 9 shown, a human action recognition device 900 based on a boundary-sensitive network is provided, including: a first processing module 901, a construction module 902, a second processing module 903, and a determination module 904, where:

[0194] The first processing module 901 is configured to obtain video data and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of a target object performing an action.

[0195] The construction module 902 is configured to construct temporal position feature data of the target object based on a human pose estimation model, the target image data, and action category data.

[0196] The second processing module 903 is configured to perform classification processing on the temporal position feature data according to a sequence model to obtain an initial classification result corresponding to the target image data.

[0197] The determination module 904 is configured to determine the target action category of the target object according to each initial classification result.

[0198] In an exemplary embodiment, the first processing module 901 includes a first acquisition sub-module and a first processing sub-module. Among them, the first processing sub-module is used for:

[0199] The first recognition sub-module is configured to perform image recognition on the video data to obtain an action start sequence, an action end sequence, and an action probability sequence of the target object; the video data includes image data of each frame.

[0200] The first screening sub-module is configured to screen each candidate target image data from the image data of each frame based on the action start sequence and the action end sequence.

[0201] The second screening sub-module is configured to screen the target image data from each candidate target image data based on the action probability sequence and a preset action probability threshold.

[0202] In an exemplary embodiment, the construction module 902 includes:

[0203] The first determination sub-module is configured to determine each key-point heat sub-map corresponding to the target object based on the human pose estimation model and the target image data.

[0204] The first construction sub-module is configured to construct class tokens according to each key-point heat sub-map, the target image data, and the action category data.

[0205] The first splicing sub-module is configured to splice the class tokens and each key-point heat sub-map to obtain the temporal position feature data corresponding to the target object.

[0206] In an exemplary embodiment, the first determination sub-module includes:

[0207] The first extraction sub-module is configured to extract the key-point data of the target object in the target image data based on the human pose estimation model, and obtain the key-point heat map corresponding to the target object.

[0208] The first division sub-module is configured to divide the key-point heat map to obtain each initial key-point heat sub-map.

[0209] The first encoding sub-module is configured to perform position encoding on each initial key-point heat sub-map to obtain each key-point heat sub-map corresponding to the target object.

[0210] In an exemplary embodiment, the first construction sub-module includes:

[0211] The second determination sub-module is configured to determine position relationship data according to the position relationship between each key-point heat sub-map.

[0212] The third determination sub-module is configured to determine the position of the target image data in the video data to obtain the temporal position data of the target image data.

[0213] The second construction sub-module is configured to construct class tokens based on the position relationship data, the temporal position data, and the action category data; the action category data is the action feature data of the action category to be recognized.

[0214] In an exemplary embodiment, the second processing module 903 includes:

[0215] The second processing sub-module is configured to perform classification processing on the temporal position feature data through a sequence model to obtain an output result; the output result includes the probability of the action type.

[0216] The first judgment sub-module is configured to judge whether the probability of the action category is greater than or equal to a preset probability threshold to obtain a judgment result.

[0217] The fourth determination sub-module is configured to determine the initial classification result corresponding to the target image data based on the judgment result.

[0218] In an exemplary embodiment, the determination module 904 includes:

[0219] A third screening sub-module, configured to screen each initial classification result to obtain each initial target classification result; the initial target classification result characterizes the action category.

[0220] A fifth determination sub-module, configured to determine the proportion of each action category according to each initial target classification result.

[0221] A sixth determination sub-module, configured to determine the action category with the highest proportion as the target action category.

[0222] Each module in the above-mentioned human action recognition device based on the boundary-sensitive network can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0223] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a human action recognition method based on the boundary-sensitive network. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0224] Those skilled in the art can understand, Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0225] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0226] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0227] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0228] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0229] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope recorded in this application.

[0230] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A human action recognition method based on a boundary sensitive network, characterized in that: The method comprises: Acquire video data, and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of the target object taking action; Constructing temporal position feature data of the target object based on a human posture estimation model, the target image data and the action category data; Classify the temporal position feature data according to the sequence model to obtain an initial classification result corresponding to the target image data; The target action category of the target object is determined according to each of the initial classification results.

2. The method according to claim 1, characterized in that The step of performing frame extraction processing on the video data to obtain target image data includes: Performing image recognition on the video data to obtain an action start sequence, an action end sequence and an action probability sequence of the target object; the video data includes image data of each frame; Based on the action start sequence and the action end sequence, screening each candidate target image data in the image data of each frame; The target image data is screened from each of the candidate target image data based on the action probability sequence and a preset action probability threshold.

3. The method according to claim 1, characterized in that The method of constructing the temporal position feature data of the target object based on the human posture estimation model, the target image data and the action category data includes: Based on the human body posture estimation model and the target image data, determining the thermal subgraphs of the key points corresponding to the target object; Constructing a class token according to each of the key point heat sub-graphs, the target image data and the action category data; The class tokens and the heat sub-maps of each key point are concatenated to obtain the temporal position feature data corresponding to the target object.

4. The method according to claim 3, characterized in that The step of determining the key point heat subgraphs corresponding to the target object based on the human body posture estimation model and the target image data includes: Extract key point data of the target object in the target image data based on a human posture estimation model to obtain a key point heat map corresponding to the target object; Dividing the key point heat map to obtain each initial key point heat submap; Position encoding is performed on each of the initial key point heat sub-graphs to obtain each key point heat sub-graph corresponding to the target object.

5. The method according to claim 3, characterized in that: The step of constructing a class token according to each of the key point heat sub-graphs, the target image data and the action category data comprises: Determining positional relationship data according to the positional relationship between the key point thermal sub-graphs; Determine the position of the target image data in the video data to obtain time sequence position data of the target image data; A class token is constructed based on the position relationship data, the temporal position data and the action category data; the action category data is action feature data of the action category that needs to be identified.

6. The method according to claim 1, characterized in that The classifying process of the temporal position feature data according to the sequence model to obtain the initial classification result corresponding to the target image data includes: Classify the time series position feature data through a sequence model to obtain an output result; the output result includes the probability of the action type; Determine whether the probability of the action category is greater than or equal to a preset probability threshold, and obtain a determination result; An initial classification result corresponding to the target image data is determined based on the judgment result.

7. The method according to claim 1, characterized in that Determining the target action category of the target object according to each of the initial classification results includes: Filtering each of the initial classification results to obtain each of the initial target classification results; the initial target classification results represent the action category; Determining the proportion of each action category according to each of the initial target classification results; The action category with the highest ratio is determined as the target action category.

8. A human motion recognition device based on a boundary sensitive network, characterized in that: The device comprises: The first processing module is used to acquire video data and perform frame extraction processing on the video data to obtain target image data; the target image data is image data of the target object taking action; A construction module, used to construct the temporal position feature data of the target object based on the human posture estimation model, the target image data and the action category data; A second processing module is used to classify the temporal position feature data according to the sequence model to obtain an initial classification result corresponding to the target image data; A determination module is used to determine the target action category of the target object according to each of the initial classification results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Behavior identification method and device

    CN112163480A

  • End-to-end multi-view three-dimensional human body posture estimation method and system and storage medium

    CN112560757A

  • Time sequence nomination generation method and device, equipment and medium

    CN114120180A

  • Gesture recognition method and system based on BVP and WiTransform by using WiFi

    CN115393955A

  • Gait feature extraction and training method, gait recognition method and related equipment

    CN115578749A