Action understanding method, device, computer device and storage medium
Through the long-term and short-term memory neural network model, the problem of low accuracy in continuous correlation image recognition in image processing of neural network models is solved, and the timing correlation of action understanding and prediction and warning of dangerous actions are realized.
Patent Information
- Application Number
- CN201910100446.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-01-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-01-31
AI Technical Summary
The existing neural network models have low accuracy in judging continuous correlation images in the field of image processing and lack coherence.
The long-term and short-term memory neural network model is used to analyze the human limb movement images. By extracting key point information and performing image analysis, the timing correlation of action understanding is achieved using its memory.
It improves the recognition accuracy of continuous correlation images, and can predict and alert when identifying dangerous actions, which enhances safety.
Smart Images

Figure CN111507137B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing, and in particular, to a method and device for action understanding, a computer device, and a storage medium. Background Art
[0002] Since the advent of the mathematical method that simulates the actual neural network of humans, people have gradually become accustomed to directly calling this artificial neural network a neural network. Neural networks have broad and attractive prospects in the fields of system identification, pattern recognition, intelligent control, etc. Especially in intelligent control, people are particularly interested in the self-learning function of neural networks, and regard this important feature of neural networks as one of the key keys to solving the problem of the adaptability of controllers in automatic control.
[0003] In the prior art, neural network models have good performance in the field of image processing. By repeatedly training the neural network model with a large number of pictures of the same type, the neural network model learns the ability to recognize one or more image categories. The neural network model can classify the input pictures relatively accurately. However, the understanding of images by the neural network model is independent and not coherent. Therefore, the accuracy of the neural network model in judging continuous and related images is relatively low. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, computer device, and storage medium for action understanding that can understand input images coherently based on the input time sequence of the images.
[0005] To solve the above technical problems, an embodiment of the present invention uses a technical solution: providing an action understanding method, including:
[0006] Obtain a target image to be recognized, where the target image includes a limb action image of a target user;
[0007] Extract key point information in the limb action image in the target image;
[0008] Input the key point information into a preset action analysis model, where the action analysis model is a long short-term memory neural network model that has been pre-trained to a convergence state and is used for image analysis of human limb actions;
[0009] Read the classification result output by the action analysis model, where the classification result includes the understanding information of the limb action image.
[0010] Optionally, the extracting the key point information in the limb action image in the target image includes:
[0011] Input the target image into a preset image extraction model, where the image extraction model is a neural network model that has been pre-trained to a convergent state and is used to extract key point information in the image;
[0012] Read the feature information output by the image extraction model, where the feature information includes the key point information of the limb movement image.
[0013] Optionally, after reading the classification result output by the action analysis model, it includes:
[0014] Feed the classification result back into the input interface of the action analysis model, so that the action analysis model passes the classification result to the next understanding node of action understanding, making the action understanding coherent in time sequence.
[0015] Optionally, the classification result is the prediction result of the human body's limb movement in the future time sequence. After feeding the classification result back into the input interface of the action analysis model, it includes:
[0016] Obtain a preset action mapping list, where the action mapping list records the mapping relationship between action behaviors and risk values;
[0017] Use the classification result as a retrieval condition to search in the action mapping list for the risk value that has a mapping relationship with the action behavior;
[0018] Identify whether the action of the target user in the future time sequence is dangerous according to the risk value. When the action of the target user in the future time sequence is dangerous, execute a preset warning instruction.
[0019] Optionally, the training method of the action analysis model includes;
[0020] Obtain training sample data marked with classification reference information, where the training sample data includes a number of human key point images;
[0021] Input the training sample data into an initialized long short-term memory neural network model to obtain the classification judgment information of the training sample data;
[0022] Compare whether the classification reference information in the same human key point image in the training sample data is consistent with the classification judgment information;
[0023] When the classification reference information is not consistent with the classification judgment information, repeatedly update the weights in the long short-term memory neural network model in a cyclic iteration until the classification reference information is consistent with the classification judgment information.
[0024] Optionally, the plurality of human key point images are temporally coherent, the classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb movement information represented by the human key point image at the next time series node.
[0025] Optionally, when the classification reference information is inconsistent with the classification judgment information, the weights in the long short-term memory neural network model are repeatedly updated iteratively until the classification reference information is consistent with the classification judgment information. After that, it includes:
[0026] Statistical accuracy of the classification judgment information output by the long short-term memory neural network model;
[0027] Compare the accuracy rate with a set first threshold;
[0028] When the accuracy rate is greater than the first threshold, the long short-term memory neural network model is trained to a convergent state.
[0029] To solve the above technical problems, an embodiment of the present invention further provides an action understanding device, including:
[0030] An acquisition module, configured to acquire a target image to be recognized, where the target image includes a limb movement image of a target user;
[0031] An extraction module, configured to extract key point information in the limb movement image of the target image;
[0032] A processing module, configured to input the key point information into a preset action analysis model, where the action analysis model is a long short-term memory neural network model that has been pre-trained to a convergent state and is used for image analysis of human limb movements;
[0033] An execution module, configured to read the classification result output by the action analysis model, where the classification result includes the understanding information of the limb movement image.
[0034] Optionally, the action understanding device further includes:
[0035] A first processing sub-module, configured to input the target image into a preset image extraction model, where the image extraction model is a neural network model that has been pre-trained to a convergent state and is used for extracting key point information in an image;
[0036] A first execution sub-module, configured to read the feature information output by the image extraction model, where the feature information includes the key point information of the limb movement image.
[0037] Optionally, the action understanding device further includes:
[0038] A second processing sub-module, configured to feedback and input the classification result to the input interface of the action analysis model, so that the action analysis model transfers the classification result to the understanding node of the next action understanding, making the action understanding coherent in time sequence.
[0039] Optionally, the action understanding device further includes:
[0040] A first acquisition sub-module, configured to acquire a preset action mapping list, where the action mapping list records the mapping relationship between action behaviors and danger values;
[0041] A first search sub-module, configured to search for the danger value having a mapping relationship with the action behavior in the action mapping list using the classification result as a retrieval condition;
[0042] A second execution sub-module, configured to identify whether the action of the target user in the future time sequence is dangerous according to the danger value, and execute a preset warning instruction when the action of the target user in the future time sequence is dangerous.
[0043] Optionally, the action understanding device further includes:
[0044] A second acquisition sub-module, configured to acquire training sample data marked with classification reference information, where the training sample data includes a plurality of human key point images;
[0045] A third processing sub-module, configured to input the training sample data into an initialized long short-term memory neural network model to obtain classification judgment information of the training sample data;
[0046] A first comparison sub-module, configured to compare whether the classification reference information in the same human key point image in the training sample data is consistent with the classification judgment information;
[0047] A third execution sub-module, configured to, when the classification reference information is inconsistent with the classification judgment information, repeatedly and iteratively update the weights in the long short-term memory neural network model until the classification reference information is consistent with the classification judgment information.
[0048] Optionally, the plurality of human key point images are coherent in time sequence, the classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb action information represented by the human key point image at the next time sequence node.
[0049] Optionally, the action understanding device further includes:
[0050] A fourth processing sub-module, configured to calculate the accuracy rate of the classification judgment information output by the long short-term memory neural network model;
[0051] A second comparison sub-module, configured to compare the accuracy rate with a set first threshold;
[0052] A fourth execution sub-module, configured to, when the accuracy rate is greater than the first threshold, train the long short-term memory neural network model to a convergent state.
[0053] To solve the above technical problems, an embodiment of the present invention further provides a computer device, including a memory and a processor. A computer-readable instruction is stored in the memory. When the computer-readable instruction is executed by the processor, the processor performs the steps of the above-mentioned action understanding method.
[0054] To solve the above technical problems, an embodiment of the present invention further provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the steps of the above-mentioned action understanding method.
[0055] The beneficial effects of the embodiments of the present invention are as follows: When performing user body movement recognition, key point information in the user body movement is extracted to reduce the overall data volume of the image data, reduce the difficulty of subsequent processing, and improve the efficiency of image recognition. Then, an action analysis model trained by a long short-term memory neural network model is used to process the key point information to obtain a classification result of the body movement image. Since the long short-term memory neural network model has memory when processing images, when judging and recognizing continuous user actions, it can remember the processing result of the previous recognition node and perform correlation recognition with the content of the target image currently being processed, making the image recognition have correlation in time series and improving the accuracy rate of the action analysis model for continuous and related image recognition. Description of the Drawings
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0057] Figure 1 It is a basic process schematic diagram of the action understanding method in the embodiments of the present invention;
[0058] Figure 2 It is a process schematic diagram of extracting key point information through a neural network model in the embodiments of the present invention;
[0059] Figure 3 Schematic flowchart for warning users of dangerous actions in an embodiment of the present invention;
[0060] Figure 4 Schematic flowchart for training an action analysis model in an embodiment of the present invention;
[0061] Figure 5 Schematic flowchart for verifying a long short-term memory neural network model in an embodiment of the present invention;
[0062] Figure 6 Schematic diagram of the basic structure of an action understanding device in an embodiment of the present invention;
[0063] Figure 7 Basic structural block diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0064] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0065] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0066] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] Those skilled in the art of the present technology can understand that the "terminal" and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive without transmission, and devices with receiving and transmitting hardware that have the receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices, which may have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which may combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm-top computers or other devices, which are conventional laptop and / or palm-top computers or other devices with and / or including a radio frequency receiver. The "terminal" and "terminal device" used herein may be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to operate locally, and / or in a distributed manner, at any other location on the earth and / or in space. The "terminal" and "terminal device" used herein may also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or may also be a smart TV, a set-top box, and other devices.
[0068] For details, please refer to Figure 1 , Figure 1 which is the basic flowchart of the action understanding method of this embodiment.
[0069] As Figure 1 shown, an action understanding method includes:
[0070] S1100. Obtain a target image to be recognized, where the target image includes a limb movement image of a target user;
[0071] Obtain a target image to be recognized, where the target image is an image of a limb movement of a target user collected. In addition to recording the limb movement image of the user, the target image also includes a background image. In some embodiments, the target image is a video frame image extracted from a continuous video file.
[0072] The target user refers to any one of the human images that appear in the target image in this example, and is not specifically limited to a designated person. However, in some embodiments, when the action understanding method of this embodiment is used for image tracking, in the case of specifying the person to be tracked, the target user specifically refers to the selected person to be tracked.
[0073] In this embodiment, the limb action image includes a face image, a body image, or an overall human body image.
[0074] S1200. Extract the key point information in the limb action image of the target image;
[0075] Extract the key point information in the limb action image of the target image. Among them, the key point information refers to the key point coordinates of the facial features, shoulders, elbows, hands, chest, waist and hips, knees, and feet of the target user in the target image, as well as the coordinates of the connection lines between the above key points.
[0076] In some embodiments, a neural network model is used to extract the key point information. For example, an image extraction model is used to extract the key point information. Among them, the image extraction model is a neural network model that has been trained to a convergent state and is used to extract the key point coordinates in the human body image.
[0077] In this embodiment, the image extraction model can be a convolutional neural network model (CNN) that has been trained to a convergent state. However, the image extraction model can also be: a deep neural network model (DNN), a recurrent neural network model (RNN), or a deformed model of the above three network models.
[0078] When the image extraction model is trained, a large number of human body images are used for training the extraction of key point coordinates. After being trained to a convergent state, it can accurately extract the key point coordinates in the human body image.
[0079] The extracted key point information can generate a key point coordinate matrix. For example, the key point coordinates of each human body part in the key point information are arranged according to a set rule to generate a key point coordinate matrix. However, it is not limited to this. In some embodiments, the extracted key point information is a key point image, and the key point image is composed of human body key points and the connection lines between the key points.
[0080] S1300. Input the key point information into a preset action analysis model. Among them, the action analysis model is a long short-term memory neural network model that has been pre-trained to a convergent state and is used for image analysis of the limb actions of the human body;
[0081] Convert the extracted key point information into a key point coordinate matrix or a key point image, and input the key point coordinate matrix or the key point image into a preset action analysis model. Among them, the action analysis model is a long short-term memory neural network model that has been pre-trained to a converged state and is used for image analysis of human body limb movements. The image analysis method is to analyze the key point coordinate matrix or the key point image.
[0082] The long short-term memory neural network model is a type of time-recurrent neural network, suitable for processing and predicting important events with relatively long intervals and delays in time series. The convolutional layer for feature extraction in the long short-term memory neural network model is defined as a "neuron cell". In the neuron cell, the feature vector extracted at the previous time series node is selectively inherited and carried over to the current key point coordinate matrix or key point image feature extraction. Due to its recursive inheritance characteristics, the action analysis model can learn the relevance of the user's actions at different time series. For example, the limb movement feature x of the target user extracted by the action analysis model at time t t is formed by the action analysis model by supplementing the change feature at time t on the basis of the feature x t- extracted at the previous action understanding time series t - 1. At time t - 1, the target user has a bent-leg action in the target image, and at time t, the target user has a jumping action in the target image. When the action analysis model learns the feature change between the target images at time t and time t - 1, the action analysis model has the ability to predict the user's action change at time t based on time t - 1.
[0083] The action analysis model performs feature extraction and classification on the input key point coordinate matrix or key point image to obtain the meaning expressed by the limb movement of the target user in the target image, or predicts what kind of action or result the action of the target user in the current target image will trigger in the future.
[0084] S1400. Read the classification result output by the action analysis model, where the classification result includes the understanding information of the limb movement image.
[0085] Read the classification result output by the action analysis model. This classification result is the understanding information of the user's action in the icon image by the action analysis model. This understanding information is the meaning expressed by the limb movement of the target user in the target image, or predicts what kind of action or result the action of the target user in the current target image will trigger in the future.
[0086] When recognizing the user's limb movements in the above embodiments, key point information in the user's limb movements is extracted to reduce the overall data volume of the image data, reduce the difficulty of subsequent processing, and improve the efficiency of image recognition. Then, an action analysis model trained by a long short-term memory neural network model is used to process the key point information to obtain the classification result of the limb movement image. Since the long short-term memory neural network model has memory when processing images, when judging and recognizing continuous user actions, it can remember the processing result of the previous recognition node and perform associative recognition with the content of the target image being currently processed, making the image recognition have relevance in time series and improving the accuracy of the action analysis model in recognizing continuous associated images.
[0087] In some embodiments, a neural network model is used to extract key point information from the target image. Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of extracting key point information by the neural network model in this embodiment.
[0088] As Figure 2 shown, Figure 1 the step S1200 shown includes:
[0089] S1211. Input the target image into a preset image extraction model, where the image extraction model is a neural network model that has been pre-trained to a convergent state and is used to extract key point information in the image;
[0090] Use the image extraction model to extract key point information, where the image extraction model is a neural network model that has been trained to a convergent state and is used to extract the key point coordinates in the human body image.
[0091] In this embodiment, the image extraction model can be a convolutional neural network model (CNN) that has been trained to a convergent state, but the image extraction model can also be: a deep neural network model (DNN), a recurrent neural network model (RNN), or a variant model of the above three network models.
[0092] When the image extraction model is trained, a large number of human body images are used for training on key point coordinate extraction. During training, first, the key points in each human body highlight are manually labeled to generate label information, and the human body image is input into the neural network model. The neural network model outputs the classification information of the human body image, and compares whether the classification information is consistent with the label information. When they are inconsistent, the backpropagation algorithm is used to correct the weights of the neural network model so that the classification information tends to be consistent with the label information. Through repeated training with a large number of human body images, after training to a convergent state, the neural network model can accurately extract the key point coordinates in the human body image. At this time, the neural network model is defined as the image extraction model.
[0093] S1212. Read the feature information output from the image extraction model, where the feature information includes the key point information of the limb movement image.
[0094] Read the feature information output from the image extraction model, where the feature information includes the key point information of the limb movement image. The feature information is the feature vector output by the last convolutional layer of the image extraction model.
[0095] The extracted key point information can generate a key point coordinate matrix. For example, the key point coordinates of each human body part in the key point information are arranged according to a set rule to generate a key point coordinate matrix. However, it is not limited to this. In some embodiments, the extracted key point information is a key point image, which is composed of human body key points and the connecting lines between the key points.
[0096] By using a neural network model to extract the key information in the target image, the extraction efficiency and accuracy of the key point information are improved.
[0097] In some embodiments, the classification result calculated at the current moment is used to calculate the limb movement of the target user at a future moment.
[0098] Such as Figure 1 After the S1400 step shown, it includes:
[0099] S1410. Feed the classification result back to the input interface of the action analysis model, so that the action analysis model passes the classification result to the next understanding node of action understanding, making the action understanding coherent in time sequence.
[0100] The long short-term memory neural network model is a type of time-recurrent neural network, suitable for processing and predicting important events with relatively long intervals and delays in time series. The convolutional layer for feature extraction in the long short-term memory neural network model is defined as a "neuron cell". The neuron cell will selectively inherit the feature vector extracted at the previous time sequence node and inherit it to the current key point coordinate matrix or key point image feature extraction. Due to its recursive inheritance feature, the action analysis model can learn the correlation of the user's actions in different time sequences. For example, the limb movement feature x of the target user extracted by the action analysis model at time t t , when predicting the action at the next future understanding time sequence t + 1, the limb movement feature x t is fed back to the input interface of the next understanding time sequence, making the action analysis model's understanding of the user's actions coherent.
[0101] For example, at time t, the target user has a leg-lifting action in the target image, and at time t+1, the target user has a running action in the target image. After the action analysis model learns the feature changes between the target images at time t and time t+1, the action analysis model learns the correlation between the user's leg-lifting and running actions, thereby enabling the action analysis model to have predictive ability.
[0102] In some embodiments, the classification result output by the action analysis model is used to predict the risk level of the user's action, so as to give a warning in advance when the user performs a dangerous action. Please refer to Figure 3 , Figure 3 which is a schematic flowchart of warning the user of dangerous actions in this embodiment.
[0103] As Figure 3 shown, after step S1410, it includes:
[0104] S1421. Obtain a preset action mapping list, where the mapping relationship between action behaviors and risk values is recorded in the action mapping list;
[0105] In this embodiment, there is a preset action mapping list, and the mapping relationship between action behaviors and risk values is recorded in the action mapping list. The value range of the risk value is between 0 and 100, but the value range of the risk value is not limited to this. According to different specific application scenarios, in some embodiments, the representation method of the risk value can be: through language text, color or sound tone.
[0106] The action mapping list defines the risk levels of various actions of the user and defines the risk values of various user behaviors.
[0107] S1422. Use the classification result as a retrieval condition to search in the action mapping list for the risk value that has a mapping relationship with the action behavior;
[0108] After the action analysis model outputs the classification result of the target image, use this classification result as the retrieval keyword, and search in the action mapping list in a traversal manner for the risk value of the user action represented by the classification result. For example, when the action analysis model determines that the walking action of the target user is about to cause the user to fall, the risk value of the user's current walking action is 80, which belongs to a high-risk behavior.
[0109] S1423. Identify whether the action of the target user in the future time series is dangerous according to the risk value. When the action of the target user in the future time series is dangerous, execute a preset warning instruction.
[0110] The risk value retrieved according to the classification result identifies whether the current limb movement of the user will cause a risk in the future time series. The judgment of the risk depends on the magnitude of the risk value. The larger the risk value, the greater the risk of the target user in the future time series. Conversely, the smaller the risk.
[0111] When it is determined that the behavior of the target user will cause a risk in the future time series, a warning is issued to the target user. The judgment of the risk needs to be made through a risk threshold. In some embodiments, the risk threshold is 60. When the risk value is greater than or equal to 60, a warning is given to the target user. However, the setting of the risk threshold is not limited to this. In some embodiments, the risk threshold is set in a gradient manner. When the risk value is in different gradient intervals, different warnings are given to the target user.
[0112] The way to issue a user warning is to execute a preset warning instruction. The execution result of the warning instruction is to remind the user of the possible future risks by voice. However, the warning method is not limited to this. According to the different specific application scenarios, in some embodiments, the user is warned by a warning light or text information. In some embodiments, the warning method varies with the different risk values. The larger the risk value, the higher the warning level.
[0113] By warning, the target user is reminded to avoid unknown risks in the future, which can effectively ensure the safety of the target user.
[0114] In some embodiments, in order to enhance the prediction ability and accuracy of the action analysis model, when training the action analysis model, it is necessary to consciously train the prediction ability of the action analysis model. Please refer to Figure 4 , Figure 4 which is a schematic flow chart for training the action analysis model.
[0115] As Figure 4 shown, the training method of the action analysis model is as follows:
[0116] S1010. Obtain training sample data marked with classification reference information, where the training sample data includes a number of human key point images;
[0117] The training sample data is the constituent unit of the entire training set. The training set is composed of several training sample data. The training sample data includes: several human key point images. The several human key point images are coherent in time series. The classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb movement information represented by the human key point image at the next time series node. That is, the training sample data is several human key point images of a series of continuous human actions, and the several human key point images are arranged in time series.
[0118] The classification reference information is obtained by manually observing several human key point images and then manually annotating each key point image. The annotation result is defined as the classification reference information. The content recorded in the classification reference information is the limb movement of the current human key point image at the next time series node. The interval between time series nodes is 1 second, but the setting of the time interval is not limited to this. According to different specific application scenarios, the time interval can be set shorter or longer. The setting of the time interval determines the length of the prediction time of the action analysis module. For example, if the current human key point image shows that the human has a leg-lifting action, and the human key point image at the next time series node shows that the human has a dumping action, then the classification reference information of the current human key point image is defined as "dumping".
[0119] S1020. Input the training sample data into the initialized long short-term memory neural network model to obtain the classification judgment information of the training sample data;
[0120] Input the training sample set into the long short-term memory neural network model in sequence. The long short-term memory neural network model extracts features and classifies the human key point images. The classification result of each human key point image output by the long short-term memory neural network model is defined as the classification judgment information of the human key point image.
[0121] The classification judgment information is the excitation data output by the long short-term memory neural network model according to the input human key point images. Before the long short-term memory neural network model is trained to convergence, the classification judgment information is a value with a relatively large discreteness. After the long short-term memory neural network model is trained to convergence, the classification judgment information is relatively stable data.
[0122] S1030. Compare whether the classification reference information and the classification judgment information in the same human key point image in the training sample data are consistent;
[0123] The loss function is a detection function configured to detect whether the classification judgment information in the long short-term memory neural network model is consistent with the expected classification reference information. When the classification judgment information output by the long short-term memory neural network model is inconsistent with the expected result of the classification reference information, it is necessary to correct the weights in the long short-term memory neural network model so that the classification judgment information of the long short-term memory neural network model is the same as the classification judgment information.
[0124] S1040. When the classification reference information is inconsistent with the classification judgment information, repeatedly and iteratively update the weights in the long short-term memory neural network model until the classification reference information is consistent with the classification judgment information.
[0125] When the classification judgment information output by the long short-term memory neural network model is inconsistent with the classification reference information, it is necessary to correct the weights in the long short-term memory neural network model so that the classification judgment information of the long short-term memory neural network model is the same as the classification judgment information. When the classification judgment information is the same as the classification judgment information, stop training the human key point image. During training, multiple training sample data are used for training (for example, 100,000 training sample data of continuous actions). Through repeated training and correction, when the classification data output by the long short-term memory neural network model is compared with the classification reference information of each training sample and reaches (not limited to) 96%, the training ends.
[0126] After the training ends, the long short-term memory neural network model trained to the convergence state is defined as the action analysis model.
[0127] Defining the classification reference information of the human key point image at the current moment as the limb movement information at the future moment during training can enhance the prediction ability of the long short-term memory neural network model and improve the accuracy of the classification result predicted by the long short-term memory neural network model.
[0128] In some embodiments, it is necessary to verify the long short-term memory neural network model to determine whether the long short-term memory neural network model is trained to the convergence state. Please refer to Figure 5 , Figure 5 which is the schematic flowchart for verifying the long short-term memory neural network model in this embodiment.
[0129] As Figure 5 shown, Figure 4 after the S1040 step shown, it includes:
[0130] S1051. Statistically calculate the accuracy rate of the classification judgment information output by the long short-term memory neural network model;
[0131] When training the long short-term memory neural network model, the number of accurate times of the classification judgment information output by the long short-term memory neural network model is counted, and at the same time, the number of training times is counted. Then, the accuracy rate of the classification judgment information is calculated according to the ratio of the number of accurate times of the classification judgment information to the number of training times.
[0132] S1052. Compare the accuracy rate with a set first threshold;
[0133] Compare the accuracy rate with the set first threshold. The first threshold is a set value used to measure the accuracy rate of the output of the long short-term memory neural network model. In this embodiment, the first threshold is 95%. However, the set value of the first threshold is not limited to this, and according to different specific application scenarios, the value of the first threshold can be set larger or smaller.
[0134] S1053. When the accuracy rate is greater than the first threshold, the long short-term memory neural network model is trained to a convergent state.
[0135] When the accuracy rate is greater than the first threshold, it is confirmed that the long short-term memory neural network model is trained to a convergent state. At this time, the training of the long short-term memory neural network model to a convergent state is completed.
[0136] To solve the above technical problems, an embodiment of the present invention further provides an action understanding device.
[0137] Specifically, please refer to Figure 6 , Figure 6 which is a schematic diagram of the basic structure of the action understanding device in this embodiment.
[0138] As Figure 6 shown, an action understanding device includes: an acquisition module 2100, an extraction module 2200, a processing module 2300, and an execution module 2400. Among them, the acquisition module 2100 is used to acquire a target image to be recognized, where the target image includes a limb action image of a target user; the extraction module 2200 is used to extract key point information in the limb action image in the target image; the processing module 2300 is used to input the key point information into a preset action analysis model, where the action analysis model is a long short-term memory neural network model that has been pre-trained to a convergent state and is used for image analysis of human limb actions; the execution module 2400 is used to read the classification result output by the action analysis model, where the classification result includes the understanding information of the limb action image.
[0139] When the action understanding device performs user limb movement recognition, it extracts key point information in the user's limb movements to reduce the overall data volume of the image data, reduce the difficulty of subsequent processing, and improve the efficiency of image recognition. Then, it uses an action analysis model trained by a long short-term memory neural network model to process the key point information to obtain the classification result of the limb movement image. Since the long short-term memory neural network model has memory when processing images, when judging and recognizing continuous user actions, it can remember the processing result of the previous recognition node and perform associative recognition with the content of the target image being currently processed, making the image recognition have relevance in time series and improving the accuracy of the action analysis model for continuous associated image recognition.
[0140] In some embodiments, the action understanding device further includes: a first processing sub-module and a first execution sub-module. Among them, the first processing sub-module is used to input the target image into a preset image extraction model, where the image extraction model is a neural network model that has been pre-trained to a convergent state and is used to extract key point information in the image; the first execution sub-module is used to read the feature information output by the image extraction model, where the feature information includes the key point information of the limb movement image.
[0141] In some embodiments, the action understanding device further includes: a second processing sub-module, which is used to feedback and input the classification result to the input interface of the action analysis model, so that the action analysis model transfers the classification result to the next understanding node of action understanding, making the action understanding have coherence in time series.
[0142] In some embodiments, the action understanding device further includes: a first acquisition sub-module, a first search sub-module, and a second execution sub-module. Among them, the first acquisition sub-module is used to acquire a preset action mapping list, where the action mapping list records the mapping relationship between action behaviors and risk values; the first search sub-module is used to search for the risk value mapped to the action behavior in the action mapping list with the classification result as the retrieval condition; the second execution sub-module is used to identify whether the action of the target user in the future time series is dangerous according to the risk value, and when the action of the target user in the future time series is dangerous, execute a preset warning instruction.
[0143] In some embodiments, the action understanding device further includes: a second acquisition sub-module, a third processing sub-module, a first comparison sub-module, and a third execution sub-module. Among them, the second acquisition sub-module is configured to acquire training sample data marked with classification reference information, where the training sample data includes a plurality of human key point images; the third processing sub-module is configured to input the training sample data into an initialized long short-term memory neural network model to obtain classification judgment information of the training sample data; the first comparison sub-module is configured to compare whether the classification reference information in the same human key point image in the training sample data is consistent with the classification judgment information; the third execution sub-module is configured to, when the classification reference information is inconsistent with the classification judgment information, repeatedly and iteratively update the weights in the long short-term memory neural network model until the classification reference information is consistent with the classification judgment information.
[0144] In some embodiments, the plurality of human key point images are temporally coherent, the classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb action information represented by the human key point image at the next time sequence node.
[0145] In some embodiments, the action understanding device further includes: a fourth processing sub-module, a second comparison sub-module, and a fourth execution sub-module. Among them, the fourth processing sub-module is configured to count the accuracy rate of the classification judgment information output by the long short-term memory neural network model; the second comparison sub-module is configured to compare the accuracy rate with a set first threshold; the fourth execution sub-module is configured to, when the accuracy rate is greater than the first threshold, train the long short-term memory neural network model to a converged state.
[0146] To solve the above technical problems, an embodiment of the present invention further provides a computer device. Specifically, please refer to Figure 7 , Figure 7 which is the basic structural block diagram of the computer device in this embodiment.
[0147] As Figure 7As shown, it is a schematic internal structure diagram of a computer device. The computer device includes a processor, a non-volatile storage medium, a memory, and a network interface connected through a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement an action understanding method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute an action understanding method. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand, Figure 7 The structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0148] In this embodiment, the processor is used to execute Figure 6 the specific functions of the acquisition module 2100, the extraction module 2200, the processing module 2300, and the execution module 2400 in the figure. The memory stores the program codes and various types of data required to execute the above modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all sub-modules in the face image key point detection device. The server can call the program codes and data of the server to execute the functions of all sub-modules.
[0149] When the computer device performs user limb movement recognition, it extracts the key point information in the user limb movement to reduce the overall data volume of the image data, reduce the difficulty of subsequent processing, and improve the efficiency of image recognition. Then, an action analysis model trained by a long short-term memory neural network model is used to process the key point information to obtain the classification result of the limb movement image. Since the long short-term memory neural network model has memory when processing images, when judging and recognizing continuous user actions, it can remember the processing result of the previous recognition node and perform correlation recognition with the content of the current target image being processed, making the image recognition have correlation in time sequence and improving the accuracy of the action analysis model for continuous and related image recognition.
[0150] The present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the action understanding method in any of the above embodiments.
[0151] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0152] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
Claims
1. A method for action understanding, characterized in that, Including: Obtain a target image to be recognized, where the target image includes a limb movement image of a target user; Extract key point information in the limb movement image in the target image; Input the key point information into a preset action analysis model, where the action analysis model is pre-trained to a convergence state and is a long short-term memory neural network model for image analysis of human limb movements; Read the classification result output by the action analysis model, where the classification result includes understanding information of the limb movement image; After reading the classification result output by the action analysis model, it includes: Feed the classification result back into the input interface of the action analysis model, so that the action analysis model transfers the classification result to the understanding node of the next action understanding, making the action understanding coherent in time sequence, and the classification result output by the action analysis model is used to predict the risk level of the limb movement; Obtain a preset action mapping list, where the action mapping list records the mapping relationship between action behaviors and risk values; Search for the risk value having a mapping relationship with the action behavior in the action mapping list using the classification result as a retrieval condition; Identify whether the action of the target user in the future time sequence is dangerous according to the risk value, and when the action of the target user in the future time sequence is dangerous, execute a preset warning instruction.
2. The action understanding method according to claim 1, wherein The extracting the key point information in the limb movement image in the target image includes: Input the target image into a preset image extraction model, where the image extraction model is pre-trained to a convergence state and is a neural network model for extracting key point information in an image; Read the feature information output by the image extraction model, where the feature information includes the key point information of the limb movement image.
3. The action understanding method according to any one of claims 1 or 2, characterized in that, The training method of the action analysis model includes; Obtain training sample data marked with classification reference information, where the training sample data includes a number of human key point images; Input the training sample data into an initialized long short-term memory neural network model to obtain classification judgment information of the training sample data; Compare whether the classification reference information in the same human key point image in the training sample data is consistent with the classification judgment information; When the classification reference information is not consistent with the classification judgment information, repeatedly update the weights in the long short-term memory neural network model in a cyclic iteration until the classification reference information is consistent with the classification judgment information.
4. The action understanding method according to claim 3, wherein The number of human key point images is coherent in time sequence, the classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb movement information represented by the human key point image at the next time sequence node.
5. The action understanding method according to claim 3, characterized in that After the weights in the long short-term memory neural network model are repeatedly updated in a cyclic iteration until the classification reference information is consistent with the classification judgment information, it includes: Statistically calculate the accuracy rate of the classification judgment information output by the long short-term memory neural network model; Compare the accuracy rate with a set first threshold; When the accuracy rate is greater than the first threshold, the long short-term memory neural network model is trained to a convergent state.
6. An action understanding device, characterized in that, It includes: An acquisition module for acquiring a target image to be recognized, where the target image includes a limb movement image of a target user; An extraction module for extracting key point information in the limb movement image in the target image; A processing module for inputting the key point information into a preset action analysis model, where the action analysis model is a long short-term memory neural network model that has been pre-trained to a convergent state and is used for image analysis of human limb movements; An execution module for reading the classification result output by the action analysis model, where the classification result includes the understanding information of the limb movement image; The action understanding device further includes: A second processing sub-module for feedback inputting the classification result to the input interface of the action analysis model, so that the action analysis model transfers the classification result to the next understanding node of action understanding, making the action understanding coherent in time sequence, and the classification result output by the action analysis model is used to predict the danger level of the limb movement; A first acquisition sub-module for acquiring a preset action mapping list, where the action mapping list records the mapping relationship between action behaviors and danger values; A first search sub-module for searching in the action mapping list for the danger value having a mapping relationship with the action behavior with the classification result as the retrieval condition; A second execution sub-module for identifying whether the action of the target user in the future time sequence is dangerous according to the danger value, and when the action of the target user in the future time sequence is dangerous, executing a preset warning instruction.
7. The action understanding device according to claim 6, wherein The action understanding device further includes: A first processing sub-module for inputting the target image into a preset image extraction model, where the image extraction model is a neural network model that has been pre-trained to a convergent state and is used for extracting key point information in the image; A first execution sub-module for reading the feature information output by the image extraction model, where the feature information includes the key point information of the limb movement image.
8. The action understanding device according to any one of claims 6 or 7, characterized in that The action understanding device further includes: A second acquisition sub-module for acquiring training sample data marked with classification reference information, where the training sample data includes a plurality of human key point images; A third processing sub-module for inputting the training sample data into an initialized long short-term memory neural network model to obtain the classification judgment information of the training sample data; A first comparison sub-module for comparing whether the classification reference information in the same human key point image in the training sample data is consistent with the classification judgment information; A third execution sub-module, configured to update the weights in the long short-term memory neural network model in a repeated cyclic iteration when the classification reference information is inconsistent with the classification judgment information, and end when the classification reference information is consistent with the classification judgment information.
9. The action understanding device according to claim 8, wherein The plurality of human key point images are temporally coherent, the classification judgment information is the calibration information of each human key point image, and the calibration information of each human key point image is the limb movement information represented by the human key point image at the next time series node.
10. The action understanding device according to claim 8, wherein The action understanding device further includes: A fourth processing sub-module, configured to count the accuracy rate of the classification judgment information output by the long short-term memory neural network model; A second comparison sub-module, configured to compare the accuracy rate with a set first threshold; A fourth execution sub-module, configured to train the long short-term memory neural network model to a convergence state when the accuracy rate is greater than the first threshold.
11. A computer device, including a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the action understanding method according to any one of claims 1 to 5.
12. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the action understanding method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Human motion recognition method and device
CN108985259A