Action Recognition Method, Device, Electronic Device and Storage Medium
By using the similarity calculation method of three-dimensional key points and candidate action features in action recognition, the problem of low accuracy in action recognition in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202110852500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-07-27
AI Technical Summary
In the prior art, in action recognition, the accuracy is affected by the action coverage range in the pre-acquisitioned sample images, resulting in low accuracy in later recognition.
By determining the set of three-dimensional key points of the target object in the image to be processed, and based on the relative positional relationship of the feature set of candidate actions and the three-dimensional key points, the similarity of the set of features of the to be processed is to determine whether the to be processed is a candidate action.
The accuracy of action recognition is improved, and the problem of inaccurate recognition caused by insufficient sample collection is avoided. Due to the good robustness of the method, it is suitable for target objects of different body types.
Smart Images

Figure CN113469134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method, apparatus, electronic device, and storage medium for action recognition. Background Art
[0002] With the popularization of monitoring devices such as cameras, video materials have shown an explosive growth. Real-time analysis of the actions of objects in video images has a wide range of application prospects.
[0003] When the prior art performs action recognition, it generally first collects sample images, labels the actions in the sample images, and then inputs the sample images and the labeling information into an action recognition model for training. After the training is completed, actions are recognized based on the action recognition model. The accuracy of action recognition in the prior art is affected by the pre-collected sample images. Generally, the actions in the pre-collected sample images cannot cover all actions. Therefore, when actions are recognized using the action recognition model later, the accuracy of action recognition is low. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for action recognition, so as to solve the problem of improving the accuracy of action recognition.
[0005] Embodiments of the present invention provide a method for action recognition, and the method includes:
[0006] Determine a set of three-dimensional key points corresponding to a target object in an image to be processed; wherein each three-dimensional key point in the set of three-dimensional key points is determined corresponding to a two-dimensional key point of the target object in the image to be processed;
[0007] Based on a set of candidate action features corresponding to a candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, determine a set of to-be-processed action features corresponding to the to-be-processed action executed by the target object; the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to a historical object when the historical object executes the candidate action;
[0008] Determine the similarity between the set of candidate action features and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0009] Further, the determining whether the to-be-processed action is the candidate action according to the similarity includes:
[0010] If it is determined that the similarity is greater than a preset first similarity threshold, then determine that the to-be-processed action is the candidate action.
[0011] Furthermore, there are at least two candidate actions. Determining the similarity between the candidate action feature set and the action to be processed feature set, and determining whether the action to be processed is the candidate action according to the similarity includes:
[0012] Determine the similarity between the candidate action feature set corresponding to each candidate action among the at least two candidate actions and the action to be processed feature set;
[0013] Select the candidate action corresponding to the maximum value among the determined similarities;
[0014] Determine that the action to be processed is the selected candidate action.
[0015] Furthermore, determining that the action to be processed is the selected candidate action includes:
[0016] If it is determined that the maximum value among the similarities is greater than a preset second similarity threshold, then determine that the action to be processed is the selected candidate action.
[0017] Furthermore, determining the similarity between the candidate action feature set and the action to be processed feature set includes:
[0018] Determine the sub-similarity between each sub-candidate action feature in the candidate action feature set and the corresponding sub-action to be processed feature in the action to be processed feature set;
[0019] According to each determined sub-similarity, determine the similarity between the candidate action feature set and the action to be processed feature set.
[0020] Furthermore, each sub-candidate action feature includes the angle feature of each sub-candidate action, and each sub-action to be processed feature includes the angle feature of each sub-action to be processed;
[0021] Determining the sub-similarity between each sub-candidate action feature in the candidate action feature set and the corresponding sub-action to be processed feature in the action to be processed feature set includes:
[0022] For the angle feature of each sub-candidate action in the candidate action feature set, determine the angle difference between the angle feature of the sub-candidate action and the angle feature of the sub-action to be processed corresponding to the angle feature of the sub-candidate action; according to the angle difference, determine the sub-similarity between the angle feature of the sub-candidate action and the angle feature of the sub-action to be processed.
[0023] Furthermore, according to each determined sub-similarity, determining the similarity between the candidate action feature set and the action to be processed feature set includes:
[0024] Determine the weight value corresponding to each sub - similarity according to the preset weight value corresponding to each sub - candidate action feature.
[0025] Determine the similarity between the candidate action feature set and the to - be - processed action feature set according to each sub - similarity and the weight value corresponding to each sub - similarity.
[0026] On the other hand, an embodiment of the present invention provides an action recognition device, and the device includes:
[0027] A first determination module, configured to determine a three - dimensional key point set corresponding to a target object in an image to be processed; wherein each three - dimensional key point in the three - dimensional key point set is determined corresponding to a two - dimensional key point of the target object in the image to be processed.
[0028] A second determination module, configured to determine a to - be - processed action feature set corresponding to the to - be - processed action executed by the target object based on a candidate action feature set corresponding to a candidate action and the relative position relationship between different three - dimensional key points in the three - dimensional key point set; the candidate action feature set is determined based on the relative position relationship between different three - dimensional key points in the three - dimensional key point set when a historical object executes the candidate action.
[0029] A third determination module, configured to determine the similarity between the candidate action feature set and the to - be - processed action feature set, and determine whether the to - be - processed action is the candidate action according to the similarity.
[0030] Further, the third determination module is specifically configured to determine that the to - be - processed action is the candidate action if it is determined that the similarity is greater than a preset first similarity threshold.
[0031] Further, the third determination module is specifically configured to determine the similarity between the candidate action feature set corresponding to each candidate action in the at least two candidate actions and the to - be - processed action feature set; select the candidate action corresponding to the maximum value of the determined similarities; and determine that the to - be - processed action is the selected candidate action.
[0032] Further, the third determination module is specifically configured to determine that the to - be - processed action is the selected candidate action if it is determined that the maximum value of the similarities is greater than a preset second similarity threshold.
[0033] Further, the third determination module is specifically configured to determine the sub-similarity between each sub-candidate action feature in the candidate action feature set and each corresponding sub-processed action feature in the to-be-processed action feature set; and determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each determined sub-similarity.
[0034] Further, the third determination module is specifically configured to, for the angle feature of each sub-candidate action in the candidate action feature set, determine the angle difference between the angle feature of the sub-candidate action and the angle feature of the corresponding sub-processed action of the sub-candidate action; and determine the sub-similarity between the angle feature of the sub-candidate action and the angle feature of the sub-processed action according to the angle difference.
[0035] Further, the third determination module is specifically configured to determine the weight value corresponding to each sub-similarity according to the weight value corresponding to each sub-candidate action feature preset; and determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each sub-similarity and the weight value corresponding to each sub-similarity.
[0036] On the other hand, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0037] The memory is used to store a computer program;
[0038] The processor is configured to implement the method steps described in any one of the above when executing the program stored on the memory.
[0039] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program implements the method steps described in any one of the above when being executed by a processor.
[0040] An embodiment of the present invention provides a method, apparatus, electronic device, and storage medium for action recognition. The method includes: determining a set of three-dimensional key points corresponding to a target object in an image to be processed, where each three-dimensional key point in the set of three-dimensional key points is determined corresponding to a two-dimensional key point of the target object in the image to be processed; determining a set of to-be-processed action features corresponding to the to-be-processed action performed by the target object based on a set of candidate action features corresponding to a candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, where the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object performs the candidate action; determining the similarity between the set of candidate action features and the set of to-be-processed action features, and determining whether the to-be-processed action is the candidate action according to the similarity.
[0041] The above technical solution has the following advantages or beneficial effects:
[0042] In an embodiment of the present invention, a set of three-dimensional key points corresponding to a target object in an image to be processed is determined, and then a set of to-be-processed action features corresponding to the to-be-processed action of the target object is determined according to the relative position relationship between different three-dimensional key points in the set of three-dimensional key points. By calculating the similarity between the set of to-be-processed action features performed by the target object and the set of candidate action features corresponding to the candidate action, it is determined whether the to-be-processed action is the candidate action according to the similarity. The action recognition method provided by the embodiment of the present invention does not need to rely on an action recognition model for action recognition, so as to avoid the problem of inaccurate action recognition caused by sample collection and improve the accuracy of action recognition. Moreover, since the difference in the position relationship of key points of target objects with different body types is small, the method provided by the embodiment of the present invention for determining action features according to the position relationship of key points of the target object and then recognizing actions according to the action features has good robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0044] Figure 1 It is a schematic diagram of the action recognition process provided in Embodiment 1 of the present invention;
[0045] Figure 2 It is a flowchart of action recognition provided in Embodiment 5 of the present invention;
[0046] Figure 3Schematic structural diagram of the action recognition device provided in Embodiment 6 of the present invention;
[0047] Figure 4 Schematic structural diagram of the action recognition provided in Embodiment 7 of the present invention. Detailed implementation manners
[0048] The present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1:
[0050] Figure 1 Schematic diagram of the action recognition process provided in the embodiment of the present invention. The process includes the following steps:
[0051] S101: Determine a set of three-dimensional key points corresponding to the target object in the image to be processed; each three-dimensional key point in the set of three-dimensional key points is determined corresponding to the two-dimensional key points of the target object in the image to be processed.
[0052] S102: Based on the set of candidate action features corresponding to the candidate action and the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, determine a set of to-be-processed action features corresponding to the to-be-processed action executed by the target object; the set of candidate action features is determined based on the relative position relationship of different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object executes the candidate action.
[0053] S103: Determine the similarity between the set of candidate action features and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0054] The action recognition method provided in the embodiment of the present invention is applied to an electronic device. The electronic device can be a device such as a PC or a tablet computer, or an intelligent image acquisition device. If the electronic device is an intelligent image acquisition device, after the intelligent image acquisition device acquires the image to be processed, it directly performs the subsequent action recognition process based on the image to be processed. If the electronic device is a device such as a PC or a tablet computer, after the image acquisition device acquires the image to be processed, it first sends the image to be processed to the electronic device, and then the electronic device performs the subsequent action recognition process based on the image to be processed.
[0055] It should be noted that the action recognition method provided in the embodiment of the present invention can be used to recognize the actions of the human body or the actions of animals.
[0056] Based on the image to be processed, the electronic device determines a set of three-dimensional key points corresponding to the target object in the image to be processed. Among them, the electronic device first determines the two-dimensional key points of the target object in the image to be processed, and then determines the three-dimensional key points based on the two-dimensional key points. Specifically, the electronic device can store a pre-trained key point recognition model, input the image to be processed into the pre-trained key point recognition model, and determine the two-dimensional coordinate information of each two-dimensional key point of the target object in the image to be processed based on the key point recognition model. The electronic device can store a pre-trained coordinate mapping model, input the two-dimensional coordinate information of each two-dimensional key point into the pre-trained coordinate mapping model, and determine the three-dimensional coordinate information of the corresponding three-dimensional key point based on the coordinate mapping model. Each three-dimensional key point of the target object constitutes a set of three-dimensional key points.
[0057] The training process of the key point recognition model includes: for each sample image in the first training set stored in the electronic device, input the sample image and the corresponding labeled image into the key point recognition model to train the key point recognition model; among them, the labeled image is labeled with the two-dimensional coordinate information of the two-dimensional key points in the corresponding sample image. The training process of the coordinate mapping model includes: for the two-dimensional coordinate information of each two-dimensional key point in the second training set stored in the electronic device, input the two-dimensional coordinate information and the corresponding three-dimensional coordinate information into the coordinate mapping model to train the coordinate mapping model; among them, the three-dimensional coordinate information is the coordinate information after mapping the corresponding two-dimensional coordinate information.
[0058] The embodiments of the present invention do not limit the method for determining the set of three-dimensional key points corresponding to the target object.
[0059] The electronic device stores candidate actions and a set of candidate action features corresponding to the candidate actions. Among them, the set of candidate action features is determined based on the relative position relationship of different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object performs the candidate action. The set of candidate action features can be a set of distance features, a set of angle features, etc.
[0060] Based on the set of candidate action features corresponding to the candidate action and the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, determine the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object. If the set of candidate action features is a set of distance features, then according to the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object determined is also a set of distance features. If the set of candidate action features is a set of angle features, then according to the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object determined is also a set of angle features.
[0061] After determining the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object, determine the similarity between the set of candidate action features and the set of to-be-processed action features. Then, determine whether the to-be-processed action is a candidate action according to the similarity.
[0062] In an embodiment of the present invention, a set of three-dimensional key points corresponding to a target object in a to-be-processed image is determined, and then, according to the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, a set of to-be-processed action features corresponding to the to-be-processed action of the target object is determined. By calculating the similarity between the set of to-be-processed action features performed by the target object and the set of candidate action features corresponding to a candidate action, it is determined whether the to-be-processed action is a candidate action according to the similarity. The action recognition method provided by the embodiment of the present invention does not need to rely on an action recognition model for action recognition, so as to avoid the problem of inaccurate action recognition caused by sample collection and improve the accuracy of action recognition. Moreover, since the difference in the position relationship of the key points of target objects with different body types is small, the method provided by the embodiment of the present invention for determining action features according to the position relationship of the key points of the target object and then recognizing actions according to the action features has good robustness.
[0063] Embodiment 2:
[0064] For action recognition, based on the above embodiments, in an embodiment of the present invention, the determining whether the to-be-processed action is the candidate action according to the similarity includes:
[0065] If it is determined that the similarity is greater than a preset first similarity threshold, determine that the to-be-processed action is the candidate action.
[0066] In an embodiment of the present invention, an electronic device stores a preset first similarity threshold, which may be, for example, 0.8, 0.85, etc. After the electronic device determines the similarity between the set of candidate action features and the set of to-be-processed action features, it determines whether the similarity is greater than the preset first similarity threshold. If so, it determines that the to-be-processed action is a candidate action; if not, it determines that the to-be-processed action is not a candidate action.
[0067] Embodiment 3:
[0068] The candidate actions stored in the electronic device may be multiple. For accurate action recognition, based on the above embodiments, in an embodiment of the present invention, the candidate action includes at least two, and the determining the similarity between the set of candidate action features and the set of to-be-processed action features and determining whether the to-be-processed action is the candidate action according to the similarity includes:
[0069] Determine the similarity between the candidate action feature set corresponding to each candidate action among the at least two candidate actions and the to-be-processed action feature set;
[0070] Select the candidate action corresponding to the maximum value among the determined similarities;
[0071] Determine that the to-be-processed action is the selected candidate action.
[0072] In the embodiments of the present invention, there are at least two candidate actions, and the candidate action feature set corresponding to each candidate action is different. For each candidate action, determine the similarity between the candidate action feature set corresponding to the candidate action and the to-be-processed action feature set.
[0073] After determining the similarity between the candidate action feature set corresponding to each candidate action and the to-be-processed action feature set of the to-be-processed action respectively, select the candidate action corresponding to the maximum value of the similarity. The selected candidate action has the highest similarity with the to-be-processed action. Therefore, determine that the to-be-processed action is the selected candidate action.
[0074] In order to further make the action recognition more accurate, in the embodiments of the present invention, the determining that the to-be-processed action is the selected candidate action includes:
[0075] If it is determined that the maximum value of the similarities is greater than a preset second similarity threshold, determine that the to-be-processed action is the selected candidate action.
[0076] There may be multiple candidate actions stored in the electronic device, and it is possible that the to-be-processed action is different from multiple candidate actions. Considering the above, in the embodiments of the present invention, the electronic device stores a preset second similarity threshold, where the preset second similarity threshold and the preset first similarity threshold may be the same or different. After selecting the candidate action corresponding to the maximum value of the determined similarities, determine whether the maximum value of the similarity is greater than the preset second similarity threshold. If so, determine that the to-be-processed action is the selected candidate action. If not, determine that the to-be-processed action is not the selected candidate action.
[0077] Since in the embodiments of the present invention, when there are at least two candidate actions, first determine the candidate action corresponding to the maximum value of the similarity, and then combine the limiting condition, that is, the preset second similarity threshold, to finally identify whether the to-be-processed action is the candidate action corresponding to the maximum value of the similarity. Therefore, the action recognition method provided by the embodiments of the present invention is more accurate.
[0078] Embodiment 4:
[0079] Based on the above embodiments, in order to more accurately determine the similarity between the candidate action feature set and the to-be-processed action feature set, in the embodiments of the present invention, determining the similarity between the candidate action feature set and the to-be-processed action feature set includes:
[0080] Determining the sub-similarity between each sub-candidate action feature in the candidate action feature set and each corresponding sub-to-be-processed action feature in the to-be-processed action feature set;
[0081] Determining the similarity between the candidate action feature set and the to-be-processed action feature set according to each determined sub-similarity.
[0082] Generally speaking, the candidate action feature set contains multiple sub-candidate action features, the to-be-processed action feature set contains multiple sub-to-be-processed action features, and there is a corresponding relationship between the sub-candidate action features and the sub-to-be-processed action features.
[0083] For example, the sub-candidate action features included in the candidate action feature set are the angle between the forearm and the upper arm, and the angle between the upper arm and the torso. Then the three-dimensional key point set corresponding to the candidate action is the wrist, elbow, shoulder, and waist. Among them, the forearm feature vector is determined according to the three-dimensional coordinate information of the wrist and elbow, the upper arm feature vector is determined according to the three-dimensional coordinate information of the elbow and shoulder, and a sub-candidate action feature in the candidate action feature set is determined according to the forearm feature vector and the upper arm feature vector. The upper arm feature vector is determined according to the three-dimensional coordinate information of the elbow and shoulder, the torso feature vector is determined according to the three-dimensional coordinate information of the shoulder and waist, and another sub-candidate action feature in the candidate action feature set is determined according to the upper arm feature vector and the torso feature vector.
[0084] When determining the to-be-processed action feature set of the to-be-processed action, first determine the four three-dimensional key points of the wrist, elbow, shoulder, and waist, and then determine a sub-to-be-processed action feature composed of the forearm feature vector and the upper arm feature vector in the to-be-processed action feature set, and another sub-to-be-processed action feature composed of the upper arm feature vector and the torso feature vector.
[0085] It can be found from this that there is a corresponding relationship between each sub-candidate action feature in the candidate action feature set and each corresponding sub-to-be-processed action feature in the to-be-processed action feature set. When determining the similarity between the candidate action feature set and the to-be-processed action feature set, first determine the sub-similarity between each sub-candidate action feature in the candidate action feature set and each corresponding sub-to-be-processed action feature in the to-be-processed action feature set, and then determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each determined sub-similarity. For example, the average value of each sub-similarity can be used as the similarity between the candidate action feature set and the to-be-processed action feature set.
[0086] In order to make the sub - similarity between each sub - candidate action feature in the determined candidate action feature set and the corresponding each sub - to - be - processed action feature in the to - be - processed action feature set more accurate, in an embodiment of the present invention, each of the sub - candidate action features includes an angular feature of each sub - candidate action, and each of the sub - to - be - processed action features includes an angular feature of each sub - to - be - processed action;
[0087] Determining the sub - similarity between each sub - candidate action feature in the candidate action feature set and the corresponding each sub - to - be - processed action feature in the to - be - processed action feature set includes:
[0088] For the angular feature of each sub - candidate action in the candidate action feature set, determine the angular difference between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action corresponding to the angular feature of the sub - candidate action; according to the angular difference, determine the sub - similarity between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action.
[0089] In an embodiment of the present invention, for the angular feature of each sub - candidate action in the candidate action feature set, determine the angular difference between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action corresponding to the angular feature of the sub - candidate action. The angular difference reflects the degree of difference between the angular feature of the sub - candidate action and the angular feature of the sub - candidate action, and the degree of difference is negatively correlated with the sub - similarity. The electronic device can pre - save the corresponding relationship between the angular difference and the sub - similarity, and the larger the angular difference, the smaller the corresponding sub - similarity. After determining the angular difference between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action corresponding to the angular feature of the sub - candidate action, according to the corresponding relationship between the angular difference and the sub - similarity, determine the sub - similarity between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action.
[0090] Considering that for target objects of different body types, the angular features of the same action have a small difference, therefore, in an embodiment of the present invention, determining the sub - similarity between the angular feature of the sub - candidate action and the angular feature of the sub - to - be - processed action based on the angular difference makes the subsequent action recognition more accurate and more robust.
[0091] Embodiment 5:
[0092] Considering that different actions have different degrees of attention to different parts, in order to make the action recognition more accurate, on the basis of the above - mentioned embodiments, in an embodiment of the present invention, determining the similarity between the candidate action feature set and the to - be - processed action feature set according to each determined sub - similarity includes:
[0093] According to the weight value corresponding to each sub - candidate action feature preset, determine the weight value corresponding to each sub - similarity;
[0094] Determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each of the sub-similarities and the weight values respectively corresponding to each of the sub-similarities.
[0095] In an embodiment of the present invention, for each candidate action, an electronic device pre-saves the weight values respectively corresponding to each sub-candidate action feature. The weight values respectively corresponding to each sub-candidate action feature are used as the weight values of the sub-similarities determined based on each sub-candidate action feature. According to each sub-similarity and the weight values respectively corresponding to each sub-similarity, a weighted calculation is performed to obtain the similarity between the candidate action feature set and the to-be-processed action feature set.
[0096] In an embodiment of the present invention, the similarity between the candidate action feature set and the to-be-processed action feature set can be determined, and the difference degree between the candidate action feature set and the to-be-processed action feature set can also be determined. Then subsequent action recognition is performed. If the determined value is the similarity, the greater the similarity, the higher the similarity to the candidate action. If the determined value is the difference degree, the greater the difference degree, the higher the difference degree from the candidate action.
[0097] The action recognition process provided by the embodiment of the present invention will be described in detail below.
[0098] Figure 2 This is the action recognition flowchart provided by the embodiment of the present invention.
[0099] As Figure 2 shown, the action recognition provided by the embodiment of the present invention includes two parts: First, extraction of candidate action templates. Second, measurement of the similarity between the to-be-processed action and the candidate action.
[0100] First, extraction of candidate action templates.
[0101] (1) Collect sample images, input the sample images into a key point recognition model, and based on the key point recognition model, obtain the two-dimensional coordinate information of each two-dimensional key point of the target object in the sample images. Currently, the key point recognition model can be divided into two types of methods: top-down and bottom-up. The top-down method means first detecting the human body frame and then predicting the human body skeleton key points within the detected human body frame. The bottom-up method detects the key points of the whole image of the human body and then uses a clustering method to connect the different key points of different people together. The key point recognition model does not limit the type of key point detection method.
[0102] (2) Input the two-dimensional key point input coordinates into the coordinate mapping model to obtain the three-dimensional coordinate information of the corresponding three-dimensional key points. The coordinate mapping model can be divided into two types of methods according to the key point input type: single-frame image two-dimensional key point input and multi-frame image two-dimensional key point input. Single-frame two-dimensional key point input means inputting the two-dimensional key points of the current frame into the coordinate mapping model to determine the three-dimensional key points of the current frame. Multi-frame two-dimensional key point input means inputting the two-dimensional key points of multiple frames before and after the current frame into the coordinate mapping model to predict the three-dimensional key points of the current frame in time series. The embodiments of the present invention do not limit the method for determining three-dimensional key points.
[0103] (3) Determine the candidate action feature set corresponding to the candidate action according to the relative position relationship of the three-dimensional key points. Calculate n artificially set sub-candidate action features of interest according to the candidate action type. The sub-candidate action features include: the vector included angle Ai (i∈{1,2,…,n}) between different key points.
[0104] Taking the left wrist key point (x1, y1, z1), left elbow key point (x2, y2, z2), and left shoulder key point (x3, y3, z3) on the left arm as an example, the human body vector included angle feature is the included angle between two vectors formed by multiple key points. The calculation method of the cosine value of the included angle is the inner product of the two vectors divided by the modulus of the two vectors. The calculation process is as follows:
[0105] V1 = (x2 - x1, y2 - y1, z2 - z1);
[0106] V2 = (x3 - x2, y3 - y2, z3 - z2);
[0107] A = (V1 * V2) / (|V1| * |V2|);
[0108] Where V1 and V2 are the vectors formed between key points, and A is the cosine value of the included angle between the two vectors. The included angle value can be obtained according to the cosine value of the angle.
[0109] (4) Use the candidate feature set composed of the above sub-candidate action features as a template and save it.
[0110] II. Measurement of the similarity between the action to be processed and the candidate action.
[0111] (5) Obtain the image to be processed, input the image to be processed into the key point recognition model, and based on the key point recognition model, obtain the two-dimensional coordinate information of each two-dimensional key point of the target object in the image to be processed.
[0112] (6) Input the two-dimensional key point into the coordinate mapping model to obtain the three-dimensional coordinate information of the corresponding three-dimensional key point.
[0113] (7) Determine the action feature set to be processed according to the relative position relationship of the three-dimensional key points.
[0114] (8) Determine the similarity between the candidate action feature set and the to-be-processed action feature set, and determine whether the to-be-processed action is a candidate action according to the similarity.
[0115] That is, based on the to-be-processed image, calculate the N sub-to-be-processed action features ai that are of concern and set artificially at the current moment according to the above steps (1) to (3). Calculate the difference Di between the sub-candidate action feature Ai in the candidate action template and the sub-to-be-processed action feature ai at the current moment; that is, Di = |Ai - ai|.
[0116] According to the attention degrees of different actions to different parts, different weights Pi are assigned to each of the above sub-candidate action features. For example, for the raising hand action, the upper body features of the human body are concerned, but the arm features are more concerned. At this time, the weight of the arm should be higher than the weight of the angle of the remaining features of the upper body. Calculate the difference value S between the current action and the action template according to the weighted sum of different feature difference values; that is,
[0117] According to the different types of candidate actions, a difference degree threshold St is set. The smaller St is, the stricter the action recognition similarity discrimination is. When the above difference value S is less than the angle threshold St, it can be considered that the to-be-processed action at the current moment has a high similarity with the candidate action template, and the current action is discriminated as the template action.
[0118] Embodiment 6:
[0119] Figure 3 It is a schematic structural diagram of an action recognition device provided by an embodiment of the present invention. The device includes:
[0120] A first determination module 31, configured to determine a three-dimensional key point set corresponding to a target object in the to-be-processed image; each three-dimensional key point in the three-dimensional key point set is determined corresponding to a two-dimensional key point of the target object in the to-be-processed image;
[0121] A second determination module 32, configured to determine a to-be-processed action feature set corresponding to the to-be-processed action executed by the target object based on the candidate action feature set corresponding to the candidate action and the relative position relationship of different three-dimensional key points in the three-dimensional key point set; the candidate action feature set is determined based on the relative position relationship of different three-dimensional key points in the three-dimensional key point set corresponding to the historical object when the historical object executes the candidate action;
[0122] A third determination module 33, configured to determine the similarity between the candidate action feature set and the to-be-processed action feature set, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0123] The third determination module 33 is specifically configured to determine the to-be-processed action as the candidate action if it is determined that the similarity is greater than a preset first similarity threshold.
[0124] The third determination module 33 is specifically configured to determine the similarity between the candidate action feature set corresponding to each candidate action in the at least two candidate actions and the to-be-processed action feature set; select the candidate action corresponding to the maximum value among the determined similarities; and determine the to-be-processed action as the selected candidate action.
[0125] The third determination module 33 is specifically configured to determine the to-be-processed action as the selected candidate action if it is determined that the maximum value among the similarities is greater than a preset second similarity threshold.
[0126] The third determination module 33 is specifically configured to determine the sub-similarity between each sub-candidate action feature in the candidate action feature set and the corresponding sub-to-be-processed action feature in the to-be-processed action feature set; and determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each determined sub-similarity.
[0127] The third determination module 33 is specifically configured to, for the angle feature of each sub-candidate action in the candidate action feature set, determine the angle difference between the angle feature of the sub-candidate action and the angle feature of the corresponding sub-to-be-processed action; and determine the sub-similarity between the angle feature of the sub-candidate action and the angle feature of the sub-to-be-processed action according to the angle difference.
[0128] The third determination module 33 is specifically configured to determine the weight value corresponding to each sub-similarity according to the weight value corresponding to each sub-candidate action feature preset; and determine the similarity between the candidate action feature set and the to-be-processed action feature set according to each sub-similarity and the weight value corresponding to each sub-similarity.
[0129] Embodiment 7:
[0130] Based on the above embodiments, an electronic device is further provided in an embodiment of the present invention. As Figure 4 shown, it includes: a processor 301, a communication interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304;
[0131] A computer program is stored in the memory 303. When the program is executed by the processor 301, the processor 301 is caused to execute the following steps:
[0132] Determine a set of three-dimensional key points corresponding to the target object in the image to be processed; each three-dimensional key point in the set of three-dimensional key points is determined corresponding to the two-dimensional key points of the target object in the image to be processed;
[0133] Based on the set of candidate action features corresponding to the candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, determine the set of to-be-processed action features corresponding to the to-be-processed action executed by the target object; the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object executes the candidate action;
[0134] Determine the similarity between the set of candidate action features and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0135] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device. Since the principle of solving problems by the above electronic device is similar to the action recognition method, the implementation of the above electronic device can refer to the implementation of the method, and the repeated parts will not be described again.
[0136] The electronic device provided in the embodiment of the present invention may specifically be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), a network-side device, etc.
[0137] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0138] The communication interface 302 is used for communication between the above electronic device and other devices.
[0139] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0140] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0141] In the embodiment of the present invention, when the processor executes the program stored in the memory, it realizes determining a set of three-dimensional key points corresponding to the target object in the image to be processed; wherein each three-dimensional key point in the set of three-dimensional key points is determined corresponding to the two-dimensional key points of the target object in the image to be processed; based on the set of candidate action features corresponding to the candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, determining a set of processed action features corresponding to the processed action executed by the target object; the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object executes the candidate action; determining the similarity between the set of candidate action features and the set of processed action features, and determining whether the processed action is the candidate action according to the similarity.
[0142] In the embodiment of the present invention, a set of three-dimensional key points corresponding to the target object in the image to be processed is determined, and then, according to the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, a set of processed action features corresponding to the processed action of the target object is determined. By calculating the similarity between the set of processed action features of the target object and the set of candidate action features corresponding to the candidate action, it is determined whether the processed action is the candidate action according to the similarity. The action recognition method provided by the embodiment of the present invention does not need to rely on an action recognition model for action recognition, so as to avoid the problem of inaccurate action recognition caused by sample collection and improve the accuracy of action recognition. Moreover, since the difference in the position relationship of the key points of target objects of different body types is small, the method provided by the embodiment of the present invention for determining action features according to the position relationship of the key points of the target object and then recognizing actions according to the action features has good robustness.
[0143] Embodiment 8:
[0144] Based on the above embodiments, the embodiment of the present invention further provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, the electronic device is caused to execute the following steps when executed:
[0145] Determine a set of three-dimensional key points corresponding to the target object in the image to be processed; each three-dimensional key point in the set of three-dimensional key points is determined corresponding to the two-dimensional key point of the target object in the image to be processed;
[0146] Based on the set of candidate action feature corresponding to the candidate action and the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, determine the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object; the set of candidate action feature is determined based on the relative position relationship of different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object performs the candidate action;
[0147] Determine the similarity between the set of candidate action feature and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0148] Based on the same inventive concept, an embodiment of the present invention also provides a computer-readable storage medium. Since the principle of the processor solving problems when executing the computer program stored on the above computer-readable storage medium is similar to the action recognition method, the implementation of the processor executing the computer program stored on the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be described again.
[0149] The above computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor in the electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid state drives (SSD), etc.
[0150] In the computer-readable storage medium provided in the embodiment of the present invention, a computer program is stored. When the computer program is executed by the processor, it realizes determining a set of three-dimensional key points corresponding to the target object in the image to be processed; each three-dimensional key point in the set of three-dimensional key points is determined corresponding to the two-dimensional key point of the target object in the image to be processed; based on the set of candidate action feature corresponding to the candidate action and the relative position relationship of different three-dimensional key points in the set of three-dimensional key points, determine the set of to-be-processed action features corresponding to the to-be-processed action performed by the target object; the set of candidate action feature is determined based on the relative position relationship of different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object performs the candidate action; determine the similarity between the set of candidate action feature and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
[0151] In an embodiment of the present invention, a set of three-dimensional key points corresponding to a target object in a to-be-processed image is determined, and then, according to the relative position relationships of different three-dimensional key points in the set of three-dimensional key points, a set of to-be-processed action features corresponding to the to-be-processed action of the target object is determined. By calculating the similarity between the set of to-be-processed action features of the target object and a set of candidate action features corresponding to a candidate action, it is determined whether the to-be-processed action is the candidate action according to the similarity. The action recognition method provided by the embodiment of the present invention does not need to rely on an action recognition model for action recognition, so as to avoid the problem of inaccurate action recognition caused by sample collection, and improve the accuracy of action recognition. Moreover, since the position relationships of the key points of target objects with different body types have little difference, the method provided by the embodiment of the present invention for determining action features according to the position relationships of the key points of the target object and then recognizing actions based on the action features has good robustness.
[0152] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0155] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0156] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for action recognition, characterized in that, The method includes: Determine a set of three-dimensional key points corresponding to a target object in the image to be processed; each three-dimensional key point in the set of three-dimensional key points is determined corresponding to a two-dimensional key point of the target object in the image to be processed; wherein, based on a pre-trained key point recognition model, determine the two-dimensional coordinate information of each two-dimensional key point of the target object in the image to be processed; input the two-dimensional coordinate information of each two-dimensional key point into a pre-trained coordinate mapping model, and based on the coordinate mapping model, determine the three-dimensional coordinate information of the corresponding three-dimensional key point; each three-dimensional key point of the target object constitutes a set of three-dimensional key points; Based on the set of candidate action features corresponding to the candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points, determine a set of candidate action features corresponding to the action to be processed executed by the target object; the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object executes the candidate action; the set of candidate action features to be processed includes a set of distance features or a set of angle features; Determine the sub-similarity between each sub-candidate action feature in the set of candidate action features and the corresponding sub-candidate action feature to be processed in the set of candidate action features to be processed; according to the preset weight value corresponding to each sub-candidate action feature, determine the weight value corresponding to each sub-similarity, and according to each sub-similarity and the weight value corresponding to each sub-similarity, determine the similarity between the set of candidate action features and the set of candidate action features to be processed, and determine whether the action to be processed is the candidate action according to the similarity.
2. The method according to claim 1, characterized in that, The determining whether the action to be processed is the candidate action according to the similarity includes: If it is determined that the similarity is greater than a preset first similarity threshold, determine that the action to be processed is the candidate action.
3. The method according to claim 1, characterized in that, The candidate action includes at least two, and the determining the similarity between the set of candidate action features and the set of candidate action features to be processed, and determining whether the action to be processed is the candidate action according to the similarity includes: Determine the similarity between the set of candidate action features corresponding to each candidate action in the at least two candidate actions and the set of candidate action features to be processed; Select the candidate action corresponding to the maximum value of the determined similarities; Determine that the action to be processed is the selected candidate action.
4. The method according to claim 3, characterized in that, The determining that the action to be processed is the selected candidate action includes: If it is determined that the maximum value of the similarities is greater than a preset second similarity threshold, determine that the action to be processed is the selected candidate action.
5. The method according to claim 1, characterized in that, Each sub-candidate action feature includes the angle feature of each sub-candidate action, and each sub-candidate action feature to be processed includes the angle feature of each sub-candidate action to be processed; The determining the sub-similarity between each sub-candidate action feature in the set of candidate action features and the corresponding sub-candidate action feature to be processed in the set of candidate action features to be processed includes: For the angular feature of each sub-candidate action in the candidate action feature set, determine the angular difference between the angular feature of the sub-candidate action and the angular feature of the sub-action to be processed corresponding to the angular feature of the sub-candidate action; according to the angular difference, determine the sub-similarity between the angular feature of the sub-candidate action and the angular feature of the sub-action to be processed.
6. An action recognition device, characterized in that, The device includes: A first determination module, configured to determine a set of three-dimensional key points corresponding to a target object in a to-be-processed image; wherein each three-dimensional key point in the set of three-dimensional key points is determined corresponding to a two-dimensional key point of the target object in the to-be-processed image; wherein, based on a pre-trained key point recognition model, determine the two-dimensional coordinate information of each two-dimensional key point of the target object in the to-be-processed image; input the two-dimensional coordinate information of each two-dimensional key point into a pre-trained coordinate mapping model, and based on the coordinate mapping model, determine the three-dimensional coordinate information of the corresponding three-dimensional key point; each three-dimensional key point of the target object constitutes a set of three-dimensional key points; A second determination module, configured to determine a set of to-be-processed action features corresponding to the to-be-processed action performed by the target object based on a set of candidate action features corresponding to a candidate action and the relative position relationship between different three-dimensional key points in the set of three-dimensional key points; the set of candidate action features is determined based on the relative position relationship between different three-dimensional key points in the set of three-dimensional key points corresponding to the historical object when the historical object performs the candidate action; the set of to-be-processed action features includes a distance feature set or an angular feature set; A third determination module, configured to determine the sub-similarity between each sub-candidate action feature in the set of candidate action features and each corresponding sub-to-be-processed action feature in the set of to-be-processed action features; according to the weight value corresponding to each sub-candidate action feature preset, determine the weight value corresponding to each sub-similarity, according to each sub-similarity and the weight value corresponding to each sub-similarity, determine the similarity between the set of candidate action features and the set of to-be-processed action features, and determine whether the to-be-processed action is the candidate action according to the similarity.
7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used for storing a computer program; The processor, when executing the program stored on the memory, implements the method steps described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method steps described in any one of claims 1-5.
Citation Information
Patent Citations
Motion prompting method and device, electronic equipment and storage medium
CN111488824A
Method and apparatus for image search, and storage medium
WO2021056440A1