A behavior recognition method, device, and electronic device
By inputting the original image into the deep learning model and the humanoid detection model, acquiring and fusing the shooting equipment and object detection information, predicting the three-dimensional posture of the target object, solving the uncertainty of the two-dimensional plane to three-dimensional spatial mapping and improving the accuracy of behavior recognition.
Patent Information
- Application Number
- CN202210220925.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-03-08
AI Technical Summary
In the prior art, due to the inability to directly obtain the human body depth information during the behavior recognition process, there are many possibilities when mapping from a two-dimensional plane to a three-dimensional space, resulting in inaccurate behavior recognition.
By inputting the original image to the trained first deep learning model and humanoid detection model, the target feature information and object detection information of the shooting device are obtained, and fused to obtain the fusion feature, and then input it to the second deep learning model to predict the three-dimensional pose of the target object.
The accuracy of behavior recognition is improved, and by considering the shooting angle and human posture of the shooting device, the mapping possibility from the two-dimensional plane to the three-dimensional space is limited, making the predicted three-dimensional posture more accurate.
Smart Images

Figure CN114627552B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and particularly to a method, device and electronic device for behavior recognition. Background Art
[0002] Behavior recognition often requires the depth information of the human body as an aid to recognize the human body posture. For example, when performing fall detection, the depths of the head and feet of the human body are used; when performing sitting posture recognition, the depths of the joints of the upper body of the human body are used to determine whether there is rotation; and when recognizing the behaviors of multiple people in a multi-person scenario, the depth distance between people is used, etc.
[0003] The depth information of the human body cannot be directly obtained through a planar image, and it is necessary to map the two-dimensional plane shown in an image to a three-dimensional space. However, in the current process of judging and analyzing behavior recognition, there are multiple possibilities for mapping from the two-dimensional plane shown in an image to a three-dimensional space, and the posture of the shooting device for shooting the image will also cause errors in the posture estimation of the person, which will lead to inaccurate behavior recognition. Summary of the Invention
[0004] This application discloses a method, device and electronic device for behavior recognition to improve the accuracy of behavior recognition.
[0005] According to the first aspect of the embodiments of this application, a method for behavior recognition is provided, which at least includes:
[0006] Input the obtained original image of the target object into the trained first deep learning model to obtain the target feature information of the shooting device for shooting the original image, where the target feature information at least includes: the shooting angle of the shooting device when shooting the target object to obtain the original image, and the predicted image predicted by the shooting device shooting the target object at a specified angle;
[0007] Input the original image into the trained humanoid detection model to obtain target detection information, where the target detection information at least includes: the two-dimensional coordinates of at least one key point indicating the posture in the original image of the target object;
[0008] Fuse the target feature information and the target detection information to obtain a fused feature; the fused feature is used to predict the three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict the three-dimensional posture of the target object;
[0009] Input the fused feature into the trained second deep learning model to obtain the three-dimensional posture of the target object.
[0010] Optionally, the shooting angles when the shooting device shoots the target object to obtain the original image at least include:
[0011] The pitch angle and roll angle set when the shooting device shoots the target object to obtain the original image.
[0012] Optionally, the target detection information further includes: object features extracted from the original image for indicating the target object;
[0013] The fusion of the target feature information and the target detection information to obtain the fusion feature includes:
[0014] Inputting the original image and the object features into a trained second deep learning model to extract a corresponding object feature map from the original image according to the object features;
[0015] Fusing the target feature information, the target detection information and the object feature map to obtain the fusion feature.
[0016] Optionally, the second deep learning model obtains the three-dimensional pose at least through the following calculation layers:
[0017] The three-dimensional coordinate information prediction layer is used to predict the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system according to the fusion feature, perform a specified operation on the two-dimensional coordinates of the key points in each plane of the three-dimensional coordinate system to obtain the three-dimensional coordinate information of the key points, and output the three-dimensional coordinate information of the key points to the three-dimensional pose prediction layer; the three-dimensional coordinate system includes three planes, and the three planes are perpendicular to each other in pairs;
[0018] The three-dimensional pose prediction layer is used to predict the three-dimensional pose of the target object according to the three-dimensional coordinate information of each key point in the input target object.
[0019] Optionally, the prediction of the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system according to the fusion feature includes:
[0020] If there are the previous N consecutive video frames before the original image currently, determine the reference three-dimensional coordinate information based on the fusion features of each video frame in the previous N consecutive video frames and the fusion feature of the original image, and predict the two-dimensional coordinates of each key point in the target object in each plane of the three-dimensional coordinate system according to the reference three-dimensional coordinate information; the reference three-dimensional coordinate information at least includes: the three-dimensional coordinate information of each key point in the target object in each of the previous N consecutive video frames predicted based on the combined fusion features by combining the fusion features of each video frame in the previous N consecutive video frames and the fusion feature of the original image.
[0021] Optionally, the predicting the three-dimensional pose of the target object according to the three-dimensional coordinate information of each key point in the input target object includes:
[0022] Predict the three-dimensional pose of the target object according to the reference three-dimensional coordinate information and the three-dimensional coordinate information of each key point in the target object.
[0023] Optionally, the second deep learning model further includes: a feature map extraction layer;
[0024] The feature map extraction layer is configured to receive the input original image and the object feature, and extract the corresponding object feature map from the original image according to the object feature.
[0025] Optionally, the three-dimensional coordinate information of the key point is the three-dimensional coordinate information relative to the root node in the three-dimensional coordinate system, and the root node is a specified key point in the target object.
[0026] According to the second aspect of the embodiments of the present application, there is provided a behavior recognition device, which at least includes:
[0027] A target feature information obtaining unit, configured to input the obtained original image of the target object into the trained first deep learning model to obtain the target feature information of the shooting device for shooting the original image, where the target feature information at least includes: the shooting angle when the shooting device shoots the target object to obtain the original image, and the predicted image predicted by the shooting device shooting the target object at a specified angle;
[0028] A target detection information obtaining unit, configured to input the original image into the trained humanoid detection model to obtain target detection information, where the target detection information at least includes: the two-dimensional coordinates of at least one key point indicating the pose in the original image in the target object;
[0029] A feature fusion unit for fusing the target feature information and the target detection information to obtain a fused feature; the fused feature is used to predict the three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict the three-dimensional pose of the target object.
[0030] A three-dimensional pose prediction unit for inputting the fused feature into a trained second deep learning model to obtain the three-dimensional pose of the target object.
[0031] According to a third aspect of the embodiments of the present application, an electronic device is provided, and the electronic device includes: a processor and a memory;
[0032] The memory is used to store machine-executable instructions;
[0033] The processor is used to read and execute the machine-executable instructions stored in the memory to implement the behavior recognition method as described above.
[0034] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0035] As can be seen from the above technical solutions, the solution provided by the present application inputs the original image captured for the target object into the first deep learning model and the humanoid detection model respectively to obtain the target feature information and the target detection information of the imaging device for capturing the original image, then fuses the target feature information and the target detection information to obtain a fused feature, and finally inputs the fused feature into a trained second deep learning model to obtain the three-dimensional pose of the target object. The above target feature information includes the shooting angle when the imaging device shoots the target object. When predicting the three-dimensional pose of the target, the influence of the imaging device is also calculated, which limits the possibility of mapping the original image from a two-dimensional plane to a three-dimensional space, improves the accuracy of behavior recognition, and at the same time fuses the features obtained by multiple models, making the predicted three-dimensional pose of the target object more accurate.
[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present specification, and are used together with the specification to explain the principles of the present specification.
[0038] Figure 1 It is a flowchart of a behavior recognition method provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the pitch angle set for the imaging device provided by an embodiment of the present application;
[0040] Figure 3 Schematic diagram of the roll angle set for the photographing device provided in the embodiment of the present application;
[0041] Figure 4 Schematic diagram of the predicted image provided in the embodiment of the present application;
[0042] Figure 5 Schematic diagram of the humanoid picture of the target object in the original image provided in the embodiment of the present application;
[0043] Figure 6 Schematic diagram of a behavior recognition device provided in the embodiment of the present application;
[0044] Figure 7 Schematic diagram of an electronic device provided in the embodiment of the present application. Detailed implementation manners
[0045] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0046] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0047] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0048] In order to enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application and to make the above objects, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0049] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a behavior recognition method provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0050] Step 101: Input the obtained original image captured for the target object into the trained first deep learning model to obtain the target feature information of the shooting device that captured the original image.
[0051] In the embodiment of the present application, the original image obtained for the target object in this step 101 may be a single video frame in a video captured by a shooting device such as a camera, or the original image may also be a single image captured by a shooting device such as a camera. Among them, the shooting device that captures the original image in this embodiment is deployed for the target object. For example, a camera deployed for students in a classroom, a monitoring device deployed for pedestrians on the street, etc. In specific applications, the embodiment of the present application can establish a connection with the shooting device used to capture the target object to obtain a single image or video captured by the shooting device, or obtain a single image or video captured for the target object by receiving externally input data.
[0052] As an embodiment, since the behavior recognition of the target object is affected by the shooting angle when the shooting device captures the target object. For example, for the same object, when the shooting device captures the object at different shooting angles, the behaviors recognized for the object in the pictures at different shooting angles may be inconsistent.
[0053] Therefore, after obtaining the original image in the embodiment of the present application, the obtained original image is input into the trained first deep learning model, and the first deep learning model processes the input original image in the trained manner to obtain the target feature information of the shooting device that captured the original image. Among them, the target feature information at least includes: the shooting angle when the shooting device captures the above target object to obtain the original image, and the predicted image predicted by the shooting device when capturing the target object at a specified angle.
[0054] In this embodiment, the shooting angle set when the shooting device captures the target object to obtain the original image includes: the pitch angle and the roll angle. Among them, the pitch angle refers to the angle after the lens of the shooting device tilts upward or downward relative to the horizontally placed shooting device, and reference can be made to Figure 2 ; the roll angle refers to the angle after the lens of the shooting device rotates to the left or right relative to the shooting device with the central axis of the camera perpendicular to the ground, and reference can be made to Figure 3 .
[0055] As an example, the above predicted image is a predicted image obtained by predicting that the photographing device photographs the target object at a pitch angle of 0 and a roll angle of 0 (that is, the photographing device is placed horizontally and the central axis of the camera is perpendicular to the ground). When predicting the predicted image, start predicting from the boundary of the original image according to the currently obtained photographing angle of the photographing device. For example, when the currently obtained photographing angle of the photographing device is a top view relative to the specified angle, if the photographing device photographs at the specified angle at this time, a part of the upper edge of the original image cannot be photographed. Therefore, a predicted image relative to the original image can be obtained by predicting the part of the image that cannot be photographed when the photographing device photographs at the specified angle. Exemplarily, as Figure 4 The left figure in the middle is the original image photographed by the photographing device in the embodiment of the present application when looking up, and the right figure is the predicted image photographed by the photographing device in the embodiment of the present application at a pitch angle of 0 and a roll angle of 0. When photographing at a pitch angle of 0 and a roll angle of 0, the photographing device cannot photograph a part of the upper edge of the image (i.e., the gray part in the right figure).
[0056] Optionally, in the embodiment of the present application, the first deep learning model can be trained in the following manner: Prepare in advance an image sample set with marked photographing angles, and a predicted image sample set corresponding to each image sample in the image sample set obtained by photographing at a specified angle, and train the first deep learning model according to the image sample set and the predicted image sample set.
[0057] Step 102, input the original image into the trained human detection model to obtain target detection information.
[0058] As an example, the target detection information in step 102 at least includes: the two-dimensional coordinates of at least one key point for indicating the posture in the original image of the target object. Optionally, key points that can reflect the posture of the target object can be selected as key points, and / or, the center points of each part of the target object can also be used as key points.
[0059] Optionally, the human detection model in this step 102 can adopt a trained model in related technologies, and this human detection model can detect people existing in the original image. Since the target object included in the original image in the embodiment of the present application is a person, the position of the target object of the behavior to be recognized in the original image can be determined through the human detection model. Optionally, the human detection model in this embodiment can perform key point calculation on the 2D image of the target task through a human pose estimation network such as HRNET, ALPHAPOSE, OPENPOSE, etc. The specific calculation process can refer to related technologies and will not be elaborated here.
[0060] Optionally, the target detection information in the embodiments of the present application further includes: object features extracted from the original image for indicating the target object, such as the clothes and hats worn by the target object, the appearance features, height features, etc. of the target object, which can identify the features of the target object.
[0061] Step 103: Fuse the target feature information and the target detection information to obtain a fused feature.
[0062] In the embodiments of the present application, the fused feature is used to predict the three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict the three-dimensional pose of the target object. Optionally, when fusing the target feature information and the target detection information, it is necessary to convert both the target feature information and the target detection information into vector matrices, and then according to the maximum dimension of the converted vector matrices, convert each vector into the same dimension, and then add the converted vectors to obtain a fused vector, and this fused vector is the fused feature.
[0063] Based on the fact that the target detection information further includes object features extracted from the original image for indicating the target object, when fusing the target feature information and the target detection information in this embodiment to obtain a fused feature, the following steps can be performed: input the original image and the object features into a trained second deep learning model to extract a corresponding object feature map from the original image according to the object features; fuse the target feature information, the target detection information, and the object feature map to obtain a fused feature. Among them, the object feature map is an object feature map of a humanoid picture that enhances the target object extracted by the second deep learning model according to the object features from the original image, where the humanoid picture is as Figure 5 shown.
[0064] Exemplarily, the second deep learning model can extract multiple object feature maps from the original image according to the object features. The resolution of the object feature map will be reduced compared to the resolution of the original image. Reducing the resolution of the object feature map is to reduce the computational complexity when processing the object feature map. Fusing the object feature map with the target feature information and the target detection information can refer to the method of fusing the target feature information and the target detection information described above, which will not be elaborated here.
[0065] Step 104: Input the fused feature into a trained second deep learning model to obtain the three-dimensional pose of the target object
[0066] Optionally, the second deep learning model in the embodiments of the present application obtains the three-dimensional pose of the target object at least through the following calculation layers:
[0067] The three-dimensional coordinate information prediction layer is used to predict the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system based on the fused features, perform a specified operation on the two-dimensional coordinates of the key points in each plane of the three-dimensional coordinate system to obtain the three-dimensional coordinate information of the key points, and output the three-dimensional coordinate information of the key points to the three-dimensional pose prediction layer, where the three-dimensional coordinate system includes three planes and the three planes are perpendicular to each other in pairs.
[0068] As an example, the three-dimensional coordinate information can reflect the depth information of the human body. Therefore, the second deep learning model can identify the behavior of the target object based on the three-dimensional coordinate information of the target object.
[0069] In the embodiments of the present application, the three-dimensional coordinates in the three-dimensional coordinate information of the target object can be calculated in the following manner:
[0070] Exemplarily, a three-dimensional coordinate system can be established in the three-dimensional space mapped by the two-dimensional plane of the humanoid picture, and the two-dimensional coordinates of at least one key point of the target object in the three planes (xy plane, xz plane, yz plane) of the three-dimensional coordinate system can be obtained based on the fused features. Since at this time, for any key point, the point will have a coordinate in each of the three planes. For example, the coordinates of point A in the xy plane are (x1, y1), the coordinates of point A in the xz plane are (x2, z1), and the coordinates of point A in the yz plane are (y2, z2). Therefore, point A will have two x-axis coordinates, two y-axis coordinates, and two z-axis coordinates. The x coordinate in the three-dimensional coordinates of point A can be obtained by averaging all the two-dimensional x-axis coordinates corresponding to point A as (x1 + x2) / 2, the y coordinate in the three-dimensional coordinates of point A can be obtained by averaging all the two-dimensional y-axis coordinates corresponding to point A as (y1 + y2) / 2, and the z coordinate in the three-dimensional coordinates of point A can be obtained by averaging all the two-dimensional z-axis coordinates corresponding to point A as (z1 + z2) / 2, that is, the final three-dimensional coordinates of point A are ((x1 + x2) / 2, (y1 + y2) / 2, (z1 + z2) / 2).
[0071] The three-dimensional pose prediction layer is used to predict the three-dimensional pose of the target object based on the three-dimensional coordinate information of each key point in the input target object.
[0072] It should be noted that due to the continuity of human behavior, the postures in a single picture may mislead the behavior of the target object. For example, it is recognized from the current image for behavior recognition that the target object is running, but in fact the target object is just demonstrating a running action during walking and is not actually running. It is obvious that such behavior cannot accurately recognize the behavior of the target object by only looking at one frame of the image. In order to make the prediction of the three-dimensional posture of the target object more coherent and stable and improve the accuracy of behavior recognition of the target object, the embodiment of the present application can also optimize the three-dimensional posture prediction layer according to the following method:
[0073] In specific implementation, the three-dimensional coordinate information prediction layer predicts the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system based on the fused feature through the following steps:
[0074] As an embodiment, if the original image is a video frame in a video, if there are the previous N consecutive video frames before the current original image, the reference three-dimensional coordinate information is determined based on the fused features of each of the previous N consecutive video frames and the fused feature of the original image, and the two-dimensional coordinates of each key point in the target object in each plane of the three-dimensional coordinate system are predicted based on the reference three-dimensional coordinate information. Among them, the reference three-dimensional coordinate information at least includes: the three-dimensional coordinate information of each key point in the target object in each of the previous N consecutive video frames predicted based on the fused feature combined with the fused features of each of the previous N consecutive video frames and the fused feature of the original image.
[0075] Based on the above method for obtaining three-dimensional coordinate information, the three-dimensional posture prediction layer predicts the three-dimensional posture of the target object based on the three-dimensional coordinate information of each key point in the input target object through the following steps: predicting the three-dimensional posture of the target object based on the reference three-dimensional coordinate information and the three-dimensional coordinate information of each key point in the target object.
[0076] Exemplarily, if there are currently 12 video frames, denoted as the 0th to 11th video frames, the original image in the embodiments of the present application is the 11th frame. If the above N is 9, then the fusion features of the 2nd to 10th video frames are obtained. According to the fusion features of each of the 2nd to 10th video frames and the fusion feature of the 11th frame, the three-dimensional coordinate information of the target object in the 2nd to 10th video frames is determined. The three-dimensional coordinate information of the target object in the 2nd to 10th video frames is divided into 3 groups of 3 each: the 2nd to 4th frames, the 5th to 7th frames, and the 8th to 10th frames. Respectively, according to the three-dimensional coordinate information of the target object in the 2nd to 4th video frames, the two-dimensional coordinates of the target object in the 11th frame are predicted; according to the three-dimensional coordinate information of the target object in the 5th to 7th video frames, the two-dimensional coordinates of the target object in the 11th frame are predicted; according to the three-dimensional coordinate information of the target object in the 8th to 10th video frames, the two-dimensional coordinates of the target object in the 11th frame are predicted. Then, the average value of the two-dimensional coordinates of each key point of the target object in the 11th frame predicted above is calculated to obtain the final two-dimensional coordinates of the target object in the 11th frame. Furthermore, the three-dimensional coordinate information of each key point can be obtained by performing a specified operation on the two-dimensional coordinates of each key point of the target object in each plane of the three-dimensional coordinate system. In this embodiment, N can be selected as a multiple of 3, and according to the method of grouping N consecutive video frames as described above, the three-dimensional coordinate information of the target object is predicted.
[0077] Further, the three-dimensional pose of the target object can be predicted according to the three-dimensional coordinate information of the target object in the 2nd to 10th video frames and the three-dimensional coordinate information of the target object in the 11th frame.
[0078] In the above embodiments, the three-dimensional pose of the target object can be predicted by combining the fusion features corresponding to the target object in the N consecutive video frames before the original image and then based on the fusion features corresponding to the target object in the original image. Compared with predicting the three-dimensional pose of the target object relying only on a single image, considering the continuity of human actions, the accuracy of behavior recognition of the target object can be further improved.
[0079] As another embodiment, if the original image is a single picture, or if there are no N consecutive video frames before the original image currently, the two-dimensional coordinates of the key points of the target object are predicted based on the fusion features of the target object, and the three-dimensional pose of the target object is predicted based on the three-dimensional coordinates of the target object.
[0080] Optionally, the second deep learning model in the embodiments of the present application further includes: a feature map extraction layer. The feature map extraction layer is used to receive the input original image and object features, and extract the corresponding object feature map from the original image according to the object features.
[0081] Thus, the Figure 1 shown process is completed.
[0082] Through Figure 1 From the method embodiments described above, it can be seen that the solution provided by this application inputs the original image captured for the target object into the first deep learning model and the human detection model respectively to obtain the target feature information and target detection information of the shooting device for shooting the original image, then fuses the target feature information and the target detection information to obtain the fused feature, and finally inputs the fused feature into the trained second deep learning model to obtain the three-dimensional pose of the target object. The above-mentioned target feature information includes the shooting angle when the shooting device shoots the target object. When predicting the three-dimensional pose of the target, the influence of the shooting device is also taken into account, which limits the possibility of mapping the original image from a two-dimensional plane to a three-dimensional space, improves the accuracy of behavior recognition, and at the same time fuses the features obtained by multiple models, making the predicted three-dimensional pose of the target object more accurate.
[0083] Optionally, in the embodiments of this application, the three-dimensional coordinate information of each key point of the predicted target object is relative to the three-dimensional coordinate information of the root node in the three-dimensional coordinate system, and the root node is a specified key point in the target object.
[0084] For example, when the target object has 3 key points: A, B, and C, A is used as the root node, the depth of the position where point A is located is set to 0, the direction facing the outside of the image of point A is taken as the positive value, and the direction facing the outside of the image of point A is taken as the negative value. For example, if it is predicted that the distance between point B and the lens of the shooting device is closer than the distance between point A and the lens of the shooting device, it is determined that the depth of the position where point B is located is less than 0. By designating a key point in the target object as the root node to establish a three-dimensional coordinate system, it is more convenient for the second deep learning model to analyze the depth relationship between the corresponding key points of the target object for three-dimensional pose prediction.
[0085] In the embodiments of this application, the above-mentioned root node can also effectively reduce the missed detection rate when recognizing the specified behavior of the target object in the image. For example, in fall recognition, the missed detection rate of falls is reduced by judging the relative depth of the feet and the relative depth of the head. In the recognition of bad sitting postures, the method of designating a key point in the target object as the root node can also be used to judge whether a person is lying on the table or tilting backward, etc., effectively reducing the missed detection rate of recognition.
[0086] The above completes the introduction of the method embodiments provided by the embodiments of this application. Next, a behavior recognition device provided by the embodiments of this application will be described. As Figure 6 shown, the device at least includes:
[0087] A target feature information acquisition unit 601, configured to input an acquired original image captured for a target object into a trained first deep learning model, to obtain target feature information of a photographing device for photographing the original image, where the target feature information at least includes: a photographing angle when the photographing device photographs the target object to obtain the original image, and a predicted image predicted by the photographing device photographing the target object at a specified angle.
[0088] A target detection information acquisition unit 602, configured to input the original image into a trained humanoid detection model, to obtain target detection information, where the target detection information at least includes: two-dimensional coordinates of at least one key point for indicating a posture in the original image of the target object.
[0089] A feature fusion unit 603, configured to fuse the target feature information and the target detection information to obtain a fused feature; the fused feature is used to predict three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict a three-dimensional posture of the target object.
[0090] A three-dimensional posture prediction unit 604, configured to input the fused feature into a trained second deep learning model, to obtain the three-dimensional posture of the target object.
[0091] Optionally, the photographing angle when the photographing device photographs the target object to obtain the original image at least includes:
[0092] A pitch angle and a roll angle set when the photographing device photographs the target object to obtain the original image.
[0093] Optionally, the target detection information further includes: object features extracted from the original image for indicating the target object;
[0094] The feature fusion unit 603 fusing the target feature information and the target detection information to obtain a fused feature includes:
[0095] Inputting the original image and the object features into a trained second deep learning model, to extract a corresponding object feature map from the original image according to the object features;
[0096] Fusing the target feature information, the target detection information, and the object feature map to obtain a fused feature.
[0097] Optionally, the second deep learning model obtains the three-dimensional posture at least through the following calculation layers:
[0098] The three-dimensional coordinate information prediction layer is configured to predict the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system according to the fusion features, perform a specified operation on the two-dimensional coordinates of the key points in each plane of the three-dimensional coordinate system to obtain the three-dimensional coordinate information of the key points, and output the three-dimensional coordinate information of the key points to the three-dimensional pose prediction layer; the three-dimensional coordinate system includes three planes, and the three planes are perpendicular to each other in pairs;
[0099] The three-dimensional pose prediction layer is configured to predict the three-dimensional pose of the target object according to the three-dimensional coordinate information of each key point in the input target object.
[0100] Optionally, predicting the two-dimensional coordinates of the key points of the target object in each plane of the three-dimensional coordinate system according to the fusion features includes:
[0101] If there are the previous N consecutive video frames before the current original image, determine the reference three-dimensional coordinate information according to the fusion features of each video frame in the previous N consecutive video frames and the fusion features of the original image, and predict the two-dimensional coordinates of each key point in the target object in each plane of the three-dimensional coordinate system according to the reference three-dimensional coordinate information; the reference three-dimensional coordinate information at least includes: the three-dimensional coordinate information of each key point in the target object in each of the previous N consecutive video frames predicted based on the combined fusion features by combining the fusion features of each video frame in the previous N consecutive video frames and the fusion features of the original image;
[0102] Optionally, predicting the three-dimensional pose of the target object according to the three-dimensional coordinate information of each key point in the input target object includes:
[0103] Predict the three-dimensional pose of the target object according to the reference three-dimensional coordinate information and the three-dimensional coordinate information of each key point in the target object.
[0104] Optionally, the second deep learning model further includes: a feature map extraction layer;
[0105] The feature map extraction layer is configured to receive the input original image and the object features, and extract the corresponding object feature map from the original image according to the object features.
[0106] Optionally, the three-dimensional coordinate information of the key points is the three-dimensional coordinate information relative to the root node in the three-dimensional coordinate system, and the root node is a specified key point in the target object.
[0107] Correspondingly, an embodiment of the present application further provides a hardware structure diagram of an electronic device, specifically as Figure 7 shown, and the electronic device may be the device for implementing the above method for behavior recognition. As Figure 7As shown, the hardware structure includes: a processor and a memory.
[0108] Among them, the memory is used to store machine-executable instructions;
[0109] The processor is used to read and execute the machine-executable instructions stored in the memory to implement the method embodiments of the corresponding behavior recognition method as shown above.
[0110] As an embodiment, the memory can be any electronic, magnetic, optical or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, the memory can be: volatile memory, non-volatile memory or similar storage media. Specifically, the memory can be RAM (Random Access Memory), flash memory, storage drive (such as a hard disk drive), solid state drive, any type of storage disk (such as an optical disk, DVD, etc.), or similar storage media, or a combination thereof.
[0111] So far, the description of the Figure 7 electronic device shown is completed.
[0112] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A behavior recognition method, characterized in that, The method includes: Inputting the obtained original image captured for a target object into a trained first deep learning model to obtain target feature information of the imaging device that captured the original image, where the target feature information at least includes: the shooting angle when the imaging device captured the target object to obtain the original image, and a predicted image predicted by the imaging device when capturing the target object at a specified angle; Inputting the original image into a trained human detection model to obtain target detection information, where the target detection information at least includes: the two-dimensional coordinates in the original image of at least one key point for indicating the pose in the target object; Fusing the target feature information and the target detection information to obtain a fused feature; the fused feature is used to predict the three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict the three-dimensional pose of the target object; the three-dimensional coordinates of the key point of the target object are obtained by predicting the two-dimensional coordinates of the key point in each plane of the three-dimensional coordinate system according to the fused feature and performing a specified operation on the two-dimensional coordinates of the key point in each plane of the three-dimensional coordinate system; Inputting the fused feature into a trained second deep learning model to obtain the three-dimensional pose of the target object.
2. The method according to claim 1, wherein The shooting angle when the imaging device captured the target object to obtain the original image at least includes: The pitch angle and roll angle set when the imaging device captured the target object to obtain the original image.
3. The method according to claim 1, wherein The target detection information further includes: object features extracted from the original image for indicating the target object; The fusing the target feature information and the target detection information to obtain a fused feature includes: Inputting the original image and the object features into a trained second deep learning model to extract a corresponding object feature map from the original image according to the object features; Fusing the target feature information, the target detection information, and the object feature map to obtain a fused feature.
4. The method according to claim 1, characterized in that The second deep learning model obtains the three-dimensional pose at least through the following calculation layers: The three-dimensional coordinate information prediction layer is used to predict the two-dimensional coordinates of the key point of the target object in each plane of the three-dimensional coordinate system according to the fused feature, perform a specified operation on the two-dimensional coordinates of the key point in each plane of the three-dimensional coordinate system to obtain the three-dimensional coordinate information of the key point, and output the three-dimensional coordinate information of the key point to the three-dimensional pose prediction layer; the three-dimensional coordinate system includes three planes, and the three planes are perpendicular to each other in pairs; The three-dimensional pose prediction layer is used to predict the three-dimensional pose of the target object according to the input three-dimensional coordinate information of each key point in the target object.
5. The method according to claim 4, wherein Predicting the two-dimensional coordinates of the key point of the target object in each plane of the three-dimensional coordinate system according to the fused feature includes: If there are the previous N consecutive video frames before the original image currently, determine the reference three-dimensional coordinate information based on the fusion features of each video frame in the previous N consecutive video frames and the fusion features of the original image, and predict the two-dimensional coordinates of each key point in the target object in each plane of the three-dimensional coordinate system according to the reference three-dimensional coordinate information; the reference three-dimensional coordinate information at least includes: based on the fusion features of each video frame in the previous N consecutive video frames and the fusion features of the original image, the three-dimensional coordinate information of each key point in the target object in each video frame in the previous N consecutive video frames predicted based on the combined fusion features.
6. The method according to claim 5, wherein The predicting the three-dimensional pose of the target object according to the three-dimensional coordinate information of each key point in the input target object includes: Predict the three-dimensional pose of the target object according to the reference three-dimensional coordinate information and the three-dimensional coordinate information of each key point in the target object.
7. The method according to claim 3, wherein The second deep learning model includes: a feature map extraction layer; The feature map extraction layer is configured to receive the input original image and the object feature, and extract the corresponding object feature map from the original image according to the object feature.
8. The method according to any one of claims 1 to 7, characterized in that The three-dimensional coordinate information of the key point is the three-dimensional coordinate information relative to the root node in the three-dimensional coordinate system, and the root node is a specified key point in the target object.
9. An action recognition device, characterized in that, The device includes: A target feature information obtaining unit, configured to input the obtained original image of the target object into the trained first deep learning model to obtain the target feature information of the shooting device for shooting the original image, where the target feature information at least includes: the shooting angle when the shooting device shoots the target object to obtain the original image, and the predicted image predicted by the shooting device shooting the target object at a specified angle; A target detection information obtaining unit, configured to input the original image into the trained humanoid detection model to obtain target detection information, where the target detection information at least includes: the two-dimensional coordinates of at least one key point for indicating the pose in the original image in the target object; A feature fusion unit, configured to fuse the target feature information and the target detection information to obtain a fusion feature; the fusion feature is used to predict the three-dimensional coordinates of at least one key point of the target object, and the three-dimensional coordinates are used to predict the three-dimensional pose of the target object; the three-dimensional coordinates of the key point of the target object are obtained by predicting the two-dimensional coordinates of the key point in each plane of the three-dimensional coordinate system according to the fusion feature and performing a specified operation on the two-dimensional coordinates of the key point in each plane of the three-dimensional coordinate system; A three-dimensional pose prediction unit, configured to input the fusion feature into the trained second deep learning model to obtain the three-dimensional pose of the target object.
10. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory is configured to store machine-executable instructions; The processor is configured to read and execute the machine-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Attitude acquisition method and training method and device of key point coordinate positioning model
CN111126272A
Three-dimensional reconstruction method, device and system, model training method and storage medium
CN111862296A
Target object skeleton key point positioning method and device based on multiple camera devices
CN111951326A