Human body posture estimation method and device, equipment and storage medium

By using a three-dimensional human modeling model including ResNet feature extractor, HMR regressor and SMPL model, the problem of classic HMR framework being insensitive to pose changes is solved, and more accurate recognition of slight pose changes and human pose estimation is achieved.

CN120163874AActive Publication Date: 2025-06-17WUHAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510261158.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-17
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The three-dimensional mannequin output from the classic HMR framework is insensitive to pose changes, resulting in errors in estimating human poses on images with slight pose changes.

Method used

A three-dimensional human body modeling model including the residual network ResNet feature extractor, a hybrid multimodal representation HMR regressor and a skinned multi-person linear SMPL model was used to train multiple human body images in the data set to generate a three-dimensional human body model that is more sensitive to posture changes.

Benefits of technology

The accuracy of the recognition of slight posture changes by the three-dimensional human body modeling model is improved, and human body posture estimation can be achieved more accurately on images with slight posture changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163874A_ABST
    Figure CN120163874A_ABST
Patent Text Reader

Abstract

The invention discloses a human body posture estimation method and device, equipment and a storage medium, and belongs to the technical field of image processing. Training a three-dimensional human body modeling model by adopting the first data set; performing human body posture estimation based on a three-dimensional human body model output by the trained three-dimensional human body modeling model; wherein the three-dimensional human body modeling model comprises a ResNet feature extractor, an HMR regression device and an SMPL model which are connected in sequence, the input of the HMR regression device comprises image features of a first image and bounding box geometric features of a second human body, the output of the HMR regression device comprises joint point features, and the bounding box geometric features are used for converting the joint point features to a first coordinate system; the projection loss of the joint point features is calculated on a first coordinate system, and the first coordinate system is the coordinate system of the original camera. According to the method, human body posture estimation can be accurately carried out on the image with slight posture change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly relates to a human pose estimation method, apparatus, device, and storage medium. Background Art

[0002] Human pose estimation refers to the process of inputting an image and performing pose estimation on the human body in the image to obtain the pose information of the human body in the image.

[0003] In related technologies, the human pose estimation method includes: inputting a picture with a human body into a classical HMR framework to obtain a three-dimensional human body model of the human body, and performing human pose estimation based on the three-dimensional model of the human body. The classical HMR framework includes a convolutional neural network, an HMR regressor, and a standard SMPL (Skinned Multi-Person Linear Model) model connected in sequence.

[0004] However, the three-dimensional human body model output by the classical HMR framework is not sensitive to pose changes, resulting in errors when using the above method to perform human pose estimation on images with slight pose changes. Summary of the Invention

[0005] The present disclosure provides a human pose estimation method, apparatus, device, and storage medium, which can more accurately perform human pose estimation on images with slight pose changes. The technical solutions at least include the following: In a first aspect, a human pose estimation method is provided, including: obtaining a first data set, where the first data set includes multiple images of a first human body; training a three-dimensional human body modeling model using the first data set, and the trained three-dimensional human body modeling model is used to generate a three-dimensional human body model according to a single image input into the three-dimensional human body modeling model; performing human pose estimation based on the three-dimensional human body model output by the trained three-dimensional human body modeling model; where the three-dimensional human body modeling model includes a residual network ResNet feature extractor, a hybrid multi-modal representation HMR regressor, and a skinned multi-person linear SMPL model connected in sequence, the input of the HMR regressor includes the image features of a first image and the bounding box geometric features of a second human body, the output of the HMR regressor includes joint point features, the first image is the image input into the three-dimensional human body modeling model, the second human body is the human body in the first image, the joint point features include multiple joint point coordinates, the bounding box geometric features are used to transform the joint point features onto a first coordinate system to calculate the projection loss of the joint point features on the first coordinate system, and the first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image.

[0006] Optionally, the image of the first human body in the first dataset includes depth RGB images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360°. The ResNet feature extractor includes 6 branches, and each branch is used to extract the feature vector of the depth RGB image at one of the azimuth angles. Training the 3D human body modeling model using the first dataset includes: pre-training the 6 branches of the ResNet feature extractor based on the depth RGB images at different azimuth angles in the first dataset; after the pre-training of the 6 branches of the ResNet feature extractor is completed, freezing the parameters of the 6 branches of the ResNet feature extractor, and at the same time fine-tuning the parameters of the HMR regressor; after fine-tuning the parameters of the HMR regressor, optimizing the parameters of the HMR regressor using the total error loss to train the 3D human body modeling model, and the total error loss includes the projection loss of the joint point features calculated on the first coordinate system.

[0007] Optionally, optimizing the parameters of the HMR regressor using the total error loss includes: establishing a projection chain relationship based on the first root displacement, where the projection chain relationship is used to indicate the process of converting the 3D coordinates in the HMR coordinate system to the 3D coordinates in the first coordinate system and projecting the 3D coordinates in the first coordinate system onto the 2D panoramic image. The HMR coordinate system is the coordinate system of the HMR virtual camera, and the HMR virtual camera is the camera corresponding to the bounding box used to annotate the second human body. The first root displacement is the root displacement between the HMR coordinate system and the first coordinate system, and the first root displacement is determined based on the geometric features of the bounding box of the second human body; obtaining the first joint point features of the second human body, where the first joint point features include the coordinates of multiple joint points of the second human body predicted by the HMR regressor, and the first joint point features are the features in the HMR coordinate system; using the projection chain relationship to project the first joint point features onto the 2D panoramic image; calculating the 2D reprojection loss of the first joint point features; determining the total error loss based on the 2D reprojection loss; and optimizing the parameters of the HMR regressor based on the total error loss to train the 3D human body modeling model.

[0008] Optionally, the geometric features of the bounding box of the second human body are represented by the following formula:

[0009] where, is the geometric feature of the bounding box, is the focal length of the original camera; denotes the coordinates of the center of the bounding box for annotating the second human body in the image coordinate system, where the image coordinate system is a two-dimensional coordinate system with the center of the first image as the origin. denotes the side length of the bounding box for annotating the second human body. is the width of the first image. is the height of the first image.

[0010] Optionally, the first root displacement is obtained using the following formula:

[0011] where is the root displacement of the HMR coordinate system relative to the first coordinate system. is the translation parameter for the weak perspective projection of the HMR coordinate system. , is the focal length of the HMR virtual camera. is the resolution of the bounding box for annotating the second human body. is the scale parameter.

[0012] Optionally, the ResNet feature extractor and the HMR regressor are connected through a dual-path pooling channel attention mechanism. The dual-path pooling channel attention mechanism includes a global average pooling channel and a global maximum pooling channel. The global average pooling channel and the global maximum pooling channel are concatenated through an attention mechanism. The method further includes: inputting the feature vectors output by 6 branches of the ResNet feature extractor into the dual-path pooling channel attention mechanism to obtain the image features of the first image; concatenating the image features of the first image with the bounding box geometric features to obtain a first fusion feature, and the first fusion feature is used for inputting into the HMR regressor.

[0013] In a second aspect, a human body pose estimation device is also provided, including: an acquisition module configured to acquire a first data set, where the first data set includes multiple images of a first human body; a training module configured to train a three-dimensional human body modeling model using the first data set, and the trained three-dimensional human body modeling model is used to generate a three-dimensional human body model based on a single image input into the three-dimensional human body modeling model; a human body pose estimation module configured to perform human body pose estimation based on the three-dimensional human body model output by the trained three-dimensional human body modeling model; wherein, the three-dimensional human body modeling model includes a residual network ResNet feature extractor, a hybrid multi-modal representation HMR regressor, and a skinned multi-person linear SMPL model connected in sequence. The input of the HMR regressor includes the image features of the first image and the bounding box geometric features of a second human body. The output of the HMR regressor includes joint point features. The first image is the image input into the three-dimensional human body modeling model, the second human body is the human body in the first image, the joint point features include multiple joint point coordinates, and the bounding box geometric features are used to transform the joint point features onto a first coordinate system to calculate the projection loss of the joint point features on the first coordinate system. The first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image.

[0014] Optionally, the images of the first human body in the first data set include depth RGB images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360°. The ResNet feature extractor includes 6 branches, and each branch is used to extract a feature vector of the depth RGB image at one of the azimuth angles. The training module is further configured to pre-train the 6 branches of the ResNet feature extractor based on the depth RGB images at different azimuth angles in the first data set; after the 6 branches of the ResNet feature extractor are pre-trained, freeze the parameters of the 6 branches of the ResNet feature extractor, and at the same time fine-tune the parameters of the HMR regressor; after fine-tuning the parameters of the HMR regressor, optimize the parameters of the HMR regressor using the total error loss to train the three-dimensional human body modeling model, and the total error loss includes the projection loss of the joint point features calculated on the first coordinate system.

[0015] Optionally, the training module is further configured to establish a projection chain relationship based on the first root displacement, where the projection chain relationship is used to indicate the process of converting the three-dimensional coordinates in the HMR coordinate system to the three-dimensional coordinates in the first coordinate system and projecting the three-dimensional coordinates in the first coordinate system onto the two-dimensional panoramic image. The HMR coordinate system is the coordinate system of the HMR virtual camera, and the HMR virtual camera is the camera corresponding to the bounding box for annotating the second human body. The first root displacement is the root displacement between the HMR coordinate system and the first coordinate system, and the first root displacement is determined based on the geometric features of the bounding box of the second human body. Obtain the first joint point features of the second human body, where the first joint point features include the coordinates of multiple joint points of the second human body predicted by the HMR regressor, and the first joint point features are the features in the HMR coordinate system. Use the projection chain relationship to project the first joint point features onto the two-dimensional panoramic image. Calculate the two-dimensional reprojection loss of the first joint point features. Determine the total error loss based on the two-dimensional reprojection loss. Optimize the parameters of the HMR regressor based on the total error loss and train the three-dimensional human body modeling model.

[0016] Optionally, the ResNet feature extractor and the HMR regressor are connected through a dual-path pooling channel attention mechanism. The dual-path pooling channel attention mechanism includes a global average pooling channel and a global maximum pooling channel, and the global average pooling channel and the global maximum pooling channel are spliced through an attention mechanism. The training module is further configured to input the feature vectors output by the 6 branches of the ResNet feature extractor into the dual-path pooling channel attention mechanism to obtain the image features of the first image. Splice the image features of the first image with the geometric features of the bounding box to obtain a first fusion feature, and the first fusion feature is used to be input into the HMR regressor.

[0017] In a third aspect, a computer device is further provided, including: a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to perform the human pose estimation method described in the above embodiments.

[0018] In a fourth aspect, a computer-readable storage medium is further provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to perform the human pose estimation method described in the above embodiments.

[0019] In a fifth aspect, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method described in the first aspect is implemented.

[0020] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure at least include: When projecting the coordinates in the three-dimensional space onto the two-dimensional plane, compared with the HMR virtual camera, the original camera is more sensitive to slight changes in the human body posture. Therefore, in the embodiments of the present disclosure, the three-dimensional human body modeling model includes a bounding box geometric feature, which is used to map the joint point coordinates predicted by the HMR regressor to the first coordinate system of the original camera, so as to calculate the projection loss of the joint point features in the first coordinate system. That is, when calculating the projection loss, the projection loss is calculated after projection in the first coordinate system. In this way, the three-dimensional human body modeling model is more sensitive to slight changes in the human body posture, thereby improving the accuracy of the three-dimensional human body modeling model in recognizing slight posture changes. Based on this three-dimensional human body modeling model, it is possible to more accurately perform human body posture estimation on images with slight posture changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 Shows a flowchart of a human body posture estimation method provided by an exemplary embodiment of the present disclosure; Figure 2 Is a schematic diagram of the projection of the joint points of the human body in different postures by the HMR virtual camera and the original camera on the two-dimensional plane; Figure 3 Shows a flowchart of a human body posture estimation method provided by another exemplary embodiment of the present disclosure; Figure 4 Is a schematic diagram of the geometric meaning of the bounding box geometric feature; Figure 5 Is a schematic diagram of the structure of the trained three-dimensional human body modeling model; Figure 6 Is a schematic diagram of a single image input to the three-dimensional human body modeling model and the three-dimensional human body model output by the three-dimensional human body modeling model; Figure 7 Shows a schematic diagram of the structure of a human body posture estimation device provided by an exemplary embodiment of the present disclosure; Figure 8 Is a schematic diagram of the structure of a computer device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second", "third" and similar terms used in the specification and claims of this patent application of the disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or items appearing before "comprising" or "including" cover the elements or items listed after "comprising" or "including" and their equivalents, and do not exclude other elements or items. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0024] To make the objectives, technical solutions and advantages of this disclosure clearer, the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings.

[0025] Figure 1 The flowchart of a human pose estimation method provided by an exemplary embodiment of this disclosure is shown, and this method can be executed by a computer device. Refer to Figure 1 and this method includes: In step 101, a first data set is obtained.

[0026] The first data set includes multiple images of a first human body.

[0027] Here, the first data set is a data set for training a three-dimensional human body modeling model, and the first human body can be any human body.

[0028] Optionally, the images of the first human body in the first data set are depth RGB images. In this case, the acquisition of the first data set is achieved through the following three steps.

[0029] In the first step, multiple RGB images of the first human body are obtained.

[0030] These multiple RGB images are images captured by an RGB camera. The RGB color model is a color standard. The RGB color model obtains various colors through the changes of the three color channels of red (Red), green (Green), and blue (Blue) and the superposition of these three color channels with each other.

[0031] In the second step, the first human body is photographed by an RGBD sensor to obtain multiple RGBD images.

[0032] The RBGD (RGB Depth) sensor is a sensor with depth information, where the depth information is, for example, distance. The RBGD image captured by the RBGD sensor has depth information.

[0033] In the third step, align the depth information in the RBGD image to the RGB image to obtain an enhanced RGB image.

[0034] Optionally, aligning the depth information in the RBGD image to the RGB image includes: using the ICP (Iterative Closest Point) algorithm to achieve spatial registration of the RGB image and the RBGD image; applying the same geometric transformation matrix to the RGB image and the RBGD image to maintain the projection relationship unchanged; and finally using a hole filling algorithm based on Conditional Random Field (CRF) to repair invalid depth regions.

[0035] The process of aligning the depth information in the RBGD image to the RGB image is actually a process of enhancing the RGB image so that the RGB image has depth information.

[0036] In the related art, there is a way to obtain a depth RGB image by using an RGBD camera. However, existing RGBD cameras (such as Kinect) have a fixed depth resolution within a certain distance range, and the accuracy of the depth information may not be ideal at long distances or in complex environments.

[0037] By adopting the above first to third steps to obtain a depth RGB image, it is possible to flexibly adjust the accuracy of the RGB camera and the RGBD sensor. For example, a higher-precision RGB camera can be replaced when the accuracy of the RGB camera is insufficient, or a higher-precision RGBD sensor can be replaced when the accuracy of the RGBD sensor is insufficient. Thus, different accuracy requirements for the depth RGB image can be met.

[0038] Optionally, the following method can also be used to enhance the images in the first dataset: horizontally flipping the images, vertically flipping the images, randomly scaling the images, where the scaling ratio ranges from 0.85 to 1.25, and randomly rotating the images, where the rotation range is from -20° to 20°.

[0039] In step 102, a three-dimensional human body modeling model is trained using the first dataset.

[0040] The trained 3D human body modeling model is used to generate a 3D human body model according to a single image input into the 3D human body modeling model. The conventional HMR framework can output a 3D human body model through a single image. In the embodiments of the present disclosure, the HMR framework is improved, and the improved HMR framework is the 3D human body modeling model. Therefore, the 3D human body modeling model can generate a 3D human body model according to a single image input into the 3D human body modeling model.

[0041] Among them, the 3D human body modeling model includes a ResNet (Residual Network) feature extractor, a hybrid multi-modal representation HMR regressor, and a skinned multi-person linear SMPL model connected in sequence. The input of the HMR regressor includes the image features of the first image and the bounding box geometric features of the second human body. The output of the HMR regressor includes joint point features. The first image is the image input into the 3D human body modeling model, and the second human body is the human body in the first image. The joint point features include multiple joint point coordinates. The bounding box geometric features are used to transform the joint point features to the first coordinate system to calculate the projection loss of the joint point features in the first coordinate system. The first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image. Here, the first image is the image input into the 3D human body modeling model, and the first image includes at least one human body. That is, in the process of training and testing the 3D human body modeling model, the first image is the image in the first dataset. When performing human pose estimation based on the 3D human body model output by the trained 3D human body modeling model, the first image is the image outside the first dataset (for example, the image for which human pose estimation is required). Similarly, the second human body is the human body in the first image for which 3D human body modeling is required. In the process of training and testing the 3D human body modeling model, the second human body is the first human body; when performing human pose estimation based on the 3D human body model output by the trained 3D human body modeling model, the second human body is the human body other than the first human body.

[0042] Here, the original camera refers to the camera when the first image is taken. The original camera is on a straight line passing through the center of the first image and perpendicular to the first image, and the distance between the original camera and the center of the first image is the focal length of the original camera.

[0043] Under normal circumstances, when the HMR regressor calculates the projection loss, it does not involve the coordinate system of the original camera. Instead, a HMR virtual camera is generated for the cropped image, and after prediction based on the HMR coordinate system of the HMR virtual camera, the projection loss is directly calculated based on the prediction result. The HMR virtual camera is on a straight line passing through the center of the cropped image (i.e., the annotation box) and perpendicular to the cropped image, and the distance between the HMR virtual camera and the center of the cropped image is the focal length of the HMR virtual camera.

[0044] When projecting the coordinates in three-dimensional space onto a two-dimensional plane, compared with the HMR virtual camera, the original camera is more sensitive to slight changes in human body postures. The following combines Figure 2 to illustrate this process.

[0045] Figure 2 is a schematic diagram of the projections of the human body joint points of different postures by the HMR virtual camera and the original camera on the two-dimensional plane. Figure 2 Part (a) of Figure 2 is a schematic diagram of the first posture and the second posture. As shown in part (a) of

[0046] When the HMR regressor regresses, there must be a process of projecting the predicted three-dimensional coordinates of the human body joint points onto the two-dimensional plane. Figure 2 Part (b) of Figure 2 is a schematic diagram of the projections of the human body joint points of the first posture 201 and the second posture 202 by the HMR virtual camera and the original camera on the two-dimensional plane respectively. As shown in part (b) of

[0047] Here, Figure 2 in part (b) of

[0048] It can be seen that when projecting in the HMR coordinate system of the HMR virtual camera 203, the projection line segment a corresponding to the connection line of the shoulder joint points of the first pose 201 and the projection line segment b corresponding to the connection line of the shoulder joint points of the second pose 202 are relatively close. This situation makes the HMR virtual camera insensitive to slight pose changes, and it is very easy to recognize the first pose 201 and the second pose 202 as the same pose, resulting in the generation of two identical 3D human models in the end. When projecting in the first coordinate system of the original camera 204, there is a large difference between the projection line segment c corresponding to the connection line of the shoulder joint points of the first pose 201 and the projection line segment d corresponding to the connection line of the shoulder joint points of the second pose 202, that is, the difference between c and d is greater than the difference between a and b. That is to say, in the first coordinate system of the original camera 204, the difference brought about by slight pose changes is more obvious, or in the first coordinate system of the original camera 204, the difference brought about by slight pose changes will be amplified.

[0049] Therefore, compared with the HMR coordinate system of the HMR virtual camera, the first coordinate system of the original camera is more sensitive to slight changes in human poses, and thus can more accurately identify different poses with only slight changes. That is to say, when projecting the 3D coordinates in the first coordinate system onto the 2D plane, the spatial offset of the shoulder bone points can be presented more accurately, thereby significantly improving the view sensitivity and making up for the deficiency in geometric priors.

[0050] In this way, by inputting the bounding box geometric features into the HMR regressor, the HMR regressor can understand the difference between the HMR coordinate system and the first coordinate system when predicting joint point features. The bounding box geometric features are also used to map the joint point coordinates predicted by the HMR regressor onto the first coordinate system of the original camera to calculate the projection loss of the joint point features in the first coordinate system, that is, when calculating the projection loss, the projection is performed in the first coordinate system. In this way, the 3D human body modeling model is more sensitive to slight changes in human poses, and thus improves the accuracy of the 3D human body modeling model in recognizing slight pose changes. That is to say, the 3D human body modeling model can accurately establish a 3D human body model.

[0051] In step 103, human pose estimation is performed based on the 3D human body model output by the trained 3D human body modeling model.

[0052] After the 3D human body modeling model is trained, the 3D human body modeling model can restore a relatively accurate 3D human body model based on a single image, and accurate human pose estimation can be performed based on this accurate 3D human body model.

[0053] When projecting the coordinates in three-dimensional space onto a two-dimensional plane, compared with the HMR virtual camera, the original camera is more sensitive to slight changes in human poses. Therefore, in the embodiments of the present disclosure, the three-dimensional human body modeling model includes a bounding box geometric feature, which is used to map the joint point coordinates predicted by the HMR regressor onto the first coordinate system of the original camera, so as to calculate the projection loss of the joint point features in the first coordinate system. That is, when calculating the projection loss, the projection is performed in the first coordinate system and then the projection loss is calculated. In this way, the three-dimensional human body modeling model is more sensitive to slight changes in human poses, thereby improving the accuracy of the three-dimensional human body modeling model in recognizing slight pose changes. Based on this three-dimensional human body modeling model, it is possible to more accurately perform human pose estimation on images with slight pose changes.

[0054] Figure 3 FIG. shows a flowchart of a human pose estimation method provided by another exemplary embodiment of the present disclosure, which can be executed by a computer device. Refer to Figure 3 , the method includes: In step 301, a first data set is obtained.

[0055] For the relevant content of obtaining the first data set, refer to the foregoing step 101.

[0056] Based on the multiple enhanced RGB images in the foregoing step 101, the first human body in the RGB image can be modeled.

[0057] The RGB images captured by the RGB camera have limited perspectives, but the first human body can be accurately modeled based on these RGB images to obtain a three-dimensional human body model of the first human body. Then, the three-dimensional human body model of the first human body can be photographed using a virtual camera, and the RGB images obtained by the virtual camera photographing are also stored in the first data set (image enhancement processing is also required, such as integrating depth information into the RGB image). Since the virtual camera can be set at any angle and any focal length, theoretically, RGB images of the first human body at any angle can be obtained.

[0058] In this case, step 301 further includes: constructing a virtual camera array; using the virtual cameras in the virtual camera array to perform multi-view rendering and photographing on the three-dimensional human body model of the first human body to generate synthetic RGB images and corresponding pose information.

[0059] Among them, the parameter configuration of the virtual camera array includes: the azimuth angle ranges from 0° to 360°, and one column of virtual cameras is set every 60° within this azimuth angle range; in any column of virtual cameras, the pitch angle of this column of virtual cameras varies from -30° to 60°, and one virtual camera is set every 10° for the pitch angle. In addition, the focal length of any virtual camera in this virtual camera array can randomly vary within the range of 35mm ± 20%.

[0060] In this way, a first data set can be constructed. The first data set includes multiple images of the first human body, and each image is a depth RGB image. When divided according to the azimuth angle, the images of the first human body in the first data set include the depth RGB images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360° respectively.

[0061] Optionally, corresponding to 6 azimuth angles, the ResNet feature extractor has 6 branches, and each branch is a ResNet-50 network. Each branch is used to extract the feature vector of the depth RGB image in the case of one azimuth angle. During the process of training the three-dimensional human body modeling model using the first data set, these 6 branches are respectively used to extract the features of the images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360°. After the ResNet feature extractor is trained, the images at different angles of the ResNet feature extractor all have good feature extraction capabilities.

[0062] In step 302, based on the depth RGB images of different azimuth angles in the first data set, pre-train the 6 branches of the ResNet feature extractor.

[0063] Before performing step 302, first explain each part of the three-dimensional human body modeling model. The three-dimensional human body modeling model includes a ResNet feature extractor, an HMR regressor, and an SMPL model connected in sequence.

[0064] The ResNet feature extractor includes 6 branches, and the input of the ResNet feature extractor is the first image. When connected by a dual-path pooling channel attention mechanism between the ResNet feature extractor and the HMR regressor, the output of the ResNet feature extractor is the 6 feature vectors of the first image extracted by the 6 branches, and the 6 feature vectors of the first image are processed by the dual-path pooling channel attention mechanism to obtain the image features of the first image.

[0065] The input of the HMR regressor includes the image features of the first image and the bounding box geometric features of the second human body, and the output of the HMR regressor is the joint point features.

[0066] The input of the SMPL model is joint point features, and the output of the SMPL model is a three-dimensional human body model of a second person.

[0067] In the embodiments of the present disclosure, first, a ResNet feature extractor is pre-trained. After the pre-training of the ResNet feature extractor is completed, the parameters of the ResNet feature extractor are frozen, and the parameters of the HMR regressor are fine-tuned. After the parameters of the HMR regressor are fine-tuned, finally, the total error loss is used to optimize the parameters of the HMR regressor, so as to implement the training of the three-dimensional human body modeling model.

[0068] In this case, optionally, the loss function of the 6 branches of the pre-trained ResNet feature extractor is represented by formula (1).

[0069] (1) In formula (1), is the loss function used for each branch when pre-training the ResNet feature extractor, is the feature vector extracted by the branch corresponding to the th azimuth angle, where is the height of the feature vector, is the width of the feature vector, is the feature vector extracted by the branch corresponding to the is an integer, The value range of is from 1 to 6. The 1st azimuth angle, the 2nd azimuth angle... the 6th azimuth angle represent 0°, 60°,..., 360° respectively. represents any one of the other 5 azimuth angles except the th azimuth angle, represents the feature vector extracted by the branch corresponding to the represents calculating the and temperature-scaled cosine similarity between, represents calculating the and temperature-scaled cosine similarity between.

[0070] Exemplarily, and the temperature-scaled cosine similarity between is calculated by formula (2).

[0071] (2) In formula (2), is the temperature scaling parameter, which is a learnable parameter. Initially it can be set to 0.07. The meanings of other parameters in formula (2) are the same as those in formula (1), which are not elaborated here.

[0072] Through formula (1), the adjacent view features ( and ) maintain the maximum mutual information in the cosine similarity space, while pushing away the correlation of non-adjacent view features. Therefore, the view sensitivity problem in single-view feature extraction can be well solved by pre-training the ResNet feature extractor, ensuring that the extracted features have geometric consistency.

[0073] When using formula (1) as the loss function to pre-train the ResNet feature extractor, the gradient descent algorithm and Adam optimizer can be used to optimize the parameters of the ResNet feature extractor.

[0074] After the pre-training of the 6 branches of the ResNet feature extractor is completed, step 303 can be executed.

[0075] In step 303, the parameters of the 6 branches of the ResNet feature extractor are frozen, and at the same time, the parameters of the HMR regressor are fine-tuned.

[0076] After the pre-training of the multi-view feature extraction network, the convolutional layer parameters of the 6 branches of the ResNet feature extractor can be frozen (the batch normalization layer is reserved for updating), and then the Adam optimizer with an initial learning rate of can be used to fine-tune the parameters of the HMR regressor, and an exponential decay strategy of 2% per training cycle is implemented, that is, the decay formula is .

[0077] Optionally, in the 3D human body modeling model, the ResNet feature extractor and the HMR regressor are connected through a dual-path pooling channel attention mechanism. The dual-path pooling channel attention mechanism includes a global average pooling channel and a global maximum pooling channel, and the global average pooling channel and the global maximum pooling channel are spliced through the attention mechanism.

[0078] In this case, optionally, step 303 includes the following two steps.

[0079] First, the feature vectors output by the 6 branches of the ResNet feature extractor are input into the dual-path pooling channel attention mechanism to obtain the image features of the first image; The first step can be represented by formulas (3) to (4).

[0080] (3) (4) In formula (3), is 's weight, denotes global average pooling on ; , denotes global max pooling on ; , is the weight matrix of the dual-path pooling channel attention mechanism, which is a learnable parameter, . denotes the concatenation operation along the channel dimension, is the sigmoid activation function.

[0081] In formula (4), is the image feature of the first image. The meanings of other parameters in formulas (3) and (4) are the same as those in the aforementioned formula (1), and are not elaborated here.

[0082] Second, concatenate the image feature of the first image with the bounding box geometric feature to obtain the first fusion feature.

[0083] The first fusion feature is used as the input to the HMR regressor. Based on this first fusion feature, the HMR regressor can be fine-tuned.

[0084] In this case, the input of this HMR regressor can be expressed by formula (5).

[0085] (5) In formula (5), is the parameter input to the HMR regressor, is the weight matrix of the HMR regressor, is the bounding box geometric feature, is the first fusion feature, is the bias of the HMR regressor. , are both learnable parameters. The meanings of other parameters in formula (5) are the same as those in formula (4), and are not elaborated here.

[0086] Bounding box geometric feature can be expressed by formula (6).

[0087] (6) In formula (6), is the bounding box geometric feature, is the focal length of the original camera; Denote the coordinates of the center of the bounding box for annotating the second human body in the image coordinate system, which is a two-dimensional coordinate system with the center of the first image as the origin. Denote the side length of the bounding box for annotating the second human body. is the width of the first image. is the height of the first image.

[0088] Figure 4 is a schematic diagram of the geometric meaning of the geometric features of the bounding box. The following combines Figure 4 to explain the parameters in the geometric features of the bounding box.

[0089] As Figure 4 shown, the width of the first image 401 is , and the length is . The image coordinate system xO1y takes the center O1 of the first image 401 as the origin.

[0090] The center of the bounding box 402 for annotating the second human body is O2. The bounding box 402 is a square bounding box, and the side length of the bounding box 402 is b. In the image coordinate system xO1y, the coordinates of the center O2 of the bounding box 402 can be expressed as .

[0091] In the HMR coordinate system 403, O3 is the origin of the HMR coordinate system 403, and the position of O3 is the position of the HMR virtual camera. O3O2 is a straight line perpendicular to the bounding box 402 and passes through the center O2 of the bounding box 402. The length of O3O2 is the focal length of the HMR virtual camera.

[0092] In the first coordinate system 404, O4 is the origin of the first coordinate system 404, and the position of O4 is the position of the original camera. O4O1 is a straight line perpendicular to the first image 401 and passes through the center O1 of the first image 401. The length of O4O1 is the focal length value.

[0093] In the geometric features of the bounding box , The first two terms of and have geometric meanings in terms of angles. As shown in the figure, represents the tangent value of the angle O4O1, represents the tangent value of the angle O4O1. Through these two terms, it helps the HMR regressor understand the angular relationship between the coordinates in the bounding box 402 and the coordinates of the original camera. The third term of and the fourth term It can reflect the proportion of the bounding box 402 in the first image (in the third and fourth items which plays a role in normalization), so that in a multi-resolution scenario (for example, images come from different cameras or different video streams), the HMR regressor can better understand the proportion of the person in the image and learn additional global scale information.

[0094] In some cases, the focal length of the original camera is a known true value. At this time, the true value can be directly adopted in . When the true value of the focal length of the original camera is unknown, then an approximate estimated value of the focal length of the original camera can be adopted. This approximate estimated value is represented by formula (7).

[0095] (7) The meanings of the parameters in formula (7) are the same as those in formula (6), and the detailed description is omitted here.

[0096] By fine-tuning the HMR regressor in step 303, the parameters of a preliminary HMR regressor can be obtained. At this time, the HMR regressor can initially understand the relationship between the original camera and the HMR virtual camera. However, the joint points predicted by the HMR regressor are still the coordinates in the HMR coordinate system. Subsequently, the coordinates of the HMR coordinate system predicted by the HMR regressor need to be converted to the first coordinate system, and then based on step 304, that is, based on the coordinates in the first coordinate system, the parameters of the HMR regressor are further optimized.

[0097] In step 304, the total error loss is used to optimize the parameters of the HMR regressor.

[0098] The total error loss includes the projection loss of the joint point features calculated on the first coordinate system.

[0099] Optionally, step 304 includes the following steps a-f.

[0100] Step a, based on the first root displacement, establish a projection chain relationship.

[0101] The projection chain relationship is used to indicate the process of converting the three-dimensional coordinates in the HMR coordinate system to the three-dimensional coordinates in the first coordinate system and projecting the three-dimensional coordinates in the first coordinate system onto the two-dimensional panoramic view. The HMR coordinate system is the coordinate system of the HMR virtual camera, and the HMR virtual camera is the camera corresponding to the bounding box for annotating the second person. The first root displacement is the root displacement between the HMR coordinate system and the first coordinate system, and the first root displacement is determined based on the geometric features of the bounding box of the second person.

[0102] Optionally, the first root displacement includes root displacements on the X-axis, Y-axis, and Z-axis, and the first root displacement can be expressed as , where is the root displacement on the X-axis in the first root displacement, is the root displacement on the Y-axis in the first root displacement, is the root displacement on the Z-axis in the first root displacement. In this case, the first root displacement is represented by formula (8).

[0103] (8) In formula (8), is the root displacement of the HMR coordinate system relative to the first coordinate system, is the translation parameter for the weak perspective projection of the HMR coordinate system, , is the focal length of the HMR virtual camera, is the resolution of the bounding box used to label the second human body, is the scale parameter. The meanings of other parameters in formula (8) are the same as those in formula (6) and are not elaborated here.

[0104] Among them, and are both predefined parameters. Exemplarily, , .

[0105] In this case, the projection chain relationship is represented by formula (9).

[0106] (9) In formula (9), represents the coordinates of the second human body projected onto the two-dimensional panoramic view, represents the three-dimensional coordinates of the second human body predicted by the HMR regressor in the HMR coordinate system, which is also the first joint point feature in step b, is the first root displacement, calculated using formula (8); is the three-dimensional coordinates of the second human body in the first coordinate system, represents the process of converting the three-dimensional coordinates in the HMR coordinate system to the three-dimensional coordinates in the first coordinate system process. represents the process of projecting the three-dimensional coordinates onto the two-dimensional coordinates.

[0107] Step b, obtaining the first joint point feature of the second human body.

[0108] The first joint point feature includes the coordinates of multiple joint points of the second human body predicted by the HMR regressor, and the first joint point feature is a feature in the HMR coordinate system.

[0109] Step c: Using the projection chain relationship, project the first joint point feature onto the two-dimensional panoramic image.

[0110] That is, input the first joint point feature into Equation (9) to obtain .

[0111] Step d: Calculate the two-dimensional reprojection loss of the first joint point feature.

[0112] Optionally, Step d is represented by Equation (10).

[0113] (10) In Equation (10), is the two-dimensional reprojection loss, is the true value of the coordinates of the second human body projected onto the two-dimensional panoramic image. The meanings of other parameters in Equation (10) are the same as those in Equation (9) and are not elaborated here.

[0114] The two-dimensional reprojection loss can accurately reflect the projection deviation of the estimation error of the three-dimensional coordinates on the real imaging plane, thereby providing a supervision signal that conforms to the actual imaging geometry for the three-dimensional human body modeling model.

[0115] The first root displacement is calculated using Equation (8), and multiple parameters involved in Equation (8) are all parameters in the bounding box geometric features. It can be seen from Equation (9) that based on the first root displacement, the joint point features in the HMR coordinate system predicted by the HMR regressor can be converted to the first coordinate system, and then the two-dimensional reprojection loss can be calculated based on Equation (10). Therefore, the bounding box geometric features play an important role in converting the multiple joint point coordinates predicted by the HMR regressor to the first coordinate system and calculating the projection loss of the joint point features on the first coordinate system.

[0116] Step e: Determine the total error loss based on the two-dimensional reprojection loss.

[0117] Optionally, the total error loss is represented by Equation (11).

[0118] (11) In Equation (11), is the total error loss.

[0119] In the three-dimensional human body modeling model, the first joint point feature predicted by the HMR regressor can be input into the SMPL model to reconstruct a three-dimensional human body model. Compared with the standard SMPL model, there may be errors in the reconstructed three-dimensional human body model. That is, the error between the 3D human body model reconstructed based on the first joint point features of the 3D human body modeling model and the standard SMPL model.

[0120] The first joint point features predicted by the HMR regressor are 3D joint points, which are the predicted values of the HMR regressor. There may also be an error between the predicted values and the true values. That is, the error between the first joint point features predicted by the HMR regressor and the true values of the first joint point features.

[0121] , , are weights, where is 's weight, is 's weight, is 's weight. The values of these three weights can be set according to experience, and the sum of these three weights is 1. Exemplarily, , , .

[0122] Step f: Optimize the parameters of the HMR regressor based on the total error loss to train the 3D human body modeling model.

[0123] The total error loss is the objective function. When training the 3D human body modeling model, aiming to minimize the total error loss, the parameters of the HMR regressor can be gradually optimized. After reaching the training goal, the training can be stopped, thereby obtaining the trained 3D human body modeling model.

[0124] In terms of structure, the difference between the 3D human body modeling model in the embodiments of the present disclosure and the conventional HMR architecture is that in the embodiments of the present disclosure, the convolutional neural network is replaced by a Resnet feature extractor, and the Resnet feature extractor is connected to the HMR regressor through a dual-path pooling channel attention mechanism. And the input of the HMR regressor in the embodiments of the present disclosure includes the bounding box geometric features.

[0125] During the training process, compared with the conventional HMR architecture, the differences in the training processes of the 3D human body modeling model in the embodiments of the present disclosure include that when calculating the 2D reprojection loss in the embodiments of the present disclosure, first, based on the bounding box geometric features, the joint point features in the HMR coordinate system predicted by the HMR regressor are converted to the first coordinate system, and then the joint point features in the first coordinate system are subjected to 2D reprojection calculation. Since when projecting the coordinates in the 3D space onto the 2D plane, compared with the HMR virtual camera, the original camera is more sensitive to slight changes in the human body posture, thus this can improve the accuracy of the 3D human body modeling model in recognizing slight posture changes.

[0126] In step 305, human pose estimation is performed based on the three-dimensional human model output by the trained three-dimensional human modeling model.

[0127] Figure 5 It is a schematic structural diagram of the trained three-dimensional human modeling model. As Figure 5 shown, after a single image is input into the ResNet feature extractor 501, the ResNet feature extractor 501 first extracts a feature vector from the single image, and then the feature vector is processed by the dual-path pooling channel attention mechanism to obtain the image feature of the single image. The first fusion feature 504 obtained by splicing the image feature and the bounding box geometric feature is input into the hybrid multi-modal representation regressor 502 (i.e., the HMR regressor). The input of the HMR regressor 502 includes the first fusion feature, and the output of the HMR regressor 502 includes the predicted joint point features. The HMR regressor 502 inputs the predicted joint point features into the skinned multi-person linear model 503 (i.e., the SMPL model), and the SMPL model can output a three-dimensional human model.

[0128] Optionally, before the HMR regressor 502 inputs the predicted joint point features into the skinned multi-person linear model 503, an adaptive weighting algorithm can also be used to optimize each joint point feature predicted by the HMR regressor to further improve the reliability of the established three-dimensional human model.

[0129] There are many implementation methods of the adaptive weighting algorithm in the related art, which are omitted here for detailed description.

[0130] Figure 6 It is a schematic diagram of a single image input into the three-dimensional human modeling model and the three-dimensional human model output by the three-dimensional human modeling model. Through Figure 6 it can be seen that the three-dimensional human model established by the three-dimensional human modeling model is relatively accurate.

[0131] The following is the device embodiment of the present application. For the details not described in detail in the device embodiment, reference can be made to the above method embodiment.

[0132] Figure 7 It shows a schematic structural diagram of a human pose estimation device provided by an exemplary embodiment of the present disclosure. Refer to Figure 7 , the human pose estimation device 700 includes: an acquisition module 701, a training module 702, and a human pose estimation module 703.

[0133] The acquisition module 701 is used to acquire a first data set, and the first data set includes multiple images of a first human body.

[0134] The training module 702 is used to train a three-dimensional human body modeling model using a first data set. The trained three-dimensional human body modeling model is used to generate a three-dimensional human body model based on a single image input into the three-dimensional human body modeling model.

[0135] The human body pose estimation module 703 is used to perform human body pose estimation based on the three-dimensional human body model output by the trained three-dimensional human body modeling model. Among them, the three-dimensional human body modeling model includes a residual network ResNet feature extractor, a hybrid multi-modal representation HMR regressor, and a skinned multi-person linear SMPL model connected in sequence. The input of the HMR regressor includes the image features of the first image and the bounding box geometric features of the second human body. The output of the HMR regressor includes joint point features. The first image is the image input into the three-dimensional human body modeling model, and the second human body is the human body in the first image. The joint point features include multiple joint point coordinates. The bounding box geometric features are used to transform the joint point features onto the first coordinate system to calculate the projection loss of the joint point features on the first coordinate system. The first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image.

[0136] Optionally, the images of the first human body in the first data set include depth RGB images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360°. The ResNet feature extractor includes 6 branches, and each branch is used to extract the feature vector of the depth RGB image at one azimuth angle. The training module 702 is also used to pre-train the 6 branches of the ResNet feature extractor based on the depth RGB images at different azimuth angles in the first data set. After the 6 branches of the ResNet feature extractor are pre-trained, the parameters of the 6 branches of the ResNet feature extractor are frozen, and at the same time, the parameters of the HMR regressor are fine-tuned. After fine-tuning the parameters of the HMR regressor, the parameters of the HMR regressor are optimized using the total error loss to train the three-dimensional human body modeling model. The total error loss includes the projection loss of the joint point features calculated on the first coordinate system.

[0137] Optionally, the training module 702 is further configured to establish a projection chain relationship based on the first root displacement. The projection chain relationship is used to indicate the process of converting the three-dimensional coordinates in the HMR coordinate system to the three-dimensional coordinates in the first coordinate system and projecting the three-dimensional coordinates in the first coordinate system onto the two-dimensional panoramic image. The HMR coordinate system is the coordinate system of the HMR virtual camera, and the HMR virtual camera is the camera corresponding to the bounding box for annotating the second human body. The first root displacement is the root displacement between the HMR coordinate system and the first coordinate system, and the first root displacement is determined based on the geometric features of the bounding box of the second human body. Obtain the first joint point features of the second human body. The first joint point features include the coordinates of multiple joint points of the second human body predicted by the HMR regressor, and the first joint point features are features in the HMR coordinate system. Use the projection chain relationship to project the first joint point features onto the two-dimensional panoramic image. Calculate the two-dimensional reprojection loss of the first joint point features. Determine the total error loss based on the two-dimensional reprojection loss. Optimize the parameters of the HMR regressor based on the total error loss to train the three-dimensional human body modeling model.

[0138] Optionally, the ResNet feature extractor and the HMR regressor are connected through a dual-path pooling channel attention mechanism. The dual-path pooling channel attention mechanism includes a global average pooling channel and a global maximum pooling channel. The global average pooling channel and the global maximum pooling channel are spliced through an attention mechanism. The training module 702 is further configured to input the feature vectors output by the 6 branches of the ResNet feature extractor into the dual-path pooling channel attention mechanism to obtain the image features of the first image. Splice the image features of the first image with the geometric features of the bounding box to obtain the first fusion feature, and the first fusion feature is used to be input into the HMR regressor.

[0139] It should be noted that when the human pose estimation device provided in the above embodiment performs human pose estimation, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the human pose estimation device provided in the above embodiment and the human pose estimation method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.

[0140] The division of modules in the embodiments of the present disclosure is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present disclosure, each functional module can be integrated in one processor, or can exist separately physically, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0141] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which may be a personal computer, a mobile phone, or a communication device, etc.) or a processor to execute all or part of the steps of the method according to various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0142] Figure 8 is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. As Figure 8 shown, the computer device 800 includes: a processor 801 and a memory 802.

[0143] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0144] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one instruction for being executed by the processor 801 to implement the human pose estimation method provided in the embodiments of the present disclosure.

[0145] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0146] The embodiments of the present disclosure also provide a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the computer device, the computer device can execute the human pose estimation method provided in the embodiments of the present disclosure.

[0147] The embodiments of the present disclosure also provide a computer program product, including a computer program / instructions. When the computer program / instructions are executed by the processor, the human pose estimation method provided in the embodiments of the present disclosure is implemented.

[0148] The foregoing are only optional embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for estimating a human body posture, characterized in that: The method comprises: Acquire a first data set, the first data set comprising a plurality of images of a first human body; Using the first data set to train a three-dimensional human body modeling model, the trained three-dimensional human body modeling model is used to generate a three-dimensional human body model according to a single image input into the three-dimensional human body modeling model; Performing human body posture estimation based on the three-dimensional human body model output by the trained three-dimensional human body modeling model; Among them, the three-dimensional human body modeling model includes a residual network ResNet feature extractor, a hybrid multimodal representation HMR regressor and a skinned multi-person linear SMPL model connected in sequence, the input of the HMR regressor includes image features of a first image and bounding box geometric features of a second human body, the output of the HMR regressor includes joint point features, the first image is an image input to the three-dimensional human body modeling model, the second human body is a human body in the first image, the joint point features include multiple joint point coordinates, and the bounding box geometric features are used to convert the joint point features to a first coordinate system to calculate the projection loss of the joint point features on the first coordinate system, the first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image.

2. The method according to any one of claims 1 to 4, characterized in that: The image of the first human body in the first data set includes the depth RGB images of the first human body at azimuth angles of 0°, 60°, 120°, 180°, 240°, 300°, and 360°, respectively, and the ResNet feature extractor includes 6 branches, each branch is used to extract a feature vector of the depth RGB image at the azimuth angle, The step of using the first data set to train a three-dimensional human body modeling model comprises: Pre-training six branches of the ResNet feature extractor based on the deep RGB images of different azimuths in the first data set; After the pre-training of the six branches of the ResNet feature extractor is completed, freezing the parameters of the six branches of the ResNet feature extractor and fine-tuning the parameters of the HMR regressor; After fine-tuning the parameters of the HMR regressor, the parameters of the HMR regressor are optimized using a total error loss to train the three-dimensional human body modeling model, wherein the total error loss includes a projection loss of the joint point features calculated on the first coordinate system.

3. The method according to claim 2, characterized in that The method of optimizing the parameters of the HMR regressor using the total error loss includes: Based on the first root displacement, a projection chain relationship is established, where the projection chain relationship is used to indicate a process of converting a three-dimensional coordinate in an HMR coordinate system into a three-dimensional coordinate in the first coordinate system, and projecting the three-dimensional coordinate in the first coordinate system onto a two-dimensional panoramic image, where the HMR coordinate system is a coordinate system of an HMR virtual camera, where the HMR virtual camera is a camera corresponding to a bounding box of the second human body, where the first root displacement is a root displacement between the HMR coordinate system and the first coordinate system, and where the first root displacement is determined based on geometric features of a bounding box of the second human body; Acquire a first joint feature of the second human body, where the first joint feature includes coordinates of multiple joints of the second human body predicted by an HMR regressor, and the first joint feature is a feature in the HMR coordinate system; Using the projection chain relationship, projecting the first joint point feature onto a two-dimensional panoramic image; Calculating the two-dimensional reprojection loss of the first joint point feature; Determining a total error loss based on the two-dimensional reprojection loss; Based on the total error loss, the parameters of the HMR regressor are optimized to train the three-dimensional human body modeling model.

4. The method according to claim 3, characterized in that The geometric features of the second human body's bounding box are expressed by the following formula: in, is the geometric feature of the bounding box, is the focal length of the original camera; represents the coordinates of the center of the bounding box used to mark the second human body in the image coordinate system, where the image coordinate system is a two-dimensional coordinate system with the center of the first image as the origin, represents the side length of the bounding box used to annotate the second human body, is the width of the first image, is the height of the first image.

5. The method according to claim 4, characterized in that The first root displacement is obtained using the following formula: in, is the root displacement of the HMR coordinate system relative to the first coordinate system, is the translation parameter for weak perspective projection of the HMR coordinate system, , is the focal length of the HMR virtual camera, is the resolution of the bounding box used to annotate the second human body, is the scale parameter.

6. The method according to any one of claims 2 to 5, characterized in that: The ResNet feature extractor is connected to the HMR regressor through a dual-path pooling channel attention mechanism, wherein the dual-path pooling channel attention mechanism includes a global average pooling channel and a global maximum pooling channel, wherein the global average pooling channel and the global maximum pooling channel are spliced ​​through the attention mechanism, The method further comprises: Inputting the feature vectors output by the six branches of the ResNet feature extractor into the dual-path pooling channel attention mechanism to obtain image features of the first image; The image features of the first image are concatenated with the geometric features of the bounding box to obtain a first fused feature, where the first fused feature is input into the HMR regressor.

7. A human body posture estimation device, characterized in that: The device comprises: An acquisition module, configured to acquire a first data set, wherein the first data set includes a plurality of images of a first human body; A training module, used for training a three-dimensional human body modeling model using the first data set, wherein the trained three-dimensional human body modeling model is used for generating a three-dimensional human body model according to a single image input into the three-dimensional human body modeling model; A human body posture estimation module, used for performing human body posture estimation based on the three-dimensional human body model output by the trained three-dimensional human body modeling model; Among them, the three-dimensional human body modeling model includes a residual network ResNet feature extractor, a hybrid multimodal representation HMR regressor and a skinned multi-person linear SMPL model connected in sequence, the input of the HMR regressor includes image features of a first image and bounding box geometric features of a second human body, the output of the HMR regressor includes joint point features, the first image is an image input to the three-dimensional human body modeling model, the second human body is a human body in the first image, the joint point features include multiple joint point coordinates, and the bounding box geometric features are used to convert the joint point features to a first coordinate system to calculate the projection loss of the joint point features on the first coordinate system, the first coordinate system is the coordinate system of the original camera, and the original camera is the camera corresponding to the first image.

8. A computer device, characterized in that: The computer device comprises: a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Object attitude tracking method and device, terminal equipment and storage medium

    CN113298870A

  • Fall detection method and system based on monocular three-dimensional human body posture

    CN113378809A

  • Cross-modal weak supervision three-dimensional human body posture estimation method and system

    CN115565203A

  • Three-dimensional crowd data generation method based on single-view color image

    CN117079066A

  • Demand response program operating system for apartment customers and operating method using it

    KR1020220162424A