Human body posture estimation method and device, computer readable medium, and electronic device
By combining monocular color cameras and LiDAR sensors, combining color images and sparse depth images, low-cost and high-precision estimation of three-dimensional human postures is achieved, depth ambiguity problems and high-precision requirements are solved, and system cost and power consumption are reduced.
Patent Information
- Application Number
- CN202210213265.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-04
AI Technical Summary
In the three-dimensional human posture estimation, it is difficult to solve the depth ambiguity problem by relying on monocular color images, resulting in large errors in depth direction and difficult to meet high-precision requirements. At the same time, the deep point cloud-based solution requires high quality point clouds, making it difficult to implement projects under low cost and low power consumption.
Using a joint monocular color camera and a single LiDAR sensor, the initial three-dimensional data is estimated and reconstructed by collecting color images and sparse depth images, combining segmented labels to build a three-dimensional point cloud, and fit the reconstructed three-dimensional model into the three-dimensional point cloud to obtain the target three-dimensional data.
It realizes low-cost and high-precision three-dimensional human posture estimation, solves the problem of depth ambiguity, meets the needs of high precision, and reduces the cost and power consumption of the system.
Smart Images

Figure CN114612612B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a human body posture estimation method, a human body posture estimation device, a computer-readable medium, and an electronic device. Background Art
[0002] With the development of 3D vision and deep learning, solutions for estimating 3D human posture through monocular cameras continue to emerge and continue to improve. Using a neural network trained with a large amount of data, the 3D posture of the human body in the RGB pictures taken by an ordinary monocular color camera can be extracted. However, due to the lack of depth information in ordinary color pictures, the estimated human posture has large errors in the depth direction. In some existing technical solutions, 3D human posture estimation is completely dependent on color images combined with parameterized human models. With the help of deep learning, good results have been achieved, but relying solely on color images cannot always solve the depth ambiguity problem well. The estimated 3D human posture has large errors in the depth direction, which is difficult to meet the requirements of some tasks with high posture accuracy requirements. Other solutions based on deep point clouds have high requirements on the quality of point clouds, making it difficult to implement engineering projects under low-cost and low-power conditions.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0004] The present invention provides a human body posture estimation method, a human body posture estimation device, a computer-readable medium and an electronic device, which can realize low-cost and high-precision three-dimensional human body posture estimation by combining a monocular color camera and a single LiDAR sensor.
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0006] According to a first aspect of the present disclosure, a method for estimating a human body posture is provided, comprising:
[0007] Collecting a color image corresponding to the object to be processed and a corresponding sparse depth image;
[0008] estimating initial three-dimensional data corresponding to the color image, and reconstructing based on the initial three-dimensional data to obtain a reconstructed three-dimensional model; and
[0009] Acquire a sparse depth image containing segmentation labels, and construct a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image containing segmentation labels;
[0010] The reconstructed three-dimensional model is fitted to the three-dimensional point cloud to obtain target three-dimensional data.
[0011] According to a second aspect of the present disclosure, there is provided a human body posture estimation device, comprising:
[0012] An image acquisition module, used to acquire a color image corresponding to the object to be processed and a corresponding sparse depth image;
[0013] a three-dimensional model reconstruction module, configured to estimate initial three-dimensional data corresponding to the color image, and reconstruct the three-dimensional model based on the initial three-dimensional data; and
[0014] A three-dimensional point cloud construction module, used to obtain a sparse depth image containing segmentation labels, and construct a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image of the segmentation labels;
[0015] The fitting operation module is used to fit the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data.
[0016] According to a third aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned human body posture estimation method is implemented.
[0017] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0018] Processor; and
[0019] A memory, configured to store executable instructions of the processor;
[0020] Wherein, the processor is configured to implement the above-mentioned human body posture estimation method by executing the executable instructions.
[0021] The human body posture estimation method provided by an embodiment of the present disclosure estimates the initial three-dimensional data using the collected color image, and reconstructs it to obtain a reconstructed three-dimensional model; at the same time, the sparse depth image corresponding to the color image is used to map the corresponding three-dimensional point cloud containing segmentation labels; the reconstructed three-dimensional model is fitted to the three-dimensional point cloud, and finally an accurate three-dimensional human body and corresponding three-dimensional posture data are obtained. Thus, low-cost and high-precision three-dimensional human body posture estimation is achieved using only one monocular color camera and a single LiDAR sensor.
[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0024] Figure 1 A schematic diagram schematically illustrates a method for estimating a human body posture in an exemplary embodiment of the present disclosure;
[0025] Figure 2 A schematic diagram schematically illustrates a method for reconstructing a three-dimensional model of a human body in an exemplary embodiment of the present disclosure;
[0026] Figure 3 A schematic diagram schematically illustrates a method for processing a sparse depth image in an exemplary embodiment of the present disclosure;
[0027] Figure 4 A schematic diagram schematically illustrates a method for fitting a reconstructed three-dimensional model to a three-dimensional point cloud in an exemplary embodiment of the present disclosure;
[0028] Figure 5 A schematic diagram schematically illustrates a flow chart of a method for estimating a human body posture in an exemplary embodiment of the present disclosure;
[0029] Figure 6 A schematic diagram schematically shows the composition of a human body posture estimation device in an exemplary embodiment of the present disclosure;
[0030] Figure 7 The diagram schematically shows the components of an electronic device in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0032] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0033] In related technologies, when performing 3D human posture estimation, most of them rely entirely on color images combined with parameterized human body models for 3D human posture estimation. With the help of deep learning, good results have been achieved, but relying solely on color images cannot always solve the depth ambiguity problem well. The estimated 3D human posture has large errors in the depth direction, which is difficult to meet the requirements of some tasks with high posture accuracy requirements. In addition, some solutions based on deep point clouds have high requirements on the quality of point clouds, which makes it difficult to implement engineering projects under low-cost and low-power conditions.
[0034] In view of the above-mentioned shortcomings and deficiencies of the prior art, this exemplary embodiment provides a method for estimating a human posture. Figure 1 As shown in , the above-mentioned human body posture estimation method may include:
[0035] Step S11, collecting a color image corresponding to the object to be processed and a corresponding sparse depth image;
[0036] Step S12, estimating initial three-dimensional data corresponding to the color image, and reconstructing based on the initial three-dimensional data to obtain a reconstructed three-dimensional model; and
[0037] Step S13, obtaining a sparse depth image containing segmentation labels, and constructing a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image containing segmentation labels;
[0038] Step S14: fitting the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data.
[0039] The human body posture estimation method provided in this example implementation method uses only a monocular color camera and a single LiDAR sensor, uses the collected color image to estimate the initial three-dimensional data, and reconstructs the three-dimensional model; at the same time, the sparse depth image corresponding to the color image is used to map the corresponding three-dimensional point cloud containing segmentation labels; the reconstructed three-dimensional model is fitted to the three-dimensional point cloud, and finally an accurate three-dimensional human body and corresponding three-dimensional posture data are obtained; low-cost and high-precision three-dimensional human body posture estimation is achieved.
[0040] Below, each step of the human body posture estimation method in this example implementation will be described in more detail with reference to the accompanying drawings and embodiments.
[0041] In step S11, a color image corresponding to the object to be processed and a corresponding sparse depth image are acquired.
[0042] The above method in this example implementation can be applied to smart terminal devices such as mobile phones and tablet computers. The terminal device can be equipped with a camera and a LiDAR (Light detection and ranging) sensor. Among them, the camera is a monocular color camera. When collecting raw image data, the color image can be collected using only a monocular color camera of the terminal device; at the same time, the sparse depth image can be collected using only a LiDAR sensor of the terminal device.
[0043] In step S12, initial three-dimensional data corresponding to the color image is estimated, and reconstruction is performed based on the initial three-dimensional data to obtain a reconstructed three-dimensional model.
[0044] In this example implementation, reference Figure 2 As shown, the above step S12 may include:
[0045] Step S121, using the SMPL parameter regression network to process the color image to obtain the initial three-dimensional data;
[0046] Step S122: input the initial three-dimensional data into the SMPL model for reconstruction to obtain a reconstructed three-dimensional model including a preset number of vertices.
[0047] Specifically, for the collected RGB image, a top-down SMPL (A Skinned Multi-Person Linear Model) parameter regression network can be used for processing. The RGB image is input into the SMPL parameter regression network, and the initial three-dimensional data corresponding to the RGB graphic is output. Among them, the above-mentioned initial three-dimensional data may include the posture parameter θ and the body parameter β predicted by the network and the global offset t. The above-mentioned SMPL parameter regression network can adopt a pre-trained HRM (Human Mesh Recovery), SPIN (SMP LoPtimization IN the loop) or GCMR (Graph Convolutional Mesh Regression), and other models.
[0048] After obtaining the predicted posture parameter θ and body posture parameter β and the global offset t, they can be used as input parameters to reconstruct the human body 3D model using the differentiable SMPL model to reconstruct a 3D human body model containing a target number of vertices. For example, the reconstructed 3D model can be a 3D human body model consisting of 6890 vertices.
[0049] In some exemplary embodiments, a SMPL parameter regression network can also be pre-trained. Specifically, a certain number of RGB images and corresponding LiDAR depth maps can be collected and aligned to construct training samples; the aligned LiDAR and RGB images are spliced in the channel dimension. The number of input channels of the HMR network structure is changed from 3 channels to 4 channels, and the network structure is used for training. Synthetic data is generated by rendering a 3D scan model for use as training data; the trained SMPL parameter regression network is used to estimate the currently collected color image to obtain the corresponding initial three-dimensional data.
[0050] In step S13, a sparse depth image including segmentation labels is acquired, and a corresponding three-dimensional point cloud including segmentation labels is constructed based on the sparse depth image including segmentation labels.
[0051] In this example implementation, reference Figure 3 As shown, the above-mentioned acquisition of the sparse depth image containing the segmentation label may include:
[0052] Step S131, aligning the sparse depth image with the color image;
[0053] Step S132: Based on the alignment result of the sparse depth image and the color image, synchronously segment the sparse depth image and the color image to obtain a color image segmentation result including a segmentation label and a sparse depth image including a segmentation label.
[0054] Specifically, when reconstructing the three-dimensional model of the human body, the collected color image and sparse depth image can be synchronously aligned; for example, the sparse depth map can be aligned with the RGB image through the camera parameters of the RGB camera and the camera parameters of the LiDAR sensor. After the image is aligned, the body segmentation network can be used to segment the RGB image, and each pixel obtains the corresponding human part label. For example, the Self-Correction network pre-trained on the Pascal-Person-Part dataset can be used to segment the color image into seven categories: head, body, upper arm, lower arm, thigh, calf, and background. Since the color image and the sparse depth image are aligned in advance, the sparse depth image can be synchronously segmented at this time to obtain a sparse depth image containing segmentation labels.
[0055] In this example implementation, after obtaining the sparse depth map containing the segmentation labels, the sparse depth map can be projected into a corresponding three-dimensional point cloud. Specifically, the projection matrix contained in the intrinsic parameters of the LIDAR sensor can be used in combination with the camera parameters to convert the depth value of each pixel in the sparse depth map into a corresponding three-dimensional point, thereby completing the construction of the three-dimensional point cloud.
[0056] In step S14, the reconstructed three-dimensional model is fitted to the three-dimensional point cloud to obtain target three-dimensional data.
[0057] In this example implementation, reference Figure 4 As shown, the above step S14 may include:
[0058] Step S141, roughly aligning the reconstructed three-dimensional model with the three-dimensional point cloud;
[0059] Step S142, using a gradient descent method to optimize the initial three-dimensional parameters, and obtaining an optimized SMPL model at the end of the iteration;
[0060] Step S143: output the target three-dimensional data using the optimized SMPL model.
[0061] Specifically, the reconstructed 3D human body model obtained in the above steps can be roughly aligned with the LiDAR point cloud. Then, the posture parameter θ, body posture parameter β and global offset t are iteratively optimized using the gradient descent method. The Adam optimizer can be used in the optimization process. After the iteration, the optimized SMPL model is obtained and the corresponding 3D joint point coordinates are output.
[0062] In the iterative optimization process, the optimization process uses the Adam optimizer to minimize the loss function. The loss function includes: data item loss and segmentation category semantic loss based on segmentation labels. Specifically, the formula of the loss function can include:
[0063] E(θ,β,t)=E data +w part E part
[0064]
[0065]
[0066] Among them, E data is a data item that can be used for the nearest neighbor distance between the data item SMPL model vertex (S1) and the LiDAR point cloud (S2). dataAs shown in the formula, the nearest neighbor point can be searched in both directions; that is, from S1 to S2, and from S2 to S1, corresponding to the two terms before and after the plus sign in the formula. part are semantic items related to the categories of body parts. x and y represent the 3D points in the SMPL vertex set S1 and the LiDAR point cloud S2, respectively. I Refers to the set of three-dimensional points in the point set S that belong to category I. w is a hyperparameter; for example, the configuration w = 0.3 and so on.
[0067] In some exemplary embodiments, the image type of the color image can also be identified, and the weight w of the data item loss and the semantic item loss can be configured according to the image type. For example, the above-mentioned image types can be divided according to the lighting type, the brightness of the background, the background type, and the difficulty level of background segmentation. Due to different background brightness, different background content or lighting type, different difficulty levels and calculation amounts may be caused in the RGB image segmentation process for the human body, so the weights of the data item loss and the semantic item loss can be configured based on this. For example, in an image with natural light as the lighting type and a simple background level, the difficulty of segmenting the human body is relatively low. At this time, a weight of 0.6 can be assigned to the semantic item loss, and so on. Of course, the above-mentioned numerical values are only exemplary descriptions, and the corresponding w can be configured according to the actual situation of the image.
[0068] Alternatively, in some exemplary embodiments, during the iterative optimization process, a joint point projection loss based on two-dimensional joint point coordinates may be added to the above-mentioned loss function.
[0069] Specifically, after collecting the RGB image, the 2D joint point coordinates in the image can be estimated; for example, Openpose, Alphapose and other models can be used to estimate the 2D joint point coordinates on the RGB image. Specifically, the projection loss of the joint point can include:
[0070]
[0071] Among them, x represents the 2D key point estimated on the RGB image, and x^ represents the projection coordinates of the 3D joint point on the SMPL model projected onto the image.
[0072] By adding joint projection loss to the loss function, the feature information of human joints is increased, which can further improve the accuracy of posture recognition.
[0073] In addition, regularization terms can be added to the iterative optimization loss function to enhance the stability of the optimization and improve the overall robustness of the algorithm.
[0074] The human body posture estimation method provided in the embodiment of the present disclosure refers to Figure 5As shown, the collected RGB image 501 and LiDAR sparse depth image 502 can be used as the input of the model, and the LiDAR sparse depth image can be aligned with the RGB image through the camera parameters of the RGB camera and the LiDAR sensor. On the one hand, the RGB image is input into the SMPL parameter regression network 503 to obtain the initial three-dimensional image; on the other hand, the aligned LiDAR sparse depth image and the RGB image are input into the body segmentation network 504, so that each pixel of the RGB image obtains the corresponding human body part label, and the body segmentation of the LiDAR depth image is completed synchronously. The LiDAR depth image with the body part label is projected into a three-dimensional LiDAR point cloud 505 using the camera parameters. At the same time, the predicted initial three-dimensional parameters are input into the differentiable SMPL model to reconstruct a three-dimensional human body model 506 consisting of 6890 vertices, and the reconstructed three-dimensional human body is roughly aligned with the LiDAR point cloud. Then the gradient descent method is used to iteratively optimize the θ, β and t parameters, and the Adam optimizer is used in the optimization process; the optimized SMPL model is obtained at the end of the iteration, and the corresponding three-dimensional joint point coordinates are output.
[0075] By using RGB images and LiDAR sparse depth images in the calculation process, the RGB data and LiDAR data are combined to balance low cost and high precision. The accuracy of the predicted three-dimensional human posture is much higher than that of the algorithm based on RGB images alone; the depth ambiguity problem of three-dimensional human posture is solved to a large extent. It can meet engineering requirements with higher precision. By semantically segmenting the RGB image, the semantic information is associated with the LiDAR point cloud, making the optimization process more reliable and more accurate. In addition, due to the use of LiDAR sensors, compared with the use of high-quality depth cameras such as Kinect, it has the advantages of low power consumption and low cost, which is more conducive to deployment on mobile devices such as mobile phones and AR glasses. At the same time, due to the combination of RGB information for reasoning, the reasoning accuracy can be made no lower than that of the solution using high-quality depth cameras. In addition, compared with starting the optimization from the random or all-zero initialized SMPL parameters, this method uses the SMPL parameters predicted by RGB data as the initial value, which accelerates the optimization process and improves the optimization accuracy.
[0076] It should be noted that the above figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0077] For further reference, Figure 6As shown, in the embodiment of this example, a human posture estimation device 60 is also provided, and the device includes: an image acquisition module 601, a three-dimensional model reconstruction module 602, a three-dimensional point cloud construction module 603 and a fitting operation module 604.
[0078] The image acquisition module 601 may be used to acquire a color image corresponding to the object to be processed and a corresponding sparse depth image.
[0079] The three-dimensional model reconstruction module 602 is used to estimate the initial three-dimensional data corresponding to the color image, and reconstruct based on the initial three-dimensional data to obtain a reconstructed three-dimensional model; and
[0080] The three-dimensional point cloud construction module 603 can be used to obtain a sparse depth image containing segmentation labels, and construct a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image containing segmentation labels.
[0081] The fitting operation module 604 may be used to fit the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data.
[0082] In some exemplary embodiments, the image acquisition module 601 may include: acquiring the color image using a monocular color camera of the terminal device; and acquiring the sparse depth image using a LiDAR sensor of the terminal device.
[0083] In some exemplary embodiments, the initial three-dimensional data includes: posture parameter θ, body posture parameter β, and global offset t.
[0084] The three-dimensional model reconstruction module 602 includes: using an SMPL parameter regression network to process the color image to obtain the initial three-dimensional data; inputting the initial three-dimensional data into an SMPL model for reconstruction to obtain a reconstructed three-dimensional model containing a preset number of vertices.
[0085] In some exemplary embodiments, the three-dimensional point cloud construction module 603 may include aligning the sparse depth image with the color image; based on the alignment result of the sparse depth image and the color image, synchronously performing image segmentation on the sparse depth image and the color image to obtain a color image segmentation result containing segmentation labels and a sparse depth image containing segmentation labels.
[0086] In some exemplary embodiments, the three-dimensional point cloud construction module 603 may include: performing projection calculation on the depth image using a preset projection matrix to obtain the three-dimensional point cloud.
[0087] In some exemplary embodiments, the fitting operation module 604 may include: roughly aligning the reconstructed three-dimensional model with the three-dimensional point cloud; optimizing the initial three-dimensional parameters using a gradient descent method, and obtaining an optimized SMPL model at the end of the iteration; and outputting the target three-dimensional data using the optimized SMPL model.
[0088] In some exemplary embodiments, when the initial three-dimensional parameters are optimized using the gradient descent method, the loss function includes: data item loss, and segmentation category semantic loss based on segmentation labels.
[0089] In some exemplary embodiments, the apparatus further comprises: a weight configuration module.
[0090] The weight configuration module can be used to identify the image type of the color image and configure the weights of each part of the loss function according to the image type.
[0091] In some exemplary embodiments, the apparatus further comprises: a joint point processing module.
[0092] The joint point processing module can be used to obtain the two-dimensional joint point coordinates corresponding to the color image.
[0093] The loss function also includes: a joint point projection loss based on the two-dimensional joint point coordinates.
[0094] The specific details of each module in the above-mentioned human body posture estimation device 60 have been described in detail in the corresponding human body posture estimation method, so they will not be repeated here.
[0095] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0096] Figure 7 A schematic diagram of an electronic device suitable for implementing an embodiment of the present invention is shown.
[0097] It should be noted that Figure 7 The electronic device 1000 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0098] like Figure 7As shown, the electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 to the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0099] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that a computer program read therefrom is installed into the storage section 1008 as needed.
[0100] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 1009, and / or installed from a removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present application are executed.
[0101] Specifically, the electronic device may be a smart mobile terminal device such as a mobile phone, a tablet computer or a laptop computer, or may be a smart terminal device such as a desktop computer.
[0102] It should be noted that the computer-readable medium shown in the embodiment of the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0103] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0104] The units involved in the embodiments of the present invention may be implemented by software or hardware, and the units described may also be arranged in a processor. The names of these units do not, in some cases, limit the units themselves.
[0105] It should be noted that, as another aspect, the present application also provides a computer-readable medium, which may be included in an electronic device; or may exist independently without being installed in the electronic device. The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments. For example, the electronic device may implement the following Figure 1 The steps shown.
[0106] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0107] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
[0108] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for estimating a human body posture, characterized in that: The method comprises: Collecting a color image corresponding to the object to be processed and a corresponding sparse depth image; estimating initial three-dimensional data corresponding to the color image, and reconstructing based on the initial three-dimensional data to obtain a reconstructed three-dimensional model; and Acquire a sparse depth image containing segmentation labels, and construct a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image containing segmentation labels; the constructing of the corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image containing segmentation labels includes: using a preset projection matrix to perform projection calculation on the depth image to obtain the three-dimensional point cloud; Fitting the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data; the target three-dimensional data includes an accurate three-dimensional human body and corresponding three-dimensional posture data; The step of fitting the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data includes: Roughly aligning the reconstructed three-dimensional model with the three-dimensional point cloud; Optimizing the initial three-dimensional data using a gradient descent method, and obtaining an optimized SMPL model at the end of the iteration; The target three-dimensional data is outputted using the optimized SMPL model.
2. The method for estimating human body posture according to claim 1, characterized in that: The collecting of the color image corresponding to the object to be processed and the corresponding sparse depth image includes: Using a monocular color camera of a terminal device to collect the color image; and The sparse depth image is collected by using a LiDAR sensor of the terminal device.
3. The method for estimating human body posture according to claim 1, characterized in that: The initial three-dimensional data includes: posture parameter θ, body posture parameter β, and global offset t; The estimating the initial three-dimensional data corresponding to the color image and reconstructing based on the initial three-dimensional data to obtain a reconstructed three-dimensional model includes: Processing the color image using an SMPL parameter regression network to obtain the initial three-dimensional data; The initial three-dimensional data is input into the SMPL model for reconstruction to obtain a reconstructed three-dimensional model including a preset number of vertices.
4. The method for estimating human body posture according to claim 1, characterized in that: The step of obtaining a sparse depth image containing segmentation labels includes: Aligning the sparse depth image with the color image; Based on the alignment result of the sparse depth image and the color image, synchronously perform image segmentation on the sparse depth image and the color image to obtain a color image segmentation result including a segmentation label and a sparse depth image including a segmentation label.
5. The method for estimating human body posture according to claim 1, characterized in that: When the gradient descent method is used to optimize the initial three-dimensional data, the loss function includes: data item loss and segmentation category semantic loss based on segmentation labels.
6. The method according to claim 5, characterized in that The method also includes: identifying the image type of the color image, and configuring the weights of the loss functions of each part according to the image type.
7. The method according to claim 5, characterized in that The method further comprises: obtaining the two-dimensional joint point coordinates corresponding to the color image; The loss function also includes: a joint point projection loss based on the two-dimensional joint point coordinates.
8. A human body posture estimation device, characterized in that: The device comprises: An image acquisition module, used to acquire a color image corresponding to the object to be processed and a corresponding sparse depth image; a three-dimensional model reconstruction module, configured to estimate initial three-dimensional data corresponding to the color image, and reconstruct the three-dimensional model based on the initial three-dimensional data; and A three-dimensional point cloud construction module, used to obtain a sparse depth image containing segmentation labels, and construct a corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image of the segmentation labels; the construction of the corresponding three-dimensional point cloud containing segmentation labels based on the sparse depth image of the segmentation labels is configured to: use a preset projection matrix to perform projection calculation on the depth image to obtain the three-dimensional point cloud; A fitting operation module, used for fitting the reconstructed three-dimensional model to the three-dimensional point cloud to obtain target three-dimensional data; the target three-dimensional data includes an accurate three-dimensional human body and corresponding three-dimensional posture data; The fitting operation module is configured as follows: Roughly aligning the reconstructed three-dimensional model with the three-dimensional point cloud; Optimizing the initial three-dimensional data using a gradient descent method, and obtaining an optimized SMPL model at the end of the iteration; The target three-dimensional data is outputted using the optimized SMPL model.
9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the human body posture estimation method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the human body posture estimation method according to any one of claims 1 to 7 by executing the executable instructions.
Citation Information
Patent Citations
Human body three-dimensional model acquisition method and device, intelligent terminal and storage medium
CN113610889A