A 3D human body shape generation method and system based on multi-viewpoint matching
The method improves 3D human body shape prediction by projecting multiple-view single-image data onto a common plane and optimizing the model with a loss function, addressing precision issues in existing methods and simplifying the reconstruction process.
Patent Information
- Application Number
- CN202111369810.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-18
AI Technical Summary
In the existing three-dimensional human body reconstruction technology, monocular image information is incomplete and has low accuracy, multi-image reconstruction costs are high and rely on external device information, resulting in rough reconstruction results and inaccurate use of two-dimensional information.
The multi-view angle matching method is used to input the 3D human body morphology generation model through monocular images of multiple perspective angles, and the model is optimized using the backpropagation algorithm and stochastic gradient descent method to construct the view angle difference loss function to improve prediction accuracy.
More accurate 3D human body morphology prediction is achieved, simplifying model complexity, reducing costs and improving prediction accuracy.
Smart Images

Figure CN114067053B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a 3D human body shape generation method and system based on multi-view matching. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] In computer vision, current three-dimensional human body reconstruction is a process of reconstructing three-dimensional information based on a single monocular or multi-view image. The single-view monocular image method is simple. However, due to incomplete monocular information and low accuracy, prior knowledge is required for three-dimensional reconstruction; while the cost of collecting data for multi-view three-dimensional reconstruction is high and the model is complex.
[0004] In addition, most current three-dimensional reconstruction algorithms do not utilize two-dimensional information accurately and comprehensively enough, and the calculation process overly relies on information provided by external devices, such as depth information provided by a depth camera, or relies on the segmentation results of the target and the background, etc., resulting in a still relatively rough reconstructed result. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background art, the present invention provides a 3D human body shape generation method and system based on multi-view matching, which uses monocular information from multiple views to predict the 3D shape of the human body, and has the characteristics of high prediction accuracy and generality.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a 3D human body shape generation method based on multi-view matching, which includes:
[0008] Obtain monocular images from multiple views;
[0009] Input the obtained monocular images from multiple views into a 3D human body shape generation model to obtain a 3D human body shape;
[0010] Wherein, the 3D human body shape generation model projects the vertices of the predicted 3D human body shape of each view onto an arbitrary three-dimensional plane, and at the same time projects the 3D human body shape ground truth onto the three-dimensional plane to obtain the 3D human body ground truth shape and the predicted 3D human body shape of each view under the three-dimensional plane; and, the 3D human body shape generation model is optimized by a loss function composed of the difference between the 3D human body ground truth shape and the predicted 3D human body shape of each view under the three-dimensional plane.
[0011] Further, the 3D human body shape generation model uses the backpropagation algorithm and the stochastic gradient descent method to update the weight and bias parameters to minimize the loss function.
[0012] Further, the backpropagation algorithm uses the first-order partial derivative to calculate the gradient of the loss function with respect to the weight and bias parameters in the 3D human body shape generation model.
[0013] Further, the stochastic gradient descent method updates the weight and bias parameters by multiplying the learning rate by the gradient of the loss function.
[0014] Further, the 3D human body shape generation model predicts the 3D human body prediction shape for each view based on the monocular images from multiple views using the HMR (Human Mesh Recovery) prediction network.
[0015] Further, the 3D human body shape generation model models the 3D human body shape using pose parameters and shape parameters.
[0016] The second aspect of the present invention provides a 3D human body shape generation system based on multi-view matching, which includes:
[0017] A data acquisition module, which is configured to: acquire monocular images from multiple views;
[0018] A 3D human body shape generation module, which is configured to: input the acquired monocular images from multiple views into the 3D human body shape generation model to obtain the 3D human body shape;
[0019] Wherein, the 3D human body shape generation model projects the vertices of the 3D human body prediction shape for each view onto an arbitrary three-dimensional plane, and at the same time projects the 3D human body shape ground truth onto the three-dimensional plane to obtain the 3D human body ground truth shape and the 3D human body prediction shape for each view in the three-dimensional plane; and, the 3D human body shape generation model is optimized by a loss function formed by the difference between the 3D human body ground truth shape and the 3D human body prediction shape for each view in the three-dimensional plane.
[0020] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a 3D human body shape generation method based on multi-view matching as described above.
[0021] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a 3D human body shape generation method based on multi-view matching as described above.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] The present invention provides a 3D human body shape generation method based on multi-view matching, which predicts the 3D human body shape using monocular images from multiple views, and is more accurate than the existing method of 3D human body reconstruction using a single monocular image, and is simpler and lighter than the existing method of 3D human body reconstruction using multi-view images.
[0024] The present invention provides a 3D human body shape generation method based on multi-view matching, which optimizes the generation network by constructing the loss between the predicted 3D human body shape and the ground truth shape of each view, and improves the accuracy of 3D human body shape prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0026] Figure 1 is a flowchart of a 3D human body shape generation method based on multi-view matching according to Embodiment 1 of the present invention;
[0027] Figure 2 is a structural diagram of a 3D human body shape generation model according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0031] Embodiment 1
[0032] As Figure 1As shown in the figure, this embodiment provides a 3D human body shape generation method based on multi-view matching, which optimizes the generation model by registering the human body shapes generated from multi-view monocular images. First, for the input multi-view monocular images, use the HMR (Human Mesh Recovery) prediction model to process them to obtain the 3D human body shapes of multiple views. Among them, the 3D human body shape model used in the present invention is SMPL (Skinned Multi-Person Linear Model), and this model models the 3D human body shape through pose parameters and shape parameters. Then, project the vertices of the human body shapes obtained from each view onto a three-dimensional plane to obtain the 3D human body shape projections on the same plane. Finally, further optimize the generation network by constructing the loss between the above projections and the ground truth shape projections, and finally obtain a more accurate 3D human body shape generation model. Finally, encapsulate the algorithm for actual testing. This method has the characteristics of generality and more accurate matching. On the one hand, this method combines multi-view monocular input images, providing data support for the model to more accurately model the 3D human body shape at each angle. On the other hand, this method proposes to optimize the generation network by constructing the loss between the predicted 3D human body shapes and the ground truth shapes of each view. Specifically, it includes the following steps:
[0033] The first step: Obtain the monocular images taken from N views, denoted as where θ represents the angle between the shooting view and the front angle during shooting. The specified front angle is the state when the human body is parallel to the camera during shooting. In this state, θ = 0°. By rotating the human body, monocular images under multiple views can be obtained, and the rotation angle is denoted as θ. By combining multi-view monocular human input images, it provides data support for the model to more accurately model the 3D human body shape at each angle.
[0034] As an implementation method, input the monocular images taken from 3 views, denoted as where θ1 = 0°, θ2 = 45°, θ3 = 135°.
[0035] The second step: Use the trained 3D human body shape generation model to input the monocular images obtained from N views into the trained 3D human body shape generation model to obtain the 3D human body shapes. Project the vertices of the predicted 3D human body shapes of each view onto an arbitrary three-dimensional plane, and at the same time project the 3D human body shape ground truth onto the three-dimensional plane to obtain the 3D human body ground truth shape and the predicted 3D human body shapes of each view on the three-dimensional plane, and optimize them through the loss function composed of the difference between the 3D human body ground truth shape and the predicted 3D human body shapes of each view on the three-dimensional plane. Specifically, the training process of the 3D human body shape generation model includes:
[0036] Step S1: Input monocular images taken from N perspectives. Based on the monocular images from multiple perspectives, use the HMR prediction network to obtain the 3D human prediction morphology for each perspective.
[0037] Specifically, use the HMR (Human Mesh Recovery) prediction network. The specific structure of the network is as Figure 2 shown. This network is an end-to-end network for reconstructing the complete 3D morphology of the human body from a single image. Preprocess it to obtain the 3D human prediction morphology for N perspectives. Among them, the 3D human morphology used in the present invention is SMPL. This method can obtain the 3D human body morphology through the pose parameters and shape parameters . The 3D human body morphology consists of M vertices. The 3D human prediction morphology for N perspectives is denoted as Among them, represents the vertex coordinate information of the 3D human prediction morphology obtained under the perspective θ.
[0038] As an implementation, use the HMR prediction network to obtain the 3D human prediction morphology for 3 perspectives. In this network, the SMPL human body modeling method is used to obtain the 3D human body morphology through the pose parameters and shape parameters . The 3D human body morphology contains 6890 vertices. The 3D human prediction morphology for 3 perspectives is denoted as Among them, represents the vertex coordinate information of the 3D human body morphology obtained under the perspective θ.
[0039] Among them, the HMR prediction network and SMPL are commonly used methods in the field of 3D human body morphology estimation. Figure 2 The model registration mentioned in
[0040] refers to the content of the following steps S2 - S3, and it is called model registration.
[0041] Step S2: Project the vertex coordinates of the 3D human prediction morphology for each perspective onto an arbitrary three-dimensional plane, and at the same time project the true value coordinates of the 3D human body morphology onto the same three-dimensional plane to obtain the 3D human body true value morphology and the 3D human prediction morphology for each perspective in the unified three-dimensional plane. Among them, Projecting the vertex coordinates of each obtained 3D human prediction morphology onto an arbitrary three-dimensional plane can obtain N projections, and each projection contains M coordinates, denoted asp , y p , z p ):
[0042]
[0043]
[0044]
[0045] Similarly, project the given true coordinates of the 3D human body shape onto the same coordinate plane according to the above steps, and denote the projected coordinate information as where GT i represents the projected coordinate information of the i-th vertex in the true shape, and the 3D human body true coordinate projection also contains M projected coordinates.
[0046] As an implementation, project the vertex coordinates of each obtained 3D human body predicted shape onto an arbitrary three-dimensional plane, and 3 projections can be obtained, and each projection contains 6,890 coordinates, denoted as where represents the coordinate information after the vertex projection of the 3D human body predicted shape obtained under the viewing angle θ. The projection formula is as follows. Assume that the general equation of the plane is x + 2y + 3z + 4 = 0. Assume that the vertex coordinates of the above 3D human body predicted shape are (x0, y0, z0), and its projected point coordinates on the plane are (x p , y p , z p ):
[0047]
[0048]
[0049]
[0050] Similarly, project the given true coordinates of the 3D human body shape onto the same coordinate plane according to the above steps, and denote the projected coordinate information as
[0051] Step S3: The 3D human body shape generation model is optimized by a loss function composed of the difference between the 3D human body true shape and the 3D human body predicted shape at each viewing angle under a unified three-dimensional plane.
[0052] Specifically, construct an additional loss function L, and the formula is as follows. It is composed of the difference between the vertex projected coordinates and the true projected coordinates at each viewing angle, and this difference is represented by the two-norm. Among them, further optimization is achieved by constructing the loss between the 3D shape projection and the true shape projection of the human body at each viewing angle Figure 2The 3D human body shape generation model shown, and then based on this error, the weights of the 3D human body shape generation model are optimized and corrected to obtain a more accurate 3D human body shape generation model;
[0053] L = L1 + L2 + … + L N
[0054]
[0055]
[0056]
[0057] As an implementation, the loss function L is composed of the difference between the vertex projection coordinates and the true value projection coordinates from the above three perspectives. This difference is represented by the two-norm, and the formula is as follows:
[0058] L = L1 + L2 + L3
[0059]
[0060]
[0061]
[0062] Step S4: Use the backpropagation algorithm and the stochastic gradient descent method to reduce the error. After multiple iterative trainings of the 3D human body shape generation model, a trained 3D human body shape generation model is obtained. Among them, backpropagation is a common method used to train artificial neural networks in combination with optimization methods (such as the stochastic gradient descent method). This method calculates the gradient of the loss function for all weight and bias parameters in the 3D human body shape generation model. This gradient is fed back to the optimization method to update the weights and bias parameters to minimize the loss function. Backpropagation uses the first-order partial derivatives to calculate the gradient of the loss function for the weight and bias parameters in the 3D human body shape generation model, and the formula is as follows:
[0063]
[0064]
[0065] The above formulas respectively represent calculating the gradient of the loss function L for the weight W (l) and the bias parameter b (l) in a certain layer l of the model, and the calculation is carried out using the chain rule, where o represents the output layer and k represents the last layer of the hidden layer. The stochastic gradient descent method updates the weights and bias parameters through the product of the learning rate and the gradient of the loss function. Specifically, the updated parameter values are calculated through the following formula, where μ represents the learning rate, which can be understood as the rate of change of the model parameters:
[0066]
[0067]
[0068] In this experiment, the method of the present invention was not adopted. Instead, a single image was directly input into the benchmark model for human 3D shape prediction, and the prediction accuracy was 45%. Using the method of the present invention, monocular images taken from three perspectives, namely θ1 = 0°, θ2 = 45°, and θ3 = 135°, were selected as the input of the model, and the prediction model shown was trained using the introduced projection formula and the L2 norm loss in the present invention, resulting in a prediction accuracy of 70%. Through the proposed method, by using multi-view monocular data and introducing the loss of the difference between the projected data and the true shape projection, the learning data of the model was expanded, providing more information from different perspectives. Therefore, the method of the present invention achieved a relatively high accuracy in the existing human 3D shape prediction task. Figure 2 As described above, through the proposed method, by using multi-view monocular data and introducing the loss of the difference between the projected data and the true shape projection, the learning data of the model was expanded, providing more information from different perspectives. Therefore, the method of the present invention achieved a relatively high accuracy in the existing human 3D shape prediction task.
[0069] Example 2
[0070] This example provides a 3D human shape generation system based on multi-view matching, which specifically includes the following modules:
[0071] A data acquisition module, which is configured to: acquire monocular images from multiple perspectives;
[0072] A 3D human shape generation module, which is configured to: input the acquired monocular images from multiple perspectives into a 3D human shape generation model to obtain a 3D human shape;
[0073] Among them, the 3D human shape generation model projects the vertices of the predicted 3D human shape of each perspective onto an arbitrary three-dimensional plane, and at the same time projects the true 3D human shape onto the three-dimensional plane to obtain the true 3D human shape under the three-dimensional plane and the predicted 3D human shape of each perspective; and the 3D human shape generation model is optimized by a loss function formed by the difference between the true 3D human shape under the three-dimensional plane and the predicted 3D human shape of each perspective.
[0074] The 3D human shape generation model uses the backpropagation algorithm and the stochastic gradient descent method to update the weight and bias parameters to minimize the loss function.
[0075] It should be noted here that each module in this example corresponds to each step in Example 1 one by one, and the specific implementation process is the same, so it will not be repeated here.
[0076] Example 3
[0077] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a 3D human body shape generation method based on multi-viewpoint matching as described in Embodiment 1 above.
[0078] Embodiment 4
[0079] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a 3D human body shape generation method based on multi-viewpoint matching as described in Embodiment 1 above.
[0080] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an embodiment implemented in hardware, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.
[0081] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in one Figure 1One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.
[0084] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0085] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A 3D human body shape generation method based on multi-view matching, characterized in that including: Obtain monocular images from multiple perspectives; Input the obtained monocular images from multiple perspectives into a 3D human body shape generation model to obtain a 3D human body shape; Among them, the 3D human body shape generation model projects the vertices of the predicted 3D human body shape of each perspective onto an arbitrary three-dimensional plane, and at the same time projects the 3D human body shape ground truth onto the three-dimensional plane, to obtain the 3D human body ground truth shape and the predicted 3D human body shape of each perspective under the three-dimensional plane, specifically: Project the vertex coordinates of each obtained 3D human prediction form onto an arbitrary three-dimensional plane to obtain N projections, and each projection contains M coordinates, denoted as , where represents the coordinate information of the projection of vertex i of the 3D human prediction form obtained under the viewing angle . The projection formula is as follows. Assume that the general equation of the plane is . Assume that the vertex coordinates of the above 3D human prediction form are , and the coordinates of its projection point on the plane are : Similarly, project the given true 3D body shape coordinates onto the same three-dimensional coordinate plane according to the above steps, and denote the projected coordinate information as , where represents the projected coordinate information of the i-th vertex in the true shape, and the 3D body shape true coordinate projection also contains M projected coordinates; Moreover, the 3D human body shape generation model is optimized by a loss function composed of the difference between the 3D human body ground truth shape and the predicted 3D human body shape of each perspective under the three-dimensional plane.
2. The 3D human body shape generation method based on multi-viewpoint matching according to claim 1, wherein, The 3D human body shape generation model uses the backpropagation algorithm and the stochastic gradient descent method to update the weight and bias parameters to minimize the loss function.
3. The 3D human body shape generation method based on multi-viewpoint matching according to claim 2, characterized in that The backpropagation algorithm uses the first-order partial derivative to calculate the gradient of the loss function with respect to the weight and bias parameters in the 3D human body shape generation model.
4. The 3D human body shape generation method based on multi-viewpoint matching according to claim 3, wherein The stochastic gradient descent method updates the weight and bias parameters by multiplying the learning rate by the gradient of the loss function.
5. A 3D human body shape generation method based on multi-viewpoint matching according to claim 1, characterized in that, The 3D human body shape generation model predicts the predicted 3D human body shape of each perspective based on the monocular images from multiple perspectives using the HMR prediction network.
6. The 3D human body shape generation method based on multi-viewpoint matching according to claim 1, characterized in that The 3D human body shape generation model models the 3D human body shape through pose parameters and shape parameters.
7. A 3D human body shape generation system based on multi-viewpoint matching, characterized in that including: A data acquisition module, which is configured to: obtain monocular images from multiple perspectives; A 3D human body shape generation module, which is configured to: input the obtained monocular images from multiple perspectives into a 3D human body shape generation model to obtain a 3D human body shape; Among them, the 3D human body shape generation model projects the vertices of the predicted 3D human body shape of each perspective onto an arbitrary three-dimensional plane, and at the same time projects the 3D human body shape ground truth onto the three-dimensional plane, to obtain the 3D human body ground truth shape and the predicted 3D human body shape of each perspective under the three-dimensional plane, specifically: Project the vertex coordinates of each obtained 3D human prediction form onto an arbitrary three-dimensional plane to obtain N projections, and each projection contains M coordinates, denoted as , where represents the projected coordinate information of vertex i of the 3D human prediction form obtained under the viewing angle . The projection formula is as follows. Assume the general equation of the plane is . Assume the vertex coordinates of the above 3D human prediction form are , and the projected point coordinates on the plane are : Similarly, project the given true 3D body shape coordinates onto the same three-dimensional coordinate plane according to the above steps, and denote the projected coordinate information as , where represents the projected coordinate information of the i-th vertex in the true shape, and the true 3D body shape coordinate projection also contains M projected coordinates; Moreover, the 3D human body shape generation model is optimized by a loss function composed of the difference between the 3D human body ground truth shape and the predicted 3D human body shape of each perspective under the three-dimensional plane.
8. A 3D human body shape generation system based on multi-viewpoint matching according to claim 7, characterized in that The 3D human body shape generation model uses the backpropagation algorithm and the stochastic gradient descent method to update the weight and bias parameters to minimize the loss function.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a 3D human body shape generation method based on multi-perspective matching as described in any one of claims 1-6.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a 3D human body shape generation method based on multi-perspective matching as described in any one of claims 1-6.
Citation Information
Patent Citations
Method and system for identifying three-dimensional position of object
CN110910449A
Three-dimensional human body reconstruction method based on graph convolution
CN111627101A