Real-time three-dimensional reconstruction method and system based on SPIN model

CN115496862BActive Publication Date: 2026-10-09FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211300822.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2026-10-09
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

但无论是基于优化的方法还是基于回归的方法,都专注于人体姿态的准确性,而忽略了人体体型的丰富度,因此重建的人体模型大多以平均体型呈现

Benefits of technology

[0049] The SPIN-MAX system of this invention, by modifying the original SPIN model, achieves more accurate 3D human body reconstruction while maintaining excellent accuracy and speed, modeling the rich differences in body shape between individuals. Simultaneously, SPIN-MAX is more suitable for real-world applications, processing single images input from remote camera modules into SMPL human model parameters representing human posture and body shape, and outputting them to remote client applications. The application of this technology not only improves the accuracy of human body shape reconstruction in 3D human body reconstruction but also introduces new vitality into virtual reality applications, bridging algorithms and industry, and encouraging more academic innovations to be implemented in industry, thus having broad and far-reaching significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496862B_ABST
    Figure CN115496862B_ABST
Patent Text Reader

Abstract

The application discloses a real-time three-dimensional reconstruction method and system based on a SPIN model; the system comprises the following modules: an image receiving module, which receives an RGB picture from a remote camera module; an image preprocessing module, which performs normalization and standardization processing on the RGB picture; a neural network module, which generates SMPL model parameters by using the preprocessed picture information; a result post-processing module, which rewrites and encapsulates the model parameters so that the model parameters are suitable for network transmission; and a result transmission module, which transmits the post-processed human body model parameters to a VR client. By introducing a stacked hourglass model and using a reprojection loss and a human body dressing differentiation loss, the application effectively improves the reconstruction effect of different human body shapes. Meanwhile, the application performs lightweight modification on the model, and through careful design of a data format, the model meets real-time requirements in a network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically, it relates to a real-time 3D reconstruction method based on the SPIN model. Background Technology

[0002] 3D human reconstruction algorithms have matured, resulting in numerous excellent research achievements, among which the SPIN (SMPLoPtimization in the Loop) 3D human reconstruction algorithm stands out. The SMPL (Skinned Multi-Person Linear Model) parametric model used by SPIN is one of the most representative parametric human models in recent years; SMPL is widely used due to its advantages of fewer parameters and faster reconstruction. SMPL is a generative 3D human model that transforms the surface features of the human body into parameters of body shape and pose. The SMPL topology is defined by N = 6890 vertices. Starting from a static template mesh, given the human body shape and pose parameters, the 3D offsets of the vertices are calculated using corresponding functions and added to the template, corresponding to the body shape-dependent deformation and pose-dependent deformation. Finally, the skinning function is used to perform pose transformation on the mesh to obtain the final 3D human mesh. Compared to non-parametric human models, parametric human models have fewer parameters and faster reconstruction speed.

[0003] Efficiently obtaining human body model parameters has become a hot topic in recent years. Currently, the industry generally divides methods into two categories: optimization-based methods and regression-based methods. Optimization-based methods mostly use 2D supervised information (such as joint positions and human body contours) and iteratively optimize the pose and body shape parameters of the human body model to gradually match the 3D model with the 2D supervised data. This method has higher reconstruction accuracy but is slow and sensitive to initial values. Regression-based methods, on the other hand, directly learn the relationship between the 2D image and the human body model parameters through neural networks. This method has fast reconstruction speed and strong robustness but lower accuracy. The SPIN model integrates the two methods, using optimization-based methods to generate high-precision supervised information and regression-based methods to generate reasonable initial values, improving the accuracy and speed of human body reconstruction through mutual promotion. However, both optimization-based and regression-based methods focus on the accuracy of human pose while ignoring the richness of human body shape; therefore, the reconstructed human body models mostly present an average body shape. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to propose a real-time 3D reconstruction system and method based on the SPIN model (hereinafter referred to as the SPIN-MAX system). The model used in the SPIN-MAX system focuses more on the differences between different human body shapes. By using a stacked hourglass model and introducing reprojection loss, clothing semantic segmentation loss, and vertex loss, the reconstruction effect for different human body shapes is effectively improved. Simultaneously, this invention performs lightweight model modifications and carefully designs the data format to ensure that the model still meets real-time requirements in a network environment, introducing excellent algorithms into real-world application scenarios, thus demonstrating practicality. It uses a parametric 3D human body reconstruction model and significantly improves the reconstruction effect of human body shapes, making the interaction between the virtual human body and the scene more realistic and natural.

[0005] The technical solution of the present invention is described in detail below.

[0006] A real-time 3D reconstruction method based on the SPIN model includes the following steps:

[0007] (1) Image reception

[0008] Receive RGB images from the remote camera module;

[0009] (2) Image preprocessing

[0010] After cropping, rotating, scaling and normalizing the received RGB image, it is fed into the ResNet50 network to obtain preprocessed image features;

[0011] (3) Neural Network Operation

[0012] The preprocessed image features are used as input and fed into a SPIN neural network and a stacked hourglass model, respectively. The SPIN neural network is responsible for generating human pose parameters and coarse human body shape parameters. The stacked hourglass model uses the human contour as supervision information to calculate the consistency loss between the generated contour and the real contour, thereby obtaining optimized and refined human body shape parameters.

[0013] Using the output human pose and body shape parameters, a 3D human mesh is reconstructed using the SMPL library. The SoftRas differentiable renderer is then used to project the 3D human mesh onto a 2D model, and the following three loss functions are calculated:

[0014] Reprojection loss: Calculates the pixel-wise mean square error loss between the two-dimensional projected contour and the two-dimensional true contour;

[0015] Clothing semantic segmentation loss: supervise the exposed human body parts and the human body parts under clothing separately;

[0016] Vertex loss: Calculates the vertex-by-vertex distance between the reconstructed 3D human body model and the real 3D human body model;

[0017] (4) Post-processing of results

[0018] The model parameters are rewritten and encapsulated to make them suitable for network transmission;

[0019] (5) Parameter transmission

[0020] The post-processed human model parameters are transmitted to the VR client.

[0021] In this invention, in step (3), the stacked hourglass model is composed of three sub-hourglass modules connected linearly. The output of one sub-hourglass module is the input of the next sub-hourglass module. The consistency loss applied after each sub-hourglass structure is the mean square error loss function on the two-dimensional human body contour. The stacked hourglass structure loss is expressed as:

[0022]

[0023] In the formula:

[0024] Let represent the L2 loss function, n represent the number of sub-hourglass modules, and S represent the 2D human contour image generated by the stacked hourglass model. Represents a realistic two-dimensional human body outline image.

[0025] In this invention, in step (3), the 3D human body mesh is projected onto 2D pixels using the differentiable renderer SoftRas.

[0026] Space; reprojection loss is expressed as:

[0027]

[0028] In the formula:

[0029] ∏(M) represents a 3D human body mesh, and the 2D human body contour image obtained by M-reprojection. Represents a realistic two-dimensional human body outline image.

[0030] In this invention, step (3) uses Semantic Clothing Segmentation (SCS) loss to achieve differential supervision of human clothing; for exposed human body parts, a minimum clothing (Naked, N) loss is applied to encourage close matching between the rendered SMPL body and the human body parts in the image. For human body parts under clothing, a clothing (C) loss is applied to encourage the rendered SMPL body parts to be located inside the clothing; the loss formula is as follows:

[0031] L SCS-N =∑ i,j (R i,j ·di,j (G)) / (∑ i,j R i,j ) 3 / 2 (3)

[0032]

[0033] L SCS =L SCS -N+L SCS-C (5)

[0034] In the formula:

[0035] R i,j This represents the pixel of the reprojected 2D human body contour. When it is located inside the real human body contour G, the distance is d. i,j (G) is 0, otherwise it represents the shortest Euclidean distance from the pixel to G; y i,j L represents the probability value that the current pixel belongs to different human body part labels; SCS This is the final SCS loss.

[0036] In this invention, in step (3), the vertex loss is used to calculate the vertex-by-vertex distance between the reconstructed 3D human body model and the real 3D human body model, and the formula is as follows:

[0037]

[0038] In the formula:

[0039] N is the total number of vertices. Let v represent the L1 loss function. i Represents the vertices of the reconstructed 3D human body model. The vertex represents a realistic 3D human body model.

[0040] This invention also provides a real-time 3D reconstruction system based on the SPIN model, which includes an image receiving module, an image preprocessing module, a neural network module, a result postprocessing module, and a parameter transmission module; including:

[0041] Image receiving module: Receives RGB images from the remote camera module;

[0042] Image preprocessing module: After cropping, rotating, scaling and normalizing the received RGB image, it is fed into the ResNet50 network to obtain preprocessed image features;

[0043] Neural Network Module: The preprocessed image features are used as input and fed into the SPIN neural network and the stacked hourglass model respectively. The SPIN neural network is responsible for generating human pose parameters and coarse human body shape parameters. The stacked hourglass model uses human contour as supervision information to calculate the consistency loss between the generated contour and the real contour, and obtains the optimized fine human body shape parameters. Using the output human pose and body shape parameters, the 3D human body mesh is reconstructed through the SMPL library. The 3D human body mesh is projected to 2D using the SoftRas differentiable renderer, and the following three loss functions are calculated: (1) Reprojection loss: calculate the pixel-wise mean square error loss between the 2D projected contour and the 2D real contour; (2) Clothing semantic segmentation loss: supervise the exposed human body parts and the human body parts under clothing respectively; (3) Vertex loss: calculate the vertex-wise distance between the reconstructed 3D human body model and the real 3D human body model;

[0044] Post-processing module: rewrites and encapsulates model parameters to make them suitable for network transmission;

[0045] Parameter transmission module: Transmits the post-processed human model parameters to the VR client.

[0046] In this invention, the image receiving module and the parameter transmission module use the Flask framework for data reception and transmission.

[0047] In this invention, the SMPL model parameters output by the neural network include 3D camera parameters, 10D human body shape parameters, and 216D human posture parameters. Among these, the relative rotational relationships of the joints are described by quaternions in the human posture parameters.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] The SPIN-MAX system of this invention, by modifying the original SPIN model, achieves more accurate 3D human body reconstruction while maintaining excellent accuracy and speed, modeling the rich differences in body shape between individuals. Simultaneously, SPIN-MAX is more suitable for real-world applications, processing single images input from remote camera modules into SMPL human model parameters representing human posture and body shape, and outputting them to remote client applications. The application of this technology not only improves the accuracy of human body shape reconstruction in 3D human body reconstruction but also introduces new vitality into virtual reality applications, bridging algorithms and industry, and encouraging more academic innovations to be implemented in industry, thus having broad and far-reaching significance.

[0050] The SPIN-MAX system proposed in this invention can complete three-dimensional human body reconstruction tasks in real time. Its internal neural network uses more diverse supervision methods to accurately and realistically reconstruct different human body shapes. At the same time, it is encapsulated in a network framework, which can receive remote requests and return data to the remote client after processing.

[0051] This invention uses a mainstream network framework to wrap the model, enabling it to transmit data. It improves the format of network data transmission while keeping the model lightweight, which greatly increases the speed of human body reconstruction and meets real-time requirements. Attached Figure Description

[0052] Figure 1 This is a flowchart of the real-time 3D reconstruction system based on the SPIN model of this invention.

[0053] Figure 2 This is a flowchart of a stacked hourglass model.

[0054] Figure 3 This is a flowchart of the SPIN-MAX neural network. Detailed Implementation

[0055] This invention focuses on three-dimensional human body reconstruction algorithms, using Python as the foundation and employing tool libraries such as OpenCV (image processing), PyTorch (neural network construction), and Flask (backend service construction), as well as pose representation methods such as quaternions and rotation matrices.

[0056] This invention makes SPIN more suitable for real-world virtual reality applications by modifying it in two aspects:

[0057] 1. Enhance SPIN human body shape supervision to generate more realistic human body shapes: Add a stacked hourglass model during the SPIN network iteration process and use human body contours as supervision information; add a reprojection loss that maps 3D human body mesh to 2D; add a clothing semantic segmentation loss to supervise the exposed human body parts and the human body parts under clothing separately; add a vertex loss to calculate the error at the mesh vertices.

[0058] 2. Modify the network transmission of SPIN: Address key issues including modular processing, network transmission framework, lightweight model modification, and improved data transmission. The neural network model can accept input signals from the network, process them, and output to a designated remote client, enabling it to handle network requests and respond, making it truly applicable in virtual reality scenarios.

[0059] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0060] I. Overview of Real-Time 3D Reconstruction System Process

[0061] The real-time 3D reconstruction system (SPIN-MAX system) based on the SPIN model described in this invention comprises three parts: a camera front-end, a real-time 3D reconstruction back-end, and a display front-end. The system flowchart is as follows: Figure 1 As shown.

[0062] 1. Front end of camera

[0063] (1) Acquire two-dimensional image data captured by the Kinect camera;

[0064] (2) Generate and send an HTTP request to the real-time 3D reconstruction backend.

[0065] 2. Real-time 3D Reconstruction Backend

[0066] (1) Receive the HTTP request from the camera front end and retrieve the image data;

[0067] (2) Determine if the image is valid. If the data is invalid, proceed to (6).

[0068] (3) Preprocess image data;

[0069] (4) Input the data into the neural network to generate SMPL human body model parameters;

[0070] (5) Process and encapsulate the generated parameters to meet the requirements of real-time transmission;

[0071] (6) Generate and send an HTTP response to the display front end.

[0072] 3. Showcase the front end

[0073] (1) Receive the HTTP response from the real-time 3D reconstruction backend;

[0074] (2) Unity parses SMPL human body parameters;

[0075] (3) Determine if the parameter is valid; if the parameter is invalid, discard the response.

[0076] (4) Use HoleLens to display the reconstruction results in three dimensions.

[0077] II. Real-time 3D Reconstruction System

[0078] 1. Stacked hourglass model and consistency loss

[0079] In this invention, we use a stacked hourglass model to generate human body contour information and apply consistency loss for supervision in the intermediate stage. Figure 2 This is a flowchart of a stacked hourglass model.

[0080] The stacked hourglass model consists of multiple linearly connected sub-hourglass modules, each responsible for capturing image information at multiple scales. While local features are crucial for recognizing face and hand information, the final overall contour information requires global learning of the human body. The sub-hourglass module is a concise and lightweight design that captures human contour features at different scales and aggregates these features at the end of the module to output pixel-level accurate predictions. The sub-hourglass module first performs downsampling operations through convolutional and max-pooling layers, gradually learning more global information during the downsampling process. Upon reaching a minimum resolution of 4×4, the module begins bottom-up upsampling and performs cross-scale feature combination. During upsampling, features at the same resolution as those from downsampling are merged, and upsampling continues. Since hourglass features are symmetrical, there must be a corresponding downsampling feature at the same resolution during upsampling. After reaching the output resolution, the hourglass result is used with two 1×1 convolutions to produce the final prediction. The prediction results can be viewed as a heatmap of the human body contour, where a larger value for each pixel indicates that the pixel is more likely to be located inside the human body contour.

[0081] By stacking multiple sub-hourglass structures end-to-end, the output of one sub-hourglass module serves as the input to the next, forming the final stacked hourglass model. The key to the stacked hourglass model lies in the ability to use intermediate error calculations between sub-hourglass modules to ensure the hourglass structure produces the desired output. The processed stacked hourglass result provides a bottom-up and top-down iterative inference mechanism, allowing for staged supervision of the features across the entire image. This multi-hourglass module, multi-intermediate-supervision approach enables the model to gradually deepen its understanding of the human contour structure across multiple scales and stages, mastering human details from local to global perspectives. In the specific implementation of this invention, our stacked hourglass model contains three sub-hourglass modules, and the specific intermediate error applied after each sub-hourglass structure is the L2 loss function on the two-dimensional human contour image. The consistency loss used in this invention can be expressed as:

[0082]

[0083] (1) Let represent the L2 loss function (using minimum mean squared error in this invention), n represent the number of sub-hourglass modules, and S represent the two-dimensional human contour image generated by the stacked hourglass model. Represents a realistic two-dimensional human body outline image.

[0084] 2. Reprojection loss

[0085] In this invention, to further focus on the realism of human body shape reconstruction, we added reprojection loss to the SPIN-MAX system, and the specific implementation method is as follows.

[0086] After generating human pose and body shape parameters using a neural network, a 3D human mesh is reconstructed. Next, the SoftRas differentiable renderer is used to project the 3D human mesh onto a 2D generated human contour, and the pixel-wise mean squared error loss between the 2D projected contour and the 2D real contour is calculated. Therefore, this process requires learning a decoder structure that maps the 3D human mesh to the 2D contour. In this invention, we use the SoftRas differentiable renderer to accomplish this task. It projects the 3D human mesh onto a 2D pixel space, integrates the probability contributions of all mesh triangles to the rendered 2D image, and implements backpropagation to update the gradients of the network parameters. The reprojection loss used in this invention can be expressed as:

[0087]

[0088] (2) In this context, Π(M) represents the two-dimensional human contour image obtained by reprojecting the three-dimensional human body mesh M. This represents a realistic two-dimensional human body contour image. Two-dimensional human body contour information can effectively reflect the body shape, and the body shape of the human body model can be effectively supervised through reprojection loss.

[0089] 3. Clothing semantic segmentation loss

[0090] SMPL is a parametric human body model built using scans of scanned, clothed bodies as real-world data. However, in reality, clothing lengths and thicknesses vary. Therefore, we need to introduce higher-level semantic information about clothing, i.e., calculating losses based on different clothing types and applying different penalties to exposed and covered body parts. The Semantic Clothing Segmentation (SCS) introduced in this paper can effectively accomplish this task. For exposed body parts, the minimum clothing (Naked, N) loss is applied, encouraging a close match between the rendered SMPL body and the body parts in the image. For covered body parts, the Clothing (C) loss is applied, encouraging the rendered SMPL body parts to be located inside the clothing.

[0091] Before applying this loss, we need to use the semantic clothing prior to generate pixel-level labels for the clothed and exposed parts of the images in the dataset as supervision information. The semantic clothing prior generates three types of body part labels for the current image pixels: exposed parts, parts under clothing, and the image background. After the neural network outputs the SMPL model parameters and projects the 3D human body mesh onto a 2D model, the supervision information generated by the semantic clothing prior is used to calculate the loss. The SCS loss formula is as follows:

[0092] LSCS-N =∑ i,j (R i,j ·d i,j (G)) / (∑ i,j R i,j ) 3 / 2 (3)

[0093]

[0094] L SCS =L SCS -N+L SCS-C (5)

[0095] In the formula:

[0096] R i,j This represents the pixel of the reprojected 2D human body contour. When it is located inside the real human body contour G, the distance is d. i,j (G) is 0, otherwise it represents the shortest Euclidean distance from the pixel to G; y i,j L represents the probability value that the current pixel belongs to different human body part labels; scs This is the final clothing semantic segmentation loss.

[0097] It is important to note that differential supervision of human clothing is a supplement to reprojection loss, but cannot completely replace it. Due to the uncertainty of semantic clothing priors, we need reprojection loss to constrain the reconstructed human body as a whole, preventing the generation of distorted human body shape parameters. Through differential supervision of human clothing, we can not only use the exposed parts of the human body for accurate reconstruction, but also use the parts of the human body under the clothing to further constrain the human body shape parameters, preventing reconstruction distortion caused by generating SMPL parameters that are too closely fitted to the clothing.

[0098] 4. Vertex Loss

[0099] The original SPIN model defined a vertex loss for the 3D model, which can calculate the vertex-by-vertex distance between the generated 3D human body model and the real 3D human body model. In the original SPIN training, this loss had a weight of 0, meaning it was not enabled. In this invention, we re-enable this loss, and its definition is as follows:

[0100]

[0101] In the formula, N is the total number of vertices. Let v represent the L1 loss function (mean absolute error is used in this invention). i Represents the vertices of the reconstructed human body model. Represents the vertex of a real human body model.

[0102] 5. Network data transmission

[0103] (1) Modular backend service

[0104] The SPIN-MAX system consists of multiple independent sub-modules, as defined below:

[0105] Image receiving module: Receives RGB images from the remote camera module;

[0106] Image preprocessing module: Normalizes and standardizes RGB images;

[0107] Neural network module: Generates SMPL model parameters using preprocessed image information;

[0108] Post-processing module: rewrites and encapsulates model parameters to make them suitable for network transmission;

[0109] Result transmission module: Transmits the post-processed human model parameters to the VR client.

[0110] This invention uses only predefined data formats for data transmission between different modules. Modular services reduce the coupling between different functions, improve service scalability, and allow individual modules to be replaced or expanded as needed.

[0111] (2) Flask framework

[0112] This invention uses the Flask framework to implement network data transmission. Flask is a micro-network framework written in Python and does not require specific tool libraries. In this invention, the image receiving module and the parameter transmission module need to use the Flask framework for data reception and transmission.

[0113] (3) Lightweight transformation

[0114] The SPIN-MAX system follows the SPIN model's image preprocessing approach, which involves segmenting the original image by generating the smallest bounding box of the human body to improve reconstruction accuracy. Experiments revealed that the smallest bounding box generation algorithm is extremely time-consuming, typically accounting for one-third of the backend processing time, but only resulting in a 0.2% improvement in accuracy. Therefore, we decided to remove the image segmentation preprocessing stage and perform subsequent processing directly at the original image resolution. This not only improves the backend processing speed but also, because the original image resolution contains richer spatial information, the camera parameters generated by the neural network are more accurate, thus improving the spatial accuracy of the reconstructed human body model.

[0115] (4) Selection of transmitted data

[0116] The output of the neural network in the SPIN-MAX system consists of 3D camera parameters, 10D human body shape parameters, and 216D human pose parameters. The 216-dimensional parameters are composed of rotation matrices for 24 joints (each rotation matrix contains 9 parameters), which places a significant burden on data transmission and reduces the real-time performance of the backend service. Quaternions, on the other hand, contain only 4 parameters and can be losslessly converted into rotation matrices without any loss of accuracy. Therefore, this invention uses quaternions instead of rotation matrices to describe the relative rotation relationships of joints.

[0117] Experiments in real-world applications demonstrate that the SPIN-MAX system can process 1280*720 input images from a remote camera module into precise human model parameters containing accurate posture and body shape information, and then return these parameters to the remote client. This technology introduces advanced parametric 3D human reconstruction techniques into augmented reality applications, and through modifications, achieves more accurate human body shape reconstruction, making interactions with virtual objects in VR applications more realistic. Our real-time 3D reconstruction system achieves a frame rate of 31 FPS, meeting real-time requirements.

Claims

1. A real-time 3D reconstruction method based on the SPIN model, characterized in that, Includes the following steps: (1) Image reception Receive RGB images from the remote camera module; (2) Image preprocessing After cropping, rotating, scaling and normalizing the received RGB image, it is fed into the ResNet50 network to obtain preprocessed image features; (3) Neural Network Operation The preprocessed image features are used as input and fed into the SPIN neural network and the stacked hourglass model respectively. The SPIN neural network is responsible for generating human pose parameters and coarse human body shape parameters. The stacked hourglass model uses human contour as supervision information to calculate the consistency loss between the generated contour and the real contour, and obtains the optimized fine human body shape parameters. Using the output human pose and body shape parameters, a 3D human mesh is reconstructed using the SMPL library; the 3D human mesh is projected onto a 2D model using the differentiable renderer SoftRas, and the following three loss functions are calculated: Reprojection loss: Calculates the pixel-wise mean square error loss between the two-dimensional projected contour and the two-dimensional true contour; Clothing semantic segmentation loss: supervise the exposed human body parts and the human body parts under clothing separately; Vertex loss: Calculates the vertex-by-vertex distance between the reconstructed 3D human body model and the real 3D human body model; (4) Post-processing of results The model parameters are rewritten and encapsulated to make them suitable for network transmission; (5) Parameter transmission The post-processed human model parameters are transmitted to the VR client.

2. The real-time three-dimensional reconstruction method according to claim 1, characterized in that, In step (3), the stacked hourglass model is composed of three linearly connected sub-hourglass modules. The output of one sub-hourglass module is the input of the next sub-hourglass module. The consistency loss applied after each sub-hourglass structure is the mean square error loss function on the two-dimensional human body contour. The stacked hourglass structure loss is expressed as: In the formula: Let represent the L2 loss function, n represent the number of sub-hourglass modules, and S represent the 2D human contour image generated by the stacked hourglass model. Represents a realistic two-dimensional human body outline image.

3. The real-time three-dimensional reconstruction method according to claim 1, characterized in that, In step (3), the 3D human body mesh is projected onto a 2D pixel space using the differentiable renderer SoftRas; the reprojection loss is expressed as: In the formula: Π(M) represents a 3D human body mesh, and the 2D human body contour image obtained by M-reprojection. Represents a realistic two-dimensional human body outline image.

4. The real-time three-dimensional reconstruction method according to claim 1, characterized in that, In step (3), the clothing semantic segmentation loss SCS is used to achieve differential supervision of human clothing; for exposed human body parts, the minimum clothing N loss L is applied. SCS-N It encourages a close match between the rendered SMPL body and the human body parts in the image; for the human body parts under clothing, it applies the clothing loss (L). SCS-C The rendering of SMPL body parts is encouraged to be located inside the clothing; the loss formula is as follows: L SCS-N =∑ i,j (R i,j ·d i,j (G)) / (∑ i,j R i,j ) 3 / 2 (3) L SCS L SCS-N +L SCS-C (5) In the formula: R i,j This represents the pixel of the reprojected 2D human body contour. When it is located inside the real human body contour G, the distance is d. i,j (G) is 0, otherwise it is the shortest Euclidean distance from the pixel to G; y i,j L represents the probability value that the current pixel belongs to a different human body part label; SCS This is the final clothing semantic segmentation loss.

5. The real-time three-dimensional reconstruction method according to claim 1, characterized in that, In step (3), the vertex loss is used to calculate the vertex-by-vertex distance between the reconstructed 3D human body model and the real 3D human body model. The formula is as follows: In the formula: N is the total number of vertices. Let v represent the L1 loss function. i Represents the vertices of the reconstructed 3D human body model. The vertex represents a realistic 3D human body model.

6. A real-time 3D reconstruction system based on the SPIN model, characterized in that, It includes an image receiving module, an image preprocessing module, a neural network module, a result postprocessing module, and a parameter transmission module; wherein: Image receiving module: Receives RGB images from the remote camera module; Image preprocessing module: After cropping, rotating, scaling and normalizing the received RGB image, it is fed into the ResNet50 network to obtain preprocessed image features; Neural Network Module: The preprocessed image features are used as input and fed into the SPIN neural network and the stacked hourglass model respectively. The SPIN neural network is responsible for generating human pose parameters and coarse human body shape parameters. The stacked hourglass model uses human contour as supervision information to calculate the consistency loss between the generated contour and the real contour, and obtains the optimized fine human body shape parameters. The output human pose and body shape parameters are used to reconstruct the three-dimensional human body mesh through the SMPL library. The SoftRas differentiable renderer is used to project the three-dimensional human body mesh onto two dimensions and calculate the following three loss functions: (1) Reprojection loss: calculate the pixel-wise mean square error loss between the two-dimensional projected contour and the two-dimensional real contour; (2) Clothing semantic segmentation loss: supervise the exposed human body parts and the human body parts under the clothing respectively; (3) Vertex loss: calculate the vertex-wise distance between the reconstructed three-dimensional human body model and the real three-dimensional human body model. Post-processing module: rewrites and encapsulates model parameters to make them suitable for network transmission; Parameter transmission module: Transmits the post-processed human model parameters to the VR client.

7. The real-time three-dimensional reconstruction system according to claim 6, characterized in that, The image receiving module and parameter transmission module use the Flask framework for data reception and transmission.

8. The real-time three-dimensional reconstruction system according to claim 6, characterized in that, The SMPL model parameters output by the neural network include 3D camera parameters, 10D human body shape parameters, and 216D human posture parameters. Among them, the relative rotational relationships of the joints are described by quaternions in the human posture parameters.

Citation Information

Patent Citations

  • Method for reconstructing dressed human body model from image based on image convolution

    CN113077545A

  • Method and apparatus for reconstructing three-dimensional model of human body, and storage medium

    US20210012558A1