Three-dimensional human body reconstruction method based on implicit neural network and diffusion model
Through the method based on implicit neural network and diffusion model, a three-dimensional human body model is learned from a single image, which solves the problem of insufficient information in the occlusion area and realizes high fidelity and consistency of three-dimensional human body reconstruction.
Patent Information
- Application Number
- CN202510198911.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art faces the problem of insufficient occlusion area information when reconstructing a three-dimensional mannequin from a single random angle image, resulting in poor results when generating new views that hardly overlap with the input views.
A three-dimensional human body reconstruction method based on implicit neural networks and diffusion models is adopted to learn human body neural radiation fields with identity characteristics from a single reference image. Through an improved implicit neural network and denoising diffusion probability model framework, combined with multi-scale structural similarity constraints, a high fidelity and consistency three-dimensional human body model is generated.
Effectively retain the detailed features in the original image, make reasonable predictions of unobserved areas, and generate a three-dimensional mannequin with high fidelity and consistency, improving the reconstruction quality and expanding the applicability of the model under sparse input conditions.
Smart Images

Figure CN119991967A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional reconstruction, and in particular relates to a three-dimensional human body reconstruction method based on implicit neural network and diffusion model. Background Art
[0002] In the field of virtual reality and augmented reality, efficiently deriving high-fidelity 3D human representation from 2D images remains a key challenge. Traditional methods usually use pixels, point clouds or neural network weights to represent the 3D morphology of the human surface mesh. However, these methods have limited control over the generation process, complex equipment configuration and high data capture costs. In order to solve these problems, technologies based on parametric representation of the human body have developed rapidly in recent years, among which research based on the Skinned Multi-Person Linear Model (SMPL) occupies an important position.
[0003] Although parametric 3D models excel in representing human structures, they are limited in depicting the complexity of individuals in clothing. Recently, implicit neural radiance field (NeRF) methods have demonstrated significant potential in capturing complex surface geometry, especially in modeling high-frequency textures such as hair and clothing. These NeRF-based human reconstruction methods have shown excellent performance under dense input conditions such as multi-view images or video sequences acquired through a multi-camera system.
[0004] However, the reliance of these methods on high-quality dense inputs limits their application in practical scenarios. For ordinary users, it is usually difficult and costly to obtain multi-view images or video sequences, especially when only a single image with a random angle is available in the real world. To address this problem, a generalized single-image-based character NeRF method can be used to generate avatars. However, due to the lack of information in occluded areas, these methods face great challenges in generating high-quality new views that barely overlap with the input view. Summary of the invention
[0005] In view of the above problems, this paper proposes a 3D human body reconstruction method based on implicit neural network and diffusion model, which aims to learn the human neural radiation field with identity characteristics from a single reference image, and use back propagation to determine the Pareto optimal solution, and finally realize the generation of high-fidelity 3D human body model under free viewing angle.
[0006] The overall process of the present invention mainly consists of three model components to support high-fidelity rendering:
[0007] First, the implicit neural network is improved to enable it to encode diverse conditional feature information from the reference view. This improvement significantly enhances the model's ability to represent input information and lays the foundation for subsequent generation.
[0008] Second: Considering that the credible information of the input image is usually related to the view angle, a denoising diffusion probability model framework is introduced to learn independent information that is independent of the view angle from the feature-guided density volume. This method ensures that the generated 3D model has good view angle consistency under sparse input conditions.
[0009] Third: In order to further enhance the model's optimization ability for unseen data distribution, a multi-scale structural similarity (MSSC) constraint is proposed. This constraint significantly improves the model's reconstruction accuracy for high-frequency and complex features during the diffusion process.
[0010] The technical solution adopted by the present invention is:
[0011] A 3D human body reconstruction method based on an implicit neural network and a diffusion model, the 3D human body reconstruction method based on an implicit neural network and a diffusion model comprises the following steps:
[0012] Step 1: Sample the SMPL canonical space in the target space θ, β through the inverse linear mixed montage algorithm LBS. Given a single image of the reference view, the global feature f g and local features f l Integrated into the coordinate system of the target input view to facilitate feature volume fusion;
[0013] Step 2: Intermediate density volume mapped from a series of MLP layers are fed into an identity-based diffusion model that incorporates view description embeddings from the CLIP ViT / L-14 model;
[0014] Step 3: Add the multi-scale structural similarity index MSSC when calculating the loss, and use backpropagation to determine the Pareto optimal solution.
[0015] Furthermore, in step 1, global features and local features are added, including:
[0016] First, from a single reference image I o Extract a characteristic volume V(G(I o ), the volume is used as a human body prior, with an axis-aligned triplanar feature representation, and is separated from the StyleGAN2 feature generator by the rescaled aligned space vector;
[0017] The reference image is constructed by compressing the entire scene into a compact latent code φ global (x), thereby ensuring the compactness of features while retaining their key information; using the SMPL model to extract local level features, extracting the features of each point from the visible pose vertices and merging them into the normalized appearance volume;
[0018] By projecting the 3D deformation points into the input view, we obtain the local alignment feature φ local (x);
[0019] The aggregate feature F extracted using the projection function Π is F = Π(φ global (x),φ local (x)), together with the embedding coordinates γ(x c ), which is then fed into a multi-layer perceptron network (MLP). Used to predict density volume and color coefficients.
[0020] Furthermore, in step 2, view description embedding is added, including:
[0021] After the density volume representation, the density features are mixed with a Gaussian noise vector ε~N(0,1) in T steps, using a 2D mask to improve the guidance of generative modeling. At each time step t, a noise feature is generated. This noise signature integrates the The forward process is formulated as follows:
[0022]
[0023] in is the noise coefficient, d is the viewing angle parameter of the image;
[0024] The overall loss is optimized by the reverse autoencoder. While fixing the pre-trained U-Net parameters, the NeRF and Low-Rank Adaptation parameters are updated to identify the Pareto frontier s in the feasible solution space S, where the target space R is expressed as:
[0025]
[0026] Where y refers to the guidance text embedding, the global optimal solution Including the balance coefficient w and the prediction noise U-Net∈;
[0027] Diffusion distribution of score distillation samples at time t tends to the entire marginal distribution p from the pre-trained text to image model;
[0028] We use the parameterized fraction to model the Wasserstein gradient flow and train an optimization procedure end-to-end that simultaneously optimizes the 3D parameters θ d and LoRA parameter φ; where 3D parameter θ d For pre-trained models and The Wasserstein gradient flow between φ and Gaussian distributed noise ∈.
[0029] Furthermore, in step 3, a multi-scale structural similarity index is added, including:
[0030] The core formula of the multi-scale structural similarity constraint is as follows:
[0031]
[0032] where MSSC is a random patch P extracted from the rendered image I and the ground truth i image (n) Calculated using αM, β j and γ j To adjust the relative importance of different components; the brightness component l, the contrast component c and the structure component s are calculated using a Gaussian weighting formula within a 5 × 5 pixel filter.
[0033] Furthermore, in the loss calculation, the following combination of loss functions is used to achieve high-quality 3D human body rendering:
[0034] Rendering Loss: We use the image-scale L2 loss to ensure the overall consistency between the lighting color rendering and the texture-free rendering, given the ground truth target image C(r) and the predicted image The rendering loss contributes to the identity feature distribution by progressively denoising the normally distributed variables in the backward process:
[0035]
[0036] in Represents a set of 3D query points on a ray;
[0037] MSSC loss: By reconstructing SSIM, it is used to detect the consistency between the identity space and the appearance latent space. For the convenience of numerical calculation, it is normalized to the form of 1-MSSC:
[0038]
[0039] LPIPS loss: In order to solve the problem of inconsistent details caused by partial reflection or shadow, the spatial consistency and seamless transition characteristics in human perception are captured through a convolutional neural network, and the VGG-based pre-trained network prior is used to accelerate the convergence of the model:
[0040]
[0041] The overall loss is expressed as:
[0042]
[0043] Among them, λ1=0.1, λ2=0.1.
[0044] The beneficial effect of the present invention is that compared with the traditional 3D human body reconstruction method, the present method can more effectively preserve the detailed features in the original image, and at the same time make reasonable predictions for the unobserved areas, thereby generating a 3D human body model with high fidelity and consistency. This improvement not only improves the reconstruction quality, but also expands the applicability of the model under sparse input conditions, providing a more flexible and efficient solution for practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Flowchart of the 3D human body reconstruction method based on implicit neural network and diffusion model.
[0046] Figure 2 Multi-scale structural similarity constraint (MSSC) structure diagram.
[0047] Figure 3 Quantitative results.
[0048] Figure 4 Qualitative results.
[0049] Figure 5 Quantitative results of ablation experiments.
[0050] Figure 6 Qualitative results of ablation experiments.
[0051] Figure 7 Restoration effect of random occlusion. DETAILED DESCRIPTION
[0052] like Figure 1 As shown in FIG. 1 , the process of the present invention includes three main steps: implicit neural representation, identity-based diffusion, and optimization of loss calculation. Step 1: Sample the SMPL canonical space in the target space α, β using the reverse linear mixed montage (LBS) algorithm. For a single reference view image, extract its global feature f g and local features f l , and integrate these features into the coordinate system of the target input view, thereby facilitating the fusion of feature volumes. Step 2: Intermediate density volume mapped from a series of MLP layers It is fed into an identity-based diffusion model that combines the view description embeddings from the CLIP ViT / L-14 model. Step 3: During the training process, the multi-scale structural similarity index (MSSC) is introduced in the computational loss to enhance the model's performance in detail reconstruction and global consistency. Through back-propagation optimization, this method can effectively determine the Pareto optimal solution, balancing the fidelity of the generated model and the rationality of the prediction.
[0053] Through the above process, the present invention realizes high-quality three-dimensional human body reconstruction driven by a single image, taking into account both the preservation of input features and the prediction effect of unseen areas.
[0054] The following is a detailed description of the specific steps of the system of the present invention:
[0055] 1. Implicit Neural Rendering Controlled by Single View Condition
[0056] In the process of 3D human body rendering, directly projecting unseen rays from the center of the camera to each pixel position on the image plane often produces certain uncertainties. To solve this problem, the present invention adopts a NeRF-based image conditional rendering method, in which the implicit neural field is modeled through a multi-layer perceptron (MLP) using the geometric and appearance information extracted from the input image.
[0057] Specifically, we first start from a single reference image I o Extract a characteristic volume V(G(I o )), which is used as a human body prior. The present invention adopts an axis-aligned triplane feature representation and separates the StyleGAN2 feature generator from the rescaled aligned space vector, so that the method can clearly capture high-level attributes as well as random variations, thus achieving richer visual effects.
[0058] In addition, in order to ensure the compactness of the features and effectively retain their essential expression, the present invention compresses the reference image into a concise latent code φ by compressing the entire scene global (x), thus ensuring the compactness of features while retaining their key information.
[0059] The present invention not only relies on global information, but also uses the SMPL model to extract local level features to further enrich the intrinsic information extracted from a single input image. Specifically, the present invention extracts the features of each point from the visible pose vertices and merges them into the normalized appearance volume. By projecting the 3D deformed points into the input view, the present invention obtains the local alignment feature φ local (x),These features help capture fine-grained local details, which are crucial for compensating for partially missing data and achieving global understanding.
[0060] As shown in formula (1), the aggregate feature F extracted by the projection function ∏ is F=∏(φ global (x),φ local (x)), together with the embedding coordinates γ(x c ), which is then fed into a multi-layer perceptron network (MLP). For predicting density volume and color coefficients. By utilizing a layered geometric representation, the present invention effectively models the body contour and its components, thereby achieving a coherent 3D perceptual generalization.
[0061]
[0062] Here, a and β control the topological deformation in the observation space.
[0063] 2. Diffusion model based on identity characteristics
[0064] In the real world, dealing with the 3D consistency problem between different views of the same scene usually relies on rich view information. However, existing single-view 3D human reconstruction methods often have difficulty generating high-quality domain-specific images under extreme view differences (e.g., from rear-view input and front-view output) because the human identity distribution itself is more biased towards the reference view. This limitation makes it impossible for existing methods to accurately model under extreme view differences.
[0065] To solve this problem, the present invention introduces a diffusion model based on identity features, which can balance the visible features from the reference character and the topological rationality from the unseen parts under extreme view difference conditions. The method of the present invention simultaneously processes the probability distribution across the implicit density space and the semantic information space, which can be regarded as a Pareto optimal problem.
[0066] After the density volume representation, the present invention mixes the density features with a Gaussian noise vector ε~N(0,1) in the total time step T. In order to emphasize the accurate data distribution, the present invention uses a 2D mask to improve the guidance of generative modeling. For each time step t, the noise feature is generated It integrates The masked region from and the unmasked region from pure noise. The forward process is formulated as follows:
[0067]
[0068] in is the noise coefficient, and d is the viewing angle parameter of the image.
[0069] Subsequently, in order to optimize the model and learn the denoising process, the present invention optimizes the overall loss through the reverse autoencoder. Since the optimization of the reverse autoencoder depends on the variational 3D generation parameters, the NeRF and Low-Rank Adaptation (LoRA) parameters are updated while the pre-trained U-Net parameters are fixed. The ultimate goal is to identify the Pareto frontier s in the feasible solution space S, where the target space R is expressed as:
[0070]
[0071] Where y refers to the guidance text embedding, the global optimal solution Including the balancing coefficient w and the prediction noise U-Net∈.
[0072] Diffusion distribution of score distillation sampling (SDS) at time t tends to the entire marginal distribution p from the pre-trained text to image model. Obviously, SDS may cause over-saturation and low diversity problems due to the single Gaussian distribution.
[0073] The present invention is inspired by variational distribution (VSD), but different from it, the present invention uses parameterized scores to model Wasserstein gradient flow and trains the optimization process end-to-end. More specifically, the present invention simultaneously optimizes the 3D parameters θ d (Pre-trained model and Wasserstein gradient flow between) and the LoRA parameter φ (in the variational distribution ∈ φ and Gaussian distribution noise ∈). A feasible compact set decision is heuristically found in the metric space, so that the present invention can obtain satisfactory results by rendering the image in one go.
[0074] By solving the Pareto optimality problem, the model of the present invention can guide data samples to a low-dimensional manifold perturbation space with an occlusion space distribution assumption. This optimization strategy not only improves the credibility of missing areas in local perspectives, but also enhances the performance of overall rendering. Based on the non-equilibrium optimization strategy, the present invention can compensate for missing areas from partial views and improve the global consistency of generating three-dimensional human bodies. Overall, the diffusion model based on identity features provides an effective solution for 3D human body reconstruction under extreme viewing conditions, significantly improving rendering quality and topological consistency.
[0075] 3. Multi-scale structural similarity constraints
[0076] In the field of image generation, mean squared error (MSE) is a classic loss function that is widely used because of its simplicity and ease of implementation. MSE is used to minimize the pixel-level error between synthetic images and ground truth images by quantifying the squared difference between the predicted value and the true value. However, MSE loss treats each pixel as equally important, and this uniform treatment may cause the model to be overly sensitive to outliers (such as inaccurate backgrounds or blurred areas). With single-pixel supervision, rendering images of the human body with complex or diverse clothing often falls into suboptimal solutions.
[0077] In order to solve this problem, the present invention designs a Figure 2The loss function shown, Multi-Scale Structural Similarity Constraint (MSSC), is used to ensure spatial alignment and global consistency. Different from the traditional MSE, MSSC balances multiple rendered human regions by integrating multi-scale structural information. Its core idea is based on a pyramid-shaped random patch information extractor, which regularizes the quality of multiple sampled light batches to ensure that the rendering results have higher structural fidelity. The core formula of the multi-scale structural similarity constraint is as follows:
[0078]
[0079] Among them, MSSC is obtained from the rendered image I and the ground truth A random patch P extracted from the image (n) Calculated. Using αM, β j and γ j The relative importance of different components is adjusted. In addition, the brightness component l, the contrast component c, and the structure component s are calculated using a Gaussian weighted formula within a 5×5 pixel filter. The present invention constructs a feature pyramid, applies M low-pass filters layer by layer, and downsamples the filtered image semantic information. According to the experimental results of the present invention, it is observed that a two-layer pyramid with both high-level and low-level weights set to 0.5 produces the best results.
[0080] A higher MSSC value indicates a higher structural similarity between the rendered image and the real image, which fully captures the correlation between adjacent pixels. This constraint shows significant superiority in overcoming the image degradation problem that is easily caused by the traditional L2 loss.
[0081] 4. Loss Function
[0082] In order to achieve high-quality 3D human rendering, this study uses a combination of multiple loss functions in training to ensure identity consistency, appearance fidelity, and topological rationality. Specifically, it includes the following three losses:
[0083] Rendering loss: Considering the uncontrollable 3D surface appearance information, we first use the image-scale L2 loss to ensure the overall consistency between the lighting color rendering and the texture-free rendering. Given the ground truth target image C(r) and the predicted image The rendering loss contributes to the identity feature distribution by progressively denoising the normally distributed variables in the backward process:
[0084]
[0085] in Represents a collection of 3D query points on a ray.
[0086] MSSC loss: By reconstructing SSIM, this loss is used to detect the consistency between the identity space and the appearance latent space. In order to facilitate numerical calculation, the present invention normalizes it to the form of 1-MSSC:
[0087]
[0088] LPIPS loss: To solve the problem of inconsistent details caused by partial reflection or shadow, LPIPS loss captures the spatial consistency and seamless transition characteristics in human perception through convolutional neural networks. The present invention uses the VGG-based pre-trained network prior to accelerate the convergence of the model:
[0089]
[0090] This study uses an end-to-end training approach and the overall loss is expressed as:
[0091]
[0092] Among them, λ1=0.1, λ2=0.1.
[0093] Experimental results:
[0094] 1. Dataset and evaluation metrics
[0095] The present invention conducts experiments in single-view configuration on two widely used human datasets to comprehensively evaluate the reconstruction performance of the model in complex human action scenes. First, on the ZJU-MoCap dataset, using its rich human motion capture data, the 3D human reconstruction task in dynamic scenes is tested. The present invention divides the data into 9 subjects, 6 of which are used for training and 3 for testing. Secondly, on the THuman scan dataset, using the high-quality human scan data it provides, complex scenes with a variety of costumes and postures are covered. To ensure the comprehensiveness of the experiment, the present invention selects 90 subjects for training and uses 10 subjects as the test set.
[0096] In order to accurately measure the performance of the model, the present invention introduces multiple evaluation indicators. The peak signal-to-noise ratio (PSNR) is used to quantify the pixel-level similarity between the generated image and the ground truth image, reflecting the overall quality of the rendering; the structural similarity index (SSIM) is used to evaluate the structural information fidelity of the image, focusing on local texture and global consistency; the learning-perceptual image patch similarity (LPIPS) is a perceptual similarity indicator based on deep learning, which has been widely used in recent NeRF-based human rendering research. It can refine the evaluation of visual fidelity and provide quality evaluation consistent with human perception.
[0097] In addition, our experiments adopt a more sophisticated mask region evaluation method. Instead of directly calculating the metrics for the entire image, we follow the common practice in human NeRF methods and calculate the PSNR, SSIM, and LPIPS metrics only within the mask region by projecting the 3D bounding box of the human body on the image plane.
[0098] 2. Experimental details
[0099] The optimization of a single scene typically requires about 120,000 iterations to converge on a single NVIDIA A40 GPU. To be consistent with NeRF training, we selected the first 90 individuals in the THuman dataset and individuals [386,387,390,392,393,394] in the ZJU dataset. For NeRF, we only draw rays within the body mask for supervision, and each sample samples 48 coarse coordinates and 24 fine coordinates in the hierarchical volume representation.
[0100] The present invention uses the LoRA plugin to perform 20 end-to-end fine-tuning training iterations on the pre-trained Stable Diffusion v1.5 checkpoint. For MSSC, the patch size is 64×64, and the hyperparameter C1=(K1L) 2 , C2=(K2L) 2 , where K1=0.01 and K2=0.03 are used to calculate the three comparison metrics. The present invention uses the Adam optimizer to simultaneously optimize the implicit generation module and the identity-based diffusion component, with a learning rate of 1e-4 and a batch size of 5.
[0101] 3. Quantitative results
[0102] A comprehensive comparison is made between the proposed method and the state-of-the-art methods on THuman and ZJU-MoCap datasets. Considering the objective evaluation index obtained by neural network training, the proposed method selects LPIPS as the most representative index, which accurately reflects the human eye's perception of rendering quality. Representation: The metric results under the back view input setting when the target view is uniformly sampled from the perspective difference [0○, 360○]. The best, second and third scores are marked in red, orange and yellow respectively. The proposed method shows excellent performance in most metrics for different input view configurations.
[0103] The experiments were evaluated in two settings: arbitrary view setting and fixed input view setting. In the arbitrary view setting, the input view is random and the target view is randomly assigned; in the fixed input view setting, the input view is fixed to three predetermined views and the remaining views are tested as output views. The final results are obtained by calculating the average.
[0104] like Figure 3 As shown in the figure, the method of the present invention performs well in SSIM and LPIPS in the two datasets. In terms of the PSNR indicator, when given an arbitrary viewing angle input, the result of SHERF is slightly higher than that of the method of the present invention. However, it should be noted that PSNR mainly emphasizes the difference at the pixel level and ignores human visual perception, so it may not be able to effectively capture the fine-grained details of unseen areas. Among the three indicators, LPIPS is considered to be the indicator that best reflects the perceptual quality.
[0105] In addition, the present invention also evaluates the performance of Neural Body and MPS-NeRF in new pose synthesis tasks. Even though these methods provide multiple perspectives as input, they still fail to provide ideal performance in cases with large pose variations. For methods based on per-subject optimization such as PixelNeRF and ELICIT, the present invention used default parameters for training. It was found that these methods depend on specific datasets and perform poorly outside the training distribution. Although SHERF is good at utilizing hierarchical reference feature information, it has difficulties in inferring the distribution of occluded human body parts, limiting its ability to be generalized. In contrast, the method of the present invention ensures topological rationality through identity-based diffusion and exhibits stronger predictive ability.
[0106] Finally, we also test the performance of our method under back view input. The experimental results show that our method outperforms other state-of-the-art methods even at extreme viewing angles. This highlights the ability of our method to generate robust images under different viewing angles, emphasizing its wider applicability, especially in complex and diverse scenes.
[0107] 4. Qualitative results
[0108] exist Figure 4 In the figure, the new perspective synthesis results of the model of the present invention and other four methods under different perspective conditions are shown. In order to evaluate the effect of the method, the present invention visualizes the rendering results under small difference and large difference scenarios respectively. In the case of small difference (the first and third rows), the difference between the input perspective and the target perspective is small, and the method of the present invention can accurately restore the details, especially the small details such as clothing (such as T-shirts) and glasses are preserved with high fidelity, enhancing the overall visual perception of the subject. In addition, by combining the prior information of the diffusion model, the method of the present invention reduces the dependence on the input information and avoids simply copying unnecessary patterns in the input view.
[0109] In the case of large differences (the second and fourth rows), when the difference between the input and target views is large, the method of the present invention still performs well, successfully retaining the features in the original image and making reasonable predictions for unobserved areas. For example, the present invention can accurately restore the texture of the front of the clothing and the details of the face, which are not captured in the input view. This ability enables the method of the present invention to generate consistent and natural images in situations with large differences in viewpoints.
[0110] 5. Ablation experiment
[0111] This paper conducts ablation studies on the THuman dataset, using the same experimental settings as the previous experiments. Figure 5 and Figure 6 As shown in Figure 2, it shows how global and local features can ensure basic RGB color accuracy. Figure 6 In the first row of , the introduction of the diffusion module significantly enhances the plausibility within the occluded region, especially in the hat part. When the front view input is provided, the diffusion module successfully makes a reasonable prediction of the back region and generates a hat that cannot be seen in the input view. In this way, the diffusion module effectively fills the occluded area and ensures the coherence of the image.
[0112] In addition, the key component MSSC plays an important role in the image rendering process. It effectively reduces the appearance of artifacts, preserves high-frequency details, and improves the accuracy and naturalness of the rendering effect. Overall, the combination of MSSC loss and diffusion model enables NeRF to capture non-local structural information, further improving all performance indicators, proving the effectiveness of these new components in improving image quality.
[0113] 6. Robustness of restoring occluded images
[0114] MPS-NeRF, SHERF and the proposed method were selected to evaluate the generalization ability under occlusion conditions. The results show that the proposed method can effectively avoid serious artifacts, while other methods are prone to sampling errors in the corresponding overlapping areas, resulting in failure.
[0115] Considering that human images in the real world are often occluded, the present invention tests the robustness of the network reconstruction under damaged data conditions. Specifically, the present invention artificially damages the images by randomly occluding 1 / 25 pixels of the two test object images in the THuman dataset. Figure 7 Visualization results of comparative experiments are shown, which include a corrupted single-view image as input.
[0116] In the experiments, MPS-NeRF and SHERF showed obvious omissions in the occluded areas, indicating that they are too dependent on the input view. In particular, when the input image is partially occluded, the process of projecting features into 3D space using 2D input projection becomes invalid. In contrast, the proposed method effectively estimates the imperceptible occluded areas and human appearance content by leveraging the prior knowledge provided by the diffusion model, thereby showing stronger robustness and accuracy under occlusion conditions.
[0117] It should be further explained that the above implementation modes are only used to understand the technical solution of the present invention, and are not used to limit the protection scope of the present invention. Any obvious adjustments and modifications made to the technical solution of the present invention that belong to the technical concept of the present invention should also fall within the protection scope of the present invention.
Claims
1. A three-dimensional human body reconstruction method based on implicit neural network and diffusion model, characterized in that: The three-dimensional human body reconstruction method based on implicit neural network and diffusion model comprises the following steps: Step 1: Sample the SMPL canonical space in the target space θ, β through the inverse linear mixed montage algorithm LBS. Given a single image of the reference view, the global feature f g and local features f l Integrated into the coordinate system of the target input view to facilitate feature volume fusion; Step 2: Intermediate density volume mapped from a series of MLP layers are fed into an identity-based diffusion model that incorporates view description embeddings from the CLIPViT / L-14 model; Step 3: Add the multi-scale structural similarity index MSSC when calculating the loss, and use backpropagation to determine the Pareto optimal solution.
2. The three-dimensional human body reconstruction method based on implicit neural network and diffusion model as claimed in claim 1, characterized in that: In step 1, global features and local features are added, including: First, from a single reference image I o Extract a characteristic volume V(G(I o ), the volume is used as a human body prior, with an axis-aligned triplanar feature representation, and by separating the StyleGAN2 feature generator from the rescaled aligned space vector; The reference image is constructed by compressing the entire scene into a compact latent code φ global (x), thereby ensuring the compactness of features while retaining their key information; using the SMPL model to extract local level features, extracting the features of each point from the visible pose vertices and merging them into the normalized appearance volume; By projecting the 3D deformation points into the input view, we obtain the local alignment feature φ local (x); The aggregate feature F extracted using the projection function Π is F = Π(φ global (x),φ local (x)), together with the embedding coordinates γ(s c ), which is then fed into a multi-layer perceptron network (MLP). Used to predict density volume and color coefficients.
3. The three-dimensional human body reconstruction method based on implicit neural network and diffusion model as claimed in claim 1, characterized in that: In step 2, view description embedding is added, including: After the density volume representation, the density features are mixed with a Gaussian noise vector ε~N(0, 1) in T steps, using a 2D mask to improve the guidance of generative modeling. At each time step t, a noise feature is generated. This noise signature integrates the The forward process is formulated as follows: in is the noise coefficient, d is the viewing angle parameter of the image; The overall loss is optimized by the reverse autoencoder. While fixing the pre-trained U-Net parameters, the NeRF and Low-Rank Adaptation parameters are updated to identify the Pareto frontier s in the feasible solution space S, where the target space R is expressed as: Where y refers to the guidance text embedding, the global optimal solution Including the balance coefficient w and the prediction noise U-Net∈; Diffusion distribution of score distillation samples at time t tends to the entire marginal distribution p from the pre-trained text to image model; We use the parameterized fraction to model the Wasserstein gradient flow and train an optimization procedure end-to-end that simultaneously optimizes the 3D parameters θ d and LoRA parameter φ; where 3D parameter θ d For pre-trained models and The Wasserstein gradient flow between φ and Gaussian distributed noise ∈.
4. The three-dimensional human body reconstruction method based on implicit neural network and diffusion model as claimed in claim 1, characterized in that: In step 3, a multi-scale structural similarity index is added, including: The core formula of the multi-scale structural similarity constraint is as follows: Among them, MSSC is obtained from the rendered image I and the ground truth A random patch P extracted from the image (n) Calculated using αM, β j and γ j To adjust the relative importance of different components; the brightness component l, the contrast component c and the structure component s are calculated using a Gaussian weighting formula within a 5 × 5 pixel filter.
5. The three-dimensional human body reconstruction method based on implicit neural network and diffusion model as claimed in claim 1, characterized in that: In the loss calculation, the following combination of loss functions is used to achieve high-quality 3D human body rendering: Rendering loss: We use the image-scale L2 loss to ensure the overall consistency between the lighting color rendering and the texture-free rendering, given the ground truth target image C(r) and the predicted image The rendering loss contributes to the identity feature distribution by progressively denoising the normally distributed variables in the backward process: in Represents a set of 3D query points on a ray; MSSC loss: By reconstructing SSIM, it is used to detect the consistency between the identity space and the appearance latent space. For the convenience of numerical calculation, it is normalized to the form of 1-MSSC: LPIPS loss: In order to solve the problem of inconsistent details caused by partial reflection or shadow, the spatial consistency and seamless transition characteristics in human perception are captured through a convolutional neural network, and the VGG-based pre-trained network prior is used to accelerate the convergence of the model: The overall loss is expressed as: Among them, λ1=0.1, λ2=0.1.
Citation Information
Cited By
Single-view 3D scene reconstruction method and device based on diffusion model, and medium
CN120259571A
A single-view 3D scene reconstruction method, device and medium based on diffusion model
CN120259571B
Multi-view target sample and scene image generation method
CN120259816A
A multi-view target sample and scene image generation method
CN120259816B
Method for generating printable 3D model by single photo based on generative adversarial network
CN120298208A