Three-dimensional model generation method and system based on fractional distillation sampling
By aligning the original image with the tetrahedral sphere mesh, combining fractional distillation sampling and noise residual optimization of the three-dimensional model, the problems of local distortion and deformation are solved, and a high-fidelity and topologically effective three-dimensional structure is generated.
Patent Information
- Application Number
- CN202510398600.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing three-dimensional model generation methods are prone to problems of local distortion and uncontrollable deformation.
By acquiring the original image and aligning the preset tetrahedral sphere mesh, combining fractional distillation sampling algorithm and noise residual, the gradient update direction is iteratively calculated, and using the micro-renderable algorithm and Laplace matrix constraints, the vertex position of the three-dimensional model is optimized until the preset iteration number is reached.
The generated three-dimensional model satisfies the pixel-level alignment and topological stability of the rendered image and the original image, avoiding local details loss or excessive smoothing, and improving generation efficiency and accuracy.
Smart Images

Figure CN120339545A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer graphics and 3D reconstruction, and relates to a 3D model generation method and system based on fractional diffusion sampling. Background Art
[0002] The generation of 3D models is an important research direction in computer vision and computer graphics, and is widely used in fields such as virtual reality, augmented reality, and game production.
[0003] Traditional 3D model generation methods mostly rely on manual design or depth map-based reconstruction techniques. Usually, they only optimize the model depending on the pixel differences between the rendered image and the original image, lacking direct constraints on geometric deformation, resulting in problems such as local distortion of the model and instability in the optimization process. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present application provides a 3D model generation method and system based on fractional diffusion sampling, which solves the technical problems of local distortion and uncontrollable deformation that easily occur in the prior art when generating 3D models.
[0005] To achieve the above object, in a first aspect, the present invention provides a 3D model generation method based on fractional diffusion sampling, including:
[0006] Obtain an original image, and align the original image with a preset tetrahedral sphere grid to obtain a first model;
[0007] Iteratively calculate a gradient update direction according to a preset fractional diffusion sampling algorithm and a noise residual, and adjust a second model according to the gradient update direction to obtain a third model until a preset number of iterations is reached, and output the current third model as the final model; and after each third model is obtained, update the current first model to the current third model;
[0008] Wherein, the noise residual is obtained by predicting and calculating the noise of a first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
[0009] Compared with the prior art, the embodiments of the present application have the following beneficial effects: By aligning the original image with the tetrahedral sphere grid to generate an initial model, and using the image contour information to constrain the distribution of grid vertices, the overall geometric structure (such as object size and shape) of the first model is highly adapted to the target object in the input image, thereby avoiding the subsequent wrong optimization direction caused by the deviation of the initial model, and at the same time reducing the parameter search complexity required for constructing a three-dimensional model from scratch; Further, by combining the fractional distillation sampling algorithm with the noise residual to calculate the gradient update direction, the cross-modal semantic supervision ability (such as object edge continuity and texture rationality) of the frozen two-dimensional image generation model is transformed into the optimization power of three-dimensional geometric deformation. The noise residual is used to quantify the semantic difference between the rendered image and the target image, driving the vertex position of the model to adjust in the direction consistent with the real object structure, breaking through the problem of local detail loss or over-smoothing caused by the traditional method relying on a single pixel loss; Dynamically update the model parameters in each iteration, use the optimized third model as the starting point for the next round of optimization, and gradually approach the target three-dimensional structure through a cyclic adjustment mechanism, avoiding optimization stagnation or gradient disappearance caused by fixing the initial model, and at the same time realizing the adaptive transition of the deformation amount from global coarse-grained adjustment (such as overall scaling) to local fine-grained correction (such as surface curvature optimization); Finally, output the model after a preset number of iterations, and control the consumption of computing resources in the optimization process through the termination condition, ensuring that the generated three-dimensional model simultaneously satisfies the pixel-level alignment between the rendered image and the original image, the rationality of the cross-modal semantic supervision of the two-dimensional generation model, and the stability of the tetrahedral grid geometric constraint, so as to achieve a balance between efficiency and accuracy and generate a high-fidelity and topologically valid three-dimensional structure.
[0010] In some embodiments of the first aspect of the present application, the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm, including:
[0011] Perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection;
[0012] According to the geometric projection and the preset lighting parameters, calculate the differentiable rendering gradient and synthesize the rendered image to obtain the first rendered image.
[0013] Compared with the prior art, the above embodiments have the following beneficial effects: By performing differentiable rasterization on the vertex position parameters of the first model, the continuous differentiability of the three-dimensional mesh vertex displacement is retained in the geometric projection process, ensuring that the rendering gradient (such as the influence of light changes on vertex positions) can be accurately calculated through backpropagation, and avoiding the gradient truncation problem caused by discretization operations in traditional rasterization; Further, by combining geometric projection with preset lighting parameters to calculate the differentiable rendering gradient, the interaction effects between light and the model surface (such as diffuse reflection, specular reflection) in a real physical environment are simulated, making the generated first rendered image more conform to the real scene in terms of details such as shading transitions and shadow distributions, providing input data that conforms to physical laws for subsequent optimization, thereby improving the generation robustness of the three-dimensional model under complex lighting conditions.
[0014] In some embodiments of the first aspect of the present application, the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image, including:
[0015] Calculate the rendering loss value according to a preset rendering loss function and the pixel difference between the first rendered image and the original image;
[0016] Calculate the geometric smoothness constraint value according to the vertex position parameters of the first model and a preset Laplacian matrix;
[0017] Generate a differentiable optimization target according to the rendering loss value and the geometric smoothness constraint value;
[0018] Perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization target to obtain the second model.
[0019] Compared with the prior art, the above embodiments have the following beneficial effects: By calculating the rendering loss value according to the pixel difference between the first rendered image and the original image, the local detail differences (such as edge misalignment, color deviation) between the current rendering result of the model and the target image are directly quantified, driving the vertex positions to be adjusted in the direction of pixel-level alignment, ensuring that the surface geometry of the model strictly matches the visible part of the input image; Further, by calculating the geometric smoothness constraint value based on the Laplacian matrix, the displacement amounts of adjacent vertices are forced to satisfy local smoothness (such as continuous curvature), suppressing mesh tearing, unevenness, or non-physical deformations caused by isolated vertex mutations; Finally, by combining the rendering loss and the geometric smoothness constraint to generate a differentiable optimization target, the pixel accuracy and geometric rationality are balanced during the optimization process, avoiding overfitting (such as overfitting to noisy pixels) or model distortion (such as surface wrinkles) caused by a single optimization target, thereby generating a second model that not only conforms to the image content but also maintains topological stability.
[0020] In some embodiments of the first aspect of the present application, the noise residual is obtained by predicting and calculating the noise of the first rendered image injected with random noise according to a preset frozen two-dimensional image generation model, including:
[0021] Inject random noise into the first rendered image to generate a noisy image and the corresponding actual noise value;
[0022] Input the noisy image and the original image into the frozen two-dimensional image generation model, and output the predicted noise value;
[0023] Calculate the noise residual according to the actual noise value and the predicted noise value.
[0024] Compared with the prior art, the above embodiments have the following beneficial effects: By injecting random noise into the first rendered image to generate a noisy image, a multi-step noise perturbation scenario during the training of the diffusion model is simulated, enabling the noise prediction ability of the frozen two-dimensional image generation model (such as Stable Diffusion) to adapt to the three-dimensional optimization requirements; further, by calculating the residual between the actual noise value and the noise value predicted by the model, the discrimination results of the teacher model for the regions in the rendered image that do not conform to the semantics of real objects (such as broken edges and unreasonable textures) are extracted, and the residual value is used as the supervision signal for three-dimensional geometric optimization, thereby converting the implicit knowledge of the two-dimensional model about the rationality of the image structure (such as object symmetry and continuity) into the driving force for correcting the vertex positions, enhancing the model's generation ability for complex details.
[0025] In some embodiments of the first aspect of the present application, the iterative calculation of the gradient update direction according to the preset score distillation sampling algorithm and the noise residual includes:
[0026] According to the preset score distillation sampling algorithm, combine the noise residual and the differentiable rendering gradient to calculate and generate the gradient update direction, where the calculation method is as follows:
[0027]
[0028] Where φ represents the frozen two-dimensional image generation model, x represents the first rendered image, g(θ) is the generator function of differentiable rendering, E t,ε represents the expectation with respect to the time step t and the random noise ε, is the noise predicted by the frozen two-dimensional image generation model, z t is the image after adding the noise at the t-th step, y is the original image, t is the time step, θ is the parameter of the first model, w(t) is the weighting function based on the time step, represents the gradient of the generator function with respect to the parameter θ, and L SDS represents the loss function of the score distillation sampling algorithm.
[0029] Compared with the prior art, the above embodiments have the following beneficial effects: By defining the gradient update direction formula, explicitly correlating the noise residuals with the differentiable rendering gradients, binding the cross-modal semantic supervision of the frozen two-dimensional model and the three-dimensional geometric deformation mathematically, ensuring that the optimization direction satisfies the dual constraints of two-dimensional image rationality (such as edge coherence) and three-dimensional geometric differentiability (such as smooth vertex movement) simultaneously; further by introducing a time-step weighting function, dynamically adjusting the contribution weights of gradient updates under different noise intensities, making the optimization process focus on macroscopic structure alignment (such as the overall shape of the object) in the early stage and microscopic detail correction (such as surface texture) in the later stage, thereby improving the multi-scale consistency of the model generation results.
[0030] In some embodiments of the first aspect of the present application, adjusting the second model according to the gradient update direction to obtain the third model includes:
[0031] Adjusting the vertex positions of the second model according to the gradient update direction and a preset non-inversion constraint to obtain the third model;
[0032] Among them, the non-inversion constraint is expressed as follows: Among them, represents the deformation gradient matrix, det() represents the determinant, i represents the i-th tetrahedron in the tetrahedron balls of the model, and j represents the j-th tetrahedron ball in the model.
[0033] Compared with the prior art, the above embodiments have the following beneficial effects: By constraining the determinant of the deformation gradient matrix, forcing each tetrahedron to maintain a positive local volume during the deformation process (i.e., prohibiting volume collapse or inversion), fundamentally avoiding problems such as mesh self-intersection, penetration, or topological failure caused by excessive vertex displacement; further by applying this constraint when adjusting the vertex positions according to the gradient update direction, combining semantic-driven optimization (such as the gradient direction of fractional distillation sampling) with geometric hard rules (such as non-invertibility of tetrahedrons), ensuring that the model deformation approximates the target structure while conforming to physical laws, thereby generating a three-dimensional model that satisfies both image semantics and geometric validity.
[0034] In some embodiments of the first aspect of the present application, aligning the original image with a preset tetrahedron ball mesh to obtain the first model includes:
[0035] Extracting the contour information of each target object in the original image;
[0036] Adjusting the vertex positions corresponding to the tetrahedron ball mesh according to each contour information to obtain the first model.
[0037] Compared with the prior art, the above embodiments have the following beneficial effects: By extracting the original image contour information and adjusting the positions of the vertices of the tetrahedral sphere mesh, the vertex distribution of the mesh is directly initialized using the edge features (such as the outer contour and hole structure) of the target object in the image, so that the initial geometric shape (such as aspect ratio and curvature) of the first model highly matches the macroscopic structure of the object in the input image; Further, using the aligned vertex positions as the starting point for optimization reduces the amount of deformation required in subsequent iterations, avoiding redundant calculations caused by starting optimization from scratch, thereby significantly improving the overall generation efficiency.
[0038] In a second aspect, the present invention also provides a three-dimensional model generation system based on fractional distillation sampling, including: an alignment module and an optimization module;
[0039] Among them, the alignment module is used to obtain the original image and align the original image with a preset tetrahedral sphere mesh to obtain a first model;
[0040] The optimization module is used to iteratively calculate the gradient update direction according to a preset fractional distillation sampling algorithm and noise residuals, and adjust the second model according to the gradient update direction to obtain a third model until a preset number of iterations is reached, and output the current third model as the final model; and after each third model is obtained, update the current first model to the current third model;
[0041] Among them, the noise residuals are calculated by predicting the noise of the first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
[0042] Compared with the prior art, the above embodiments of the present application have the following beneficial effects: By aligning the original image with the tetrahedral sphere mesh to generate an initial model, and using the image contour information to constrain the distribution of mesh vertices, the overall geometric structure (such as object size and shape) of the first model is highly adapted to the target object in the input image, thus avoiding the subsequent optimization direction error caused by the deviation of the initial model and reducing the parameter search complexity required for building a three-dimensional model from scratch; Further, by combining the fractional distillation sampling algorithm with the noise residual to calculate the gradient update direction, the cross-modal semantic supervision ability (such as object edge continuity and texture rationality) of the frozen two-dimensional image generation model is transformed into the optimization power of three-dimensional geometric deformation. The noise residual is used to quantify the semantic difference between the rendered image and the target image, driving the vertex position of the model to adjust in the direction consistent with the real object structure, breaking through the problems of local detail loss or over-smoothing caused by the traditional method relying on a single pixel loss; Dynamically update the model parameters in each iteration, use the optimized third model as the starting point for the next round of optimization, and gradually approach the target three-dimensional structure through a cyclic adjustment mechanism, avoiding optimization stagnation or gradient disappearance caused by fixing the initial model, and at the same time realizing the adaptive transition of the deformation amount from global coarse-grained adjustment (such as overall scaling) to local fine-grained correction (such as surface curvature optimization); Finally, output the model after a preset number of iterations, control the computational resource consumption of the optimization process through the termination condition, and ensure that the generated three-dimensional model simultaneously satisfies the pixel-level alignment between the rendered image and the original image, the rationality of the cross-modal semantic supervision of the two-dimensional generation model, and the stability of the tetrahedral mesh geometric constraint, so as to achieve a balance between efficiency and accuracy and generate a high-fidelity and topologically valid three-dimensional structure.
[0043] In some embodiments of the second aspect of the present application, the optimization module includes a rasterization processing unit and a first rendering unit;
[0044] Among them, the rasterization processing unit is used to perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection;
[0045] The first rendering unit is used to calculate a differentiable rendering gradient and synthesize a rendered image according to the geometric projection and preset lighting parameters to obtain a first rendered image.
[0046] Compared with the prior art, the above embodiments have the following beneficial effects: By performing differentiable rasterization on the vertex position parameters of the first model, the continuous differentiability of the three-dimensional mesh vertex displacement is retained in the geometric projection process, ensuring that the rendering gradient (such as the influence of light changes on vertex positions) can be accurately calculated through backpropagation, and avoiding the gradient truncation problem caused by discretization operations in traditional rasterization; Further, by combining geometric projection with preset lighting parameters to calculate the differentiable rendering gradient, the interaction effects between light and the model surface (such as diffuse reflection, specular reflection) in a real physical environment are simulated, making the generated first rendering image more conform to the real scene in terms of details such as smooth shading transitions and shadow distributions, providing input data that conforms to physical laws for subsequent optimization, thereby enhancing the generation robustness of the three-dimensional model under complex lighting conditions.
[0047] In some embodiments of the second aspect of the present application, the optimization module further includes: a differentiable loss calculation unit, a smooth constraint calculation unit, an optimization target generation unit, and a differentiable optimization unit;
[0048] Among them, the differentiable loss calculation unit is used to calculate a rendering loss value according to a preset rendering loss function and the pixel difference between the first rendering image and the original image;
[0049] The smooth constraint calculation unit is used to calculate a geometric smooth constraint value according to the vertex position parameters of the first model and a preset Laplacian matrix;
[0050] The optimization target generation unit is used to generate a differentiable optimization target according to the rendering loss value and the geometric smooth constraint value;
[0051] The differentiable optimization unit is used to perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization target to obtain a second model.
[0052] Compared with the prior art, the above embodiments have the following beneficial effects: By calculating the rendering loss value according to the pixel difference between the first rendering image and the original image, the local detail differences (such as edge misalignment, color deviation) between the current rendering result of the model and the target image are directly quantified, driving the vertex positions to adjust in the direction of pixel-level alignment, ensuring that the surface geometry of the model strictly matches the visible part of the input image; Further, by calculating the geometric smooth constraint value based on the Laplacian matrix, the displacement amounts of adjacent vertices are forced to satisfy local smoothness (such as continuous curvature), suppressing mesh tearing, unevenness, or non-physical deformations caused by isolated vertex mutations; Finally, by combining the rendering loss and the geometric smooth constraint to generate a differentiable optimization target, the pixel accuracy and geometric rationality are balanced during the optimization process, avoiding overfitting (such as excessive fitting to noisy pixels) or model distortion (such as surface wrinkles) caused by a single optimization target, thereby generating a second model that not only conforms to the image content but also maintains topological stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 : A schematic flow chart of a three-dimensional model generation method based on fractional distillation sampling provided in some embodiments of the present invention.
[0054] Figure 2 : A schematic structural diagram of a three-dimensional model generation system based on fractional distillation sampling provided in some embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1:
[0057] Please refer to Figure 1 , a three-dimensional model generation method based on fractional distillation sampling provided in an embodiment of the present invention, including steps S1 to S2:
[0058] Step S1: Obtain an original image, and align the original image with a preset tetrahedral sphere mesh to obtain a first model.
[0059] Further, the alignment operation in step S1 can be implemented through the following preferred embodiments, including steps S11 - S12, specifically as follows:
[0060] S11: Extract the contour information of each target object in the original image.
[0061] S12: Adjust the positions of the corresponding vertices of the tetrahedral sphere mesh according to each contour information to obtain a first model.
[0062] In specific implementation, the vertices of the tetrahedral sphere mesh can be expressed as: where N is the number of vertices in the mesh, and R represents the set of real numbers.
[0063] In this preferred embodiment, steps S11 - S12 directly initialize the grid vertex distribution by extracting the contour information of the original image and adjusting the positions of the vertices of the tetrahedral sphere mesh, making the initial geometric shape (such as aspect ratio, curvature) of the first model highly match the macroscopic structure of the object in the input image; further, using the aligned vertex positions as the optimization starting point reduces the level of deformation required in subsequent iterations, avoiding redundant calculations caused by optimizing from scratch, thereby significantly improving the overall generation efficiency.
[0064] Step S2: Iteratively calculate and generate a gradient update direction according to a preset fractional distillation sampling algorithm and noise residuals, and adjust the second model according to the gradient update direction to obtain a third model until a preset number of iterations is reached, and output the current third model as the final model; and after each third model is obtained, update the current first model to the current third model;
[0065] Wherein, the noise residuals are obtained by predicting and calculating the noise of the first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
[0066] Further, the first rendered image in step S2 can be obtained through the following preferred implementation manner, including steps S21 - S22, specifically as follows:
[0067] S21: Perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection;
[0068] S22: Calculate a differentiable rendering gradient and synthesize a rendered image according to the geometric projection and preset lighting parameters to obtain the first rendered image.
[0069] In this preferred embodiment, steps S21 - S22 perform differentiable rasterization processing on the vertex position parameters of the first model, retain the continuous differentiability of the three-dimensional mesh vertex displacement in the geometric projection process, ensure that the rendering gradient (such as the influence of lighting changes on vertex positions) can be accurately calculated through backpropagation, and avoid the gradient truncation problem caused by discretization operations in traditional rasterization; further, by combining the geometric projection and preset lighting parameters to calculate the differentiable rendering gradient, simulate the interaction effect of lighting and the model surface (such as diffuse reflection, specular reflection) in a real physical environment, make the generated first rendered image more conform to the real scene in details such as light and dark transitions and shadow distributions, and provide input data that conforms to physical laws for subsequent optimization, thereby improving the generation robustness of the three-dimensional model under complex lighting conditions.
[0070] Further, the second model in step S2 can be obtained through the following preferred implementation manner, including steps S23 - S26, specifically as follows:
[0071] S23: Calculate a rendering loss value according to a preset rendering loss function and the pixel difference between the first rendered image and the original image;
[0072] S24: Calculate the geometric smoothing constraint value according to the vertex position parameters of the first model and the preset Laplacian matrix;
[0073] S25: Generate a differentiable optimization objective according to the rendering loss value and the geometric smoothing constraint value;
[0074] S26: Perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization objective to obtain a second model.
[0075] In specific implementation, the differentiable optimization objective can be expressed as: where R(X) represents the rendering loss function, Φ represents the rendering loss, L represents the Laplacian matrix, λ represents a hyperparameter for adjusting the deformation gradient norm, and F x represents the deformation gradient to ensure the smoothness of the mesh deformation.
[0076] In this preferred embodiment, steps S23 - S26 calculate the rendering loss value based on the pixel difference between the first rendered image and the original image, directly quantify the local detail differences (such as edge misalignment, color deviation) between the current rendering result of the model and the target image, drive the vertex positions to adjust in the direction of pixel-level alignment, and ensure that the surface geometry of the model strictly matches the visible part of the input image; further calculate the geometric smoothing constraint value based on the Laplacian matrix, force the displacement amounts of adjacent vertices to satisfy local smoothness (such as continuous curvature), and suppress mesh tearing, unevenness, or non-physical deformation caused by isolated vertex mutations; finally, generate a differentiable optimization objective by combining the rendering loss and the geometric smoothing constraint, balance pixel accuracy and geometric rationality during the optimization process, and avoid overfitting (such as excessive fitting to noisy pixels) or model distortion (such as surface wrinkles) caused by a single optimization objective, so as to generate a second model that not only conforms to the image content but also maintains topological stability.
[0077] Furthermore, the noise residual in step S2 can be implemented through the following preferred embodiments, including steps S27 - S29, specifically as follows:
[0078] S27: Inject random noise into the first rendered image to generate a noisy image and the corresponding actual noise value;
[0079] S28: Input the noisy image and the original image into a frozen two-dimensional image generation model, and output the predicted noise value;
[0080] S29: Calculate the noise residual according to the actual noise value and the predicted noise value.
[0081] In implementation, the frozen two-dimensional image generation model is an image generation model pre-trained on a large-scale two-dimensional image dataset using the diffusion method. These models contain geometric prior information of the object from multiple perspectives and act as a role similar to the teacher model in knowledge distillation in the fractional distillation sampling, used to supervise the optimization of the tetrahedral sphere grid. Usually, open-source pre-trained two-dimensional image generation models such as Stable Diffusion can be used as the supervisory role. The frozen two-dimensional image generation model is used to predict the noise, and the gradient update direction is calculated based on the difference between the predicted noise and the actually added noise. When the prediction error is smaller, the rendered image is more accurate, and the optimization of the tetrahedral sphere grid is better.
[0082] In this preferred embodiment, steps S27 - S29 generate a noisy image by injecting random noise into the first rendered image, simulating the multi-step noise perturbation scenario during the training of the diffusion model, so that the noise prediction ability of the frozen two-dimensional image generation model (such as Stable Diffusion) adapts to the three-dimensional optimization requirements; further, by calculating the residual between the actual noise value and the model-predicted noise value, the discrimination results of the teacher model for the regions in the rendered image that do not conform to the semantics of real objects (such as broken edges, unreasonable textures) are extracted, and the residual value is used as the supervisory signal for three-dimensional geometric optimization, thereby converting the implicit knowledge of the two-dimensional model about the rationality of the image structure (such as object symmetry, continuity) into the driving force for correcting the vertex positions, enhancing the model's ability to generate complex details.
[0083] Furthermore, step S2 can be implemented through the following preferred implementation when iteratively calculating the gradient update direction, specifically:
[0084] According to the preset fractional distillation sampling algorithm, combining the noise residual and the differentiable rendering gradient, calculate the generated gradient update direction, where the calculation method is as follows:
[0085]
[0086] Among them, φ represents the frozen two-dimensional image generation model, x represents the first rendered image, g(θ) is the generator function of differentiable rendering, E t,ε represents the expectation with respect to the time step t and the random noise ε, is the noise predicted by the frozen two-dimensional image generation model, z t is the image after adding the noise at the t-th step, y is the original image, t is the time step, θ is the parameter of the first model, w(t) is the weighting function based on the time step, represents the gradient of the generator function with respect to the parameter θ, and L SDS represents the loss function of the fractional distillation sampling algorithm.
[0087] In the prior art, a large amount of high-precision 3D annotation data is often required for generating 3D models. However, the present invention creatively introduces the fractional distillation sampling technique. By freezing a large-scale pre-trained 2D image generation model as a cross-modal teacher, the geometric prior knowledge learned by it in the 2D image space is transformed into a gradient signal for driving 3D geometric optimization through the bridge established by differentiable rendering, realizing the knowledge distillation from 2D visual knowledge to the 3D parameter space and significantly reducing the dependence on 3D annotation data on the premise of ensuring the generation quality.
[0088] In this preferred embodiment, step S2 explicitly associates the noise residual with the differentiable rendering gradient by defining a gradient update direction formula, mathematically binding the cross-modal semantic supervision of the frozen 2D model and the 3D geometric deformation, ensuring that the optimization direction simultaneously satisfies the dual constraints of 2D image rationality (such as edge coherence) and 3D geometric differentiability (such as smooth vertex movement); further, by introducing a time-step weighting function, the contribution weights of gradient updates under different noise intensities are dynamically adjusted, making the optimization process focus on macroscopic structure alignment (such as the overall shape of the object) in the early stage and microscopic detail correction (such as surface texture) in the later stage, thereby improving the multi-scale consistency of the model generation result.
[0089] Furthermore, when adjusting the second model in step S2, it can be achieved through the following preferred implementation manner, specifically:
[0090] Adjust the vertex positions of the second model according to the gradient update direction and a preset non-inversion constraint to obtain a third model; where the non-inversion constraint is expressed as follows: Where represents the deformation gradient matrix, det() represents the determinant, i represents the i-th tetrahedron in the tetrahedral sphere in the model, and j represents the j-th tetrahedral sphere in the model.
[0091] In this preferred embodiment, step S2 enforces the determinant of the deformation gradient matrix to ensure that each tetrahedron maintains a positive local volume during the deformation process (i.e., prohibits volume collapse or inversion), fundamentally avoiding problems such as mesh self-intersection, penetration, or topological failure caused by excessive vertex displacement; further, by applying this constraint when adjusting the vertex positions according to the gradient update direction, the semantic-driven optimization (such as the gradient direction of fractional distillation sampling) is combined with the geometric hard rule (such as tetrahedra being non-invertible), ensuring that the model deformation approximates the target structure on the premise of conforming to physical laws, thereby generating a 3D model that satisfies both image semantics and geometric validity.
[0092] In summary, compared with the prior art, the above embodiments of the present application have the following beneficial effects: By aligning the original image with the tetrahedral sphere mesh to generate an initial model and using the image contour information to constrain the distribution of mesh vertices, the overall geometric structure (such as object size and shape) of the first model is highly adapted to the target object in the input image, thus avoiding incorrect subsequent optimization directions caused by deviations in the initial model and reducing the parameter search complexity required to build a three-dimensional model from scratch; further, by combining the fractional distillation sampling algorithm with noise residuals to calculate the gradient update direction, the cross-modal semantic supervision ability (such as object edge continuity and texture rationality) of the frozen two-dimensional image generation model is transformed into the optimization driving force for three-dimensional geometric deformation. The noise residuals are used to quantify the semantic differences between the rendered image and the target image, driving the adjustment of the model vertex positions in the direction that conforms to the real object structure, breaking through the problems of local detail loss or excessive smoothing caused by traditional methods relying on single-pixel loss; dynamically updating the model parameters in each iteration, using the optimized third model as the starting point for the next round of optimization, and gradually approaching the target three-dimensional structure through a cyclic adjustment mechanism, avoiding optimization stagnation or gradient disappearance caused by fixing the initial model, and at the same time realizing the adaptive transition of the deformation amount from global coarse-grained adjustment (such as overall scaling) to local fine-grained correction (such as surface curvature optimization); finally, outputting the model after a preset number of iterations, controlling the consumption of computing resources in the optimization process through termination conditions, ensuring that the generated three-dimensional model simultaneously satisfies the pixel-level alignment between the rendered image and the original image, the rationality of cross-modal semantic supervision of the two-dimensional generation model, and the stability of tetrahedral mesh geometric constraints, thereby achieving a balance between efficiency and accuracy and generating a high-fidelity and topologically valid three-dimensional structure.
[0093] Embodiment 2:
[0094] Please refer to Figure 2 , based on the same inventive concept, a three-dimensional model generation system disclosed in an embodiment of the present invention includes: an alignment module M1 and an optimization module M2;
[0095] Among them, the alignment module M1 is used to obtain the original image and align the original image with a preset tetrahedral sphere mesh to obtain a first model.
[0096] Further, the alignment module M1 includes: a contour extraction unit and an alignment adjustment unit;
[0097] Among them, the contour extraction unit is used to extract the contour information of each target object in the original image;
[0098] The alignment adjustment unit is used to adjust the positions of the corresponding vertices of the tetrahedral sphere mesh according to each piece of contour information to obtain a first model.
[0099] In this preferred embodiment, the contour extraction unit and the alignment adjustment unit directly initialize the grid vertex distribution by extracting the contour information of the original image and adjusting the vertex positions of the tetrahedral sphere grid, using the edge features of the target object in the image (such as the outer contour, hole structure), so that the initial geometric shape (such as aspect ratio, curvature) of the first model highly matches the macroscopic structure of the object in the input image; further, taking the aligned vertex positions as the optimization starting point, reducing the level of deformation required in subsequent iterations, and avoiding redundant calculations caused by starting optimization from scratch, thereby significantly improving the overall generation efficiency.
[0100] The optimization module M2 is used to iteratively calculate the gradient update direction according to the preset fractional distillation sampling algorithm and noise residual, and adjust the second model according to the gradient update direction to obtain the third model until the preset number of iterations is reached, and output the current third model as the final model; and after obtaining the third model each time, update the current first model to the current third model;
[0101] Among them, the noise residual is calculated by predicting the noise of the first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
[0102] Furthermore, the optimization module M2 includes: a rasterization processing unit and a first rendering unit;
[0103] Among them, the rasterization processing unit is used to perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection;
[0104] The first rendering unit is used to calculate the differentiable rendering gradient and synthesize the rendered image according to the geometric projection and the preset lighting parameters to obtain the first rendered image.
[0105] In this preferred embodiment, the optimization module M2 preserves the continuous differentiability of the three-dimensional grid vertex displacement in the geometric projection process by performing differentiable rasterization processing on the vertex position parameters of the first model, ensuring that the rendering gradient (such as the influence of lighting changes on vertex positions) can be accurately calculated through backpropagation, and avoiding the gradient truncation problem caused by discrete operations in traditional rasterization; further, by combining the geometric projection and the preset lighting parameters to calculate the differentiable rendering gradient, simulating the interaction effect between lighting and the model surface (such as diffuse reflection, specular reflection) in a real physical environment, making the generated first rendered image more conform to the real scene in details such as shading transition and shadow distribution, providing input data that conforms to physical laws for subsequent optimization, thereby improving the generation robustness of the three-dimensional model under complex lighting conditions.
[0106] Furthermore, the optimization module M2 further includes: a differentiable loss calculation unit, a smoothness constraint calculation unit, an optimization objective generation unit, and a differentiable optimization unit;
[0107] Among them, the differentiable loss calculation unit is used to calculate a rendering loss value according to a preset rendering loss function and the pixel difference between the first rendered image and the original image;
[0108] The smoothness constraint calculation unit is used to calculate a geometric smoothness constraint value according to the vertex position parameters of the first model and a preset Laplacian matrix;
[0109] The optimization objective generation unit is used to generate a differentiable optimization objective according to the rendering loss value and the geometric smoothness constraint value;
[0110] The differentiable optimization unit is used to perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization objective to obtain a second model.
[0111] In this preferred embodiment, the optimization module M2 directly quantifies the local detail differences (such as edge misalignment and color deviation) between the current rendering result of the model and the target image by calculating the rendering loss value according to the pixel difference between the first rendered image and the original image, drives the adjustment of the vertex positions in the direction of pixel-level alignment to ensure that the surface geometry of the model strictly matches the visible part of the input image; further, by calculating the geometric smoothness constraint value based on the Laplacian matrix, it enforces the displacement amounts of adjacent vertices to satisfy local smoothness (such as continuous curvature), suppresses mesh tearing, unevenness, or non-physical deformation caused by isolated vertex mutations; finally, by combining the rendering loss and the geometric smoothness constraint to generate a differentiable optimization objective, it balances pixel accuracy and geometric rationality during the optimization process, avoiding overfitting (such as excessive fitting to noisy pixels) or model distortion (such as surface wrinkles) caused by a single optimization objective, thereby generating a second model that both conforms to the image content and maintains topological stability.
[0112] Furthermore, the optimization module M2 further includes: a noise addition unit, a noise prediction unit, and a residual calculation unit;
[0113] Among them, the noise addition unit is used to inject random noise into the first rendered image to generate a noisy image and the corresponding actual noise value;
[0114] The noise prediction unit is used to input the noisy image and the original image into a frozen two-dimensional image generation model and output a predicted noise value;
[0115] The residual calculation unit is used to calculate a noise residual according to the actual noise value and the predicted noise value.
[0116] In this preferred embodiment, the optimization module M2 generates a noisy image by injecting random noise into the first rendered image, simulating the multi-step noise perturbation scenario during the training of the diffusion model, so that the noise prediction ability of the frozen two-dimensional image generation model (such as StableDiffusion) adapts to the three-dimensional optimization requirements; further, by calculating the residual between the actual noise value and the model-predicted noise value, the discrimination result of the teacher model for the regions in the rendered image that do not conform to the semantics of real objects (such as broken edges, unreasonable textures) is extracted, and the residual value is used as the supervision signal for three-dimensional geometric optimization, thereby converting the implicit knowledge of the two-dimensional model about the rationality of the image structure (such as object symmetry, continuity) into the driving force for vertex position correction, enhancing the model's generation ability for complex details.
[0117] Furthermore, the optimization module M2 can be implemented by the following preferred implementation when calculating the generation gradient update direction, specifically:
[0118] According to the preset score distillation sampling algorithm, combining the noise residual and the differentiable rendering gradient, calculate the generation gradient update direction, and the calculation method is as follows:
[0119]
[0120] Among them, φ represents the frozen two-dimensional image generation model, x represents the first rendered image, g(θ) is the generator function of differentiable rendering, E t,ε represents the expectation of the time step t and the random noise ε, is the noise predicted by the frozen two-dimensional image generation model, z t is the image after adding the noise at the t-th step, y is the original image, t is the time step, θ is the parameter of the first model, w(t) is the weighting function based on the time step, represents the gradient of the generator function with respect to the parameter θ, and L SDS represents the loss function of the score distillation sampling algorithm.
[0121] In this preferred embodiment, the optimization module M2 explicitly associates the noise residual with the differentiable rendering gradient by defining the gradient update direction formula, mathematically binding the cross-modal semantic supervision of the frozen two-dimensional model and the three-dimensional geometric deformation, ensuring that the optimization direction simultaneously satisfies the dual constraints of two-dimensional image rationality (such as edge coherence) and three-dimensional geometric differentiability (such as smooth vertex movement); further, by introducing a time step weighting function, dynamically adjusting the contribution weights of gradient updates under different noise intensities, making the optimization process focus on macroscopic structure alignment (such as the overall shape of the object) in the early stage and microscopic detail correction (such as surface texture) in the later stage, thereby improving the multi-scale consistency of the model generation results.
[0122] Furthermore, when adjusting the second model, the optimization module M2 can be implemented through the following preferred embodiments, specifically:
[0123] According to the gradient update direction and a preset non-inversion constraint, adjust the vertex positions of the second model to obtain a third model; wherein, the non-inversion constraint is expressed as follows: Wherein, represents the deformation gradient matrix, det() represents the determinant, i represents the i-th tetrahedron in the tetrahedral sphere in the model, and j represents the j-th tetrahedral sphere in the model.
[0124] In this preferred embodiment, the optimization module M2 enforces the determinant of the deformation gradient matrix to force each tetrahedron to maintain a positive local volume during the deformation process (i.e., prohibits volume collapse or inversion), fundamentally avoiding problems such as mesh self-intersection, penetration, or topological failure caused by excessive vertex displacement; further, by applying this constraint when adjusting the vertex positions according to the gradient update direction, combining semantic-driven optimization (such as the gradient direction of fractional distillation sampling) with geometric hard rules (such as tetrahedra being non-invertible), ensuring that the model deformation approximates the target structure while conforming to physical laws, thereby generating a three-dimensional model that satisfies both image semantics and geometric validity.
[0125] In summary, compared with the prior art, the embodiments of the present application have the following beneficial effects: By aligning the original image with the tetrahedral sphere grid to generate an initial model and using the image contour information to constrain the distribution of grid vertices, the overall geometric structure (such as object size and shape) of the first model is highly adapted to the target object in the input image, thus avoiding the subsequent wrong optimization direction caused by the deviation of the initial model and reducing the parameter search complexity required for constructing a three-dimensional model from scratch; Further, by combining the fractional distillation sampling algorithm with the noise residual to calculate the gradient update direction, the cross-modal semantic supervision ability (such as object edge continuity and texture rationality) of the frozen two-dimensional image generation model is transformed into the optimization power of three-dimensional geometric deformation. The semantic difference between the rendered image and the target image is quantified by the noise residual, driving the adjustment of the model vertex position in the direction consistent with the real object structure, breaking through the problems of local detail loss or over-smoothing caused by the traditional method relying on a single pixel loss; Dynamically update the model parameters in each iteration, use the optimized third model as the starting point for the next round of optimization, and gradually approach the target three-dimensional structure through a cyclic adjustment mechanism, avoiding optimization stagnation or gradient disappearance caused by fixing the initial model, and realizing the adaptive transition of the deformation amount from global coarse-grained adjustment (such as overall scaling) to local fine-grained correction (such as surface curvature optimization); Finally, output the model after a preset number of iterations, control the consumption of computing resources in the optimization process through termination conditions, ensure that the generated three-dimensional model simultaneously meets the pixel-level alignment of the rendered image and the original image, the rationality of the cross-modal semantic supervision of the two-dimensional generation model, and the stability of the tetrahedral grid geometric constraint, so as to achieve a balance between efficiency and accuracy and generate a high-fidelity and topologically valid three-dimensional structure.
[0126] For the specific working processes of the above-described modules, reference may be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein. The division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system.
[0127] The above-described specific embodiments have further elaborated the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A three-dimensional model generation method based on fractional distillation sampling, characterized in that, including: obtain an original image, and align the original image with a preset tetrahedral sphere grid to obtain a first model; iteratively calculate a gradient update direction according to a preset fractional distillation sampling algorithm and a noise residual, and adjust a second model according to the gradient update direction to obtain a third model until a preset number of iterations is reached, and output the current third model as the final model; and after each third model is obtained, update the current first model to the current third model; wherein, the noise residual is calculated by predicting and calculating the noise of a first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
2. The three-dimensional model generation method based on fractional distillation sampling according to claim 1, characterized in that, The first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm, including: perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection; calculate a differentiable rendering gradient and synthesize a rendered image according to the geometric projection and preset lighting parameters to obtain a first rendered image.
3. The three-dimensional model generation method based on fractional distillation sampling according to claim 1, characterized in that, The second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image, including: calculate a rendering loss value according to a preset rendering loss function and the pixel difference between the first rendered image and the original image; calculate a geometric smoothing constraint value according to the vertex position parameters of the first model and a preset Laplacian matrix; generate a differentiable optimization target according to the rendering loss value and the geometric smoothing constraint value; perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization target to obtain a second model.
4. The 3D model generation method based on fractional distillation sampling according to claim 3, characterized in that The noise residual is calculated by predicting and calculating the noise of a first rendered image injected with random noise according to a preset frozen two-dimensional image generation model, including: inject random noise into the first rendered image to generate a noisy image and a corresponding actual noise value; input the noisy image and the original image into the frozen two-dimensional image generation model, and output a predicted noise value; calculate a noise residual according to the actual noise value and the predicted noise value.
5. A three-dimensional model generation method based on fractional distillation sampling as described in claim 1, characterized in that, The iteratively calculating a gradient update direction according to a preset fractional distillation sampling algorithm and a noise residual includes: calculate a gradient update direction according to a preset fractional distillation sampling algorithm, in combination with the noise residual and the differentiable rendering gradient, and the calculation method is as follows: Among them, φ represents the frozen two-dimensional image generation model, x represents the first rendered image, g(θ) is the generator function of differentiable rendering, E t,ε represents the expectation over the time step t and the random noise ε, is the noise predicted by the frozen two-dimensional image generation model, z t is the image after adding the noise at the t-th step, y is the original image, t is the time step, θ is the parameter of the first model, and w(t) is the weighting function based on the time step, represents the gradient of the generator function with respect to the parameter θ, and L SDS represents the loss function of the fractional distillation sampling algorithm.
6. The 3D model generation method based on fractional distillation sampling according to claim 5, characterized in that, The adjusting the second model according to the gradient update direction to obtain a third model includes: adjust the vertex positions of the second model according to the gradient update direction and a preset non-inversion constraint to obtain a third model; Among them, the non-inversion constraint is expressed as follows: Among them, represents the deformation gradient matrix, det() represents the determinant, i represents the i-th tetrahedron in the tetrahedron balls in the model, and j represents the j-th tetrahedron ball in the model.
7. A three-dimensional model generation method based on fractional distillation sampling according to any one of claims 1-6, characterized in that, The aligning the original image with a preset tetrahedral sphere grid to obtain a first model includes: extract the contour information of each target object in the original image; adjust the vertex positions corresponding to the tetrahedral sphere grid according to each contour information to obtain a first model.
8. A three-dimensional model generation system based on fractional distillation sampling, characterized in that including: an alignment module and an optimization module; Among them, the alignment module is used to obtain the original image and align the original image with a preset tetrahedral sphere grid to obtain a first model; The optimization module is used to iteratively calculate and generate a gradient update direction according to a preset fractional distillation sampling algorithm and noise residual, and adjust the second model according to the gradient update direction to obtain a third model until a preset number of iterations is reached, and output the current third model as the final model; and after each third model is obtained, update the current first model to the current third model; Among them, the noise residual is calculated by predicting the noise of a first rendered image injected with random noise according to a preset frozen two-dimensional image generation model; the first rendered image is obtained by rendering the current first model according to a preset differentiable rendering algorithm; the second model is obtained by adjusting the current first model according to the rendering loss between the first rendered image and the original image.
9. The three-dimensional model generation system based on fractional distillation sampling according to claim 8, wherein, The optimization module includes a rasterization processing unit and a first rendering unit; Among them, the rasterization processing unit is used to perform differentiable rasterization processing on the vertex position parameters of the first model to generate a geometric projection; The first rendering unit is used to calculate a differentiable rendering gradient and synthesize a rendered image according to the geometric projection and preset lighting parameters to obtain a first rendered image.
10. The three-dimensional model generation system based on fractional distillation sampling according to claim 8, characterized in that, The optimization module further includes: a differentiable loss calculation unit, a smoothing constraint calculation unit, an optimization target generation unit, and a differentiable optimization unit; Among them, the differentiable loss calculation unit is used to calculate a rendering loss value according to a preset rendering loss function and the pixel difference between the first rendered image and the original image; The smoothing constraint calculation unit is used to calculate a geometric smoothing constraint value according to the vertex position parameters of the first model and a preset Laplacian matrix; The optimization target generation unit is used to generate a differentiable optimization target according to the rendering loss value and the geometric smoothing constraint value; The differentiable optimization unit is used to perform gradient descent optimization on the vertex positions of the first model according to the differentiable optimization target to obtain a second model.
Citation Information
Cited By
Three-dimensional grid model protection method and system based on hybrid encryption and geometric projection
CN121966833A
Method and system for protecting three-dimensional mesh model based on hybrid encryption and geometric projection
CN121966833B