A method for generating printable 3D models from a single photo based on generative adversarial networks

By combining the generative adversarial network with implicit neural representation, the problem of generating high-precision 3D models in a single image is solved, and high-precision reconstruction of global structure and local details is realized, noise interference is reduced, and the adaptability of 3D printing is improved.

CN120298208BActive Publication Date: 2025-08-19HANGZHOU SECOND LIFE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510780326.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-19
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

It is difficult for the prior art to generate high-precision, rich in detail and suitable for 3D printing from a single image. The traditional method has problems such as insufficient global structure and local details, noise interference and poor printing adaptability.

Method used

Using a method of combining generative adversarial networks with implicit neural representations, a high-precision 3D model is achieved from a single photo through multi-scale refinement and multiple technical optimizations, including conditional diffusion inversion, discrete wavelet transformation and differentiable rendering feedback correction.

Benefits of technology

The generated 3D model achieves high accuracy in global structure and local details, reduces noise interference, improves the robustness and print adaptability of the model, and meets the needs of industrial design and personalized customization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298208B_ABST
    Figure CN120298208B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a printable 3D model from a single photo based on a generative adversarial network, comprising the following steps: S1, obtaining a single input image and preprocessing the input image; S2, inputting the preprocessed image into an encoder, extracting multi-scale features, and outputting a deep feature representation; S3, generating a 3D model using a conditional generative adversarial network and optimizing the 3D model quality through comparison; S4, refining the 3D model at multiple scales using an implicit neural representation network and converting it into a continuous function representation; S5, optimizing the 3D model using an improved conditional diffusion inversion module, combining multi-scale discrete wavelet transform and rendering feedback correction; S6, optimizing the feature extraction, 3D model generation, and refinement processes using an end-to-end joint training system; and S7, post-processing the 3D model to output a 3D model that meets printing requirements. By integrating generative adversarial networks and implicit neural representation technology, the present invention achieves the rapid generation of high-precision 3D models with rich details suitable for 3D printing from a single photo.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating a printable 3D model from a single photo based on a generative adversarial network. Background Art

[0002] With the rapid development of computer vision and deep learning technologies, 3D model reconstruction technology based on deep neural networks has become an important research direction in computer vision and graphics. Traditional 3D reconstruction methods rely primarily on multi-view imaging or multiple image data collected by expensive equipment such as laser scanning, and achieve 3D reconstruction through methods such as structured light, stereo vision, or photometric consistency. While these methods have achieved certain results in accuracy and detail restoration, they are subject to numerous limitations in data acquisition, equipment cost, and processing speed, hindering their widespread application in consumer products and everyday scenarios.

[0003] In recent years, generative adversarial networks (GANs), as an emerging generative model, have been widely applied to tasks such as image generation, image conversion, and data augmentation. Generative adversarial networks use an adversarial training mechanism between a generator and a discriminator to approximate the distribution of real data, thereby generating high-quality image data. In particular, the emergence of conditional generative adversarial networks (GANs), by introducing conditional inputs, allows generative models to have greater flexibility and control over the content and style of generated results. While these technologies have achieved significant progress in the fields of image generation and conversion, direct application to 3D model generation remains challenging.

[0004] In the existing technology, the generation of 3D models using GAN mainly relies on multi-view data or depth sensor information. However, the 3D information contained in a single image is often implicit and incomplete. Due to the lack of sufficient stereo information, the traditional single-image 3D reconstruction method generates 3D models with obvious deficiencies in global structure and local details, making it difficult to meet the accuracy requirements of actual 3D printing. Although some studies have attempted to combine generative adversarial networks and implicit neural representation methods to compensate for the problem of insufficient 3D model data generated from a single image by converting discrete 3D model data into continuous function representations, these methods still have deficiencies in stability, detail recovery, and printing adaptability. The 3D models generated by some methods have noise on the surface, insufficient recovery of local structural details, and due to the lack of global topology optimization methods, the generated 3D models are prone to problems such as structural instability and printing failure during the actual printing process.

[0005] In addition, the conditional generative adversarial networks used in existing technologies, when processing the task of generating 3D models from a single image, usually only focus on global shape information and ignore local details and structural consistency, resulting in the generated model having blurred textures and missing details in local areas. At the same time, although implicit neural representation methods can represent 3D models as continuous signed distance functions and achieve surface characterization with infinite levels of detail, they are limited in their effectiveness in processing high-frequency details and noise correction, and often need to be combined with other technologies to further improve the accuracy of model refinement and optimization. Existing methods lack adaptive control means for noise correction and local structural adjustment, making it difficult to effectively retain the fine geometric information implicit in the original image during the multi-scale refinement process.

[0006] In addition, although some studies have attempted to introduce advanced technologies such as diffusion inversion, discrete wavelet transform, and graph convolutional networks to improve the refinement quality of 3D models, these methods are often used in isolation and lack an end-to-end optimization mechanism that organically combines multiple technologies. Due to the complex complementary relationship between the balance between global structural information and local details, noise correction and detail enhancement in the 3D model generation and refinement process, traditional methods cannot simultaneously meet both requirements, resulting in a large gap in the visual effects and printing adaptability of the generated 3D models. Especially in the scenario of single image input, due to data scarcity and information loss, existing technologies cannot guarantee the accuracy and robustness of the generated model and the feasibility of high-quality 3D printing.

[0007] Therefore, how to provide a method for generating a printable 3D model from a single photo based on a generative adversarial network is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0008] One objective of the present invention is to propose a method for generating a printable 3D model from a single photo based on a generative adversarial network. This method leverages several advanced technologies, including deep learning, generative adversarial networks, implicit neural representations, diffusion inversion, discrete wavelet transforms, cross-domain fusion, and differentiable rendering feedback correction, to describe in detail a comprehensive approach for automatically generating high-precision, detail-rich, and 3D-printable 3D models from a single input image. Through this method, the implicit 3D information in a single image is fully exploited and expressed, resulting in a 3D model that achieves high reconstruction accuracy in both global structure and local detail, effectively compensating for the reliance of traditional 3D reconstruction techniques on multi-view data. Furthermore, the method utilizes improved conditional diffusion inversion and multi-scale refinement techniques to effectively reduce noise interference and enhance the consistency of the model's local structure. Combined with cross-domain fusion and a differentiable rendering feedback correction module, the method further enhances the 3D model's detail representation and printability. This method offers advantages such as convenient data acquisition, low generation cost, high model accuracy, rich detail, and robustness, providing an efficient and intelligent solution for 3D printing applications in industrial design, manufacturing, and personalized customization.

[0009] According to an embodiment of the present invention, a method for generating a printable 3D model from a single photo based on a generative adversarial network includes the following steps:

[0010] S1. Obtain a single input image of the target object and preprocess the input image;

[0011] S2. Input the preprocessed image into the encoder, extract the multi-scale features of the preprocessed image, and output the deep feature representation;

[0012] S3. The deep feature representation is input into the generative adversarial network. The generator generates a 3D model based on the deep feature representation. The discriminator compares the generated model with the real model and provides feedback to guide the generator to optimize the quality of the 3D model.

[0013] S4. Inputting the 3D model into the implicit neural representation network, using the implicit representation method of the signed distance function to convert the 3D model data into a continuous function representation, and performing multi-scale refinement processing on the 3D model through the implicit neural representation network;

[0014] S5. Input the refined 3D model into the improved conditional diffusion inversion module, which combines multi-scale discrete wavelet transform with differentiable rendering feedback correction technology to optimize the global structure and local details of the 3D model;

[0015] S6. Build an end-to-end joint training system to optimize feature extraction, 3D model generation, and refinement processes;

[0016] S7. Perform mesh smoothing, topology correction, and format conversion on the optimized 3D model to generate a file format that meets the 3D printer standard, perform printing path planning, and finally output a 3D model that meets the printing requirements.

[0017] Optionally, the S3 specifically includes:

[0018] S31, the output deep feature representation O and the adaptive noise vector z are fused after linear transformation, and the fusion condition vector C is calculated:

[0019] ;

[0020] in, is the learnable transformation matrix of the noise vector, represents element-wise multiplication, is the smoothing parameter, and are exponential and logarithmic functions respectively;

[0021] S32, input the fusion condition vector C into the hierarchical decoding network, and use multi-scale upsampling and residual connection modules to generate a 3D model :

[0022] ;

[0023] Among them, L is the number of layers of residual modules in the decoding network, is the upsampling convolution kernel of the i-th layer, Represents the processing output of the fusion condition vector by the i-th layer residual module, is the fusion coefficient in the residual connection of the i-th layer, Represents the convolution operation;

[0024] S33, the generated 3D model Perform adaptive normalization processing to obtain a normalized model :

[0025] ;

[0026] in, is the mean value of the 3D model, is the variance, To prevent division by zero for small positive numbers, Represents the result of local area average pooling of the 3D model. and To balance the contribution of each part, represents the Frobenius norm;

[0027] S34, normalize the model Input the discriminator separately with the real 3D model and calculate the adversarial loss based on the mixed gradient difference The discriminator captures the differences between the normalized model and the real 3D model in local texture, edge and structural gradient information, and feeds the differences back to the generator to guide the generator to continuously improve the quality of the 3D model:

[0028] ;

[0029] in, and is the balance coefficient, is a real 3D model, tanh() is the hyperbolic tangent function, is the spatial gradient operator, It means taking the sum of the absolute values of the elements;

[0030] S35. Adopt an alternating optimization strategy to jointly update the parameters of the generator and discriminator and output a 3D model.

[0031] Optionally, the S4 specifically includes:

[0032] S41. From the generated 3D model In the , the three-dimensional point set P is extracted by uniform sampling method:

[0033] ;

[0034] in, represents the number of sampling points, represents the three-dimensional real space, Represents the i-th sampling point, where i is the index of the sampling point;

[0035] S42. Initialize the implicit neural representation network , the implicit neural representation network accepts three-dimensional coordinates as input and outputs the corresponding signed distance s;

[0036] S43, for each sampling point , using implicit networks to calculate the predicted signed distance ;

[0037] S44. Define multi-scale refinement loss function is the weighted average of the error between the predicted symbol distance and the target distance at each scale:

[0038] ;

[0039] Among them, S represents the number of scales, is the weight coefficient of the s-th scale, Representation based on preliminary 3D model The target distance calculated at scale s;

[0040] S45. Use gradient descent to minimize the multi-scale refinement loss function , optimize the parameters of the implicit neural representation network , and construct the refined 3D model by extracting the zero level set of the implicit function:

[0041] ;

[0042] in, represents the refined 3D model, Represents the gradient of the signed distance at point x, represents the 2-norm of the gradient, represents the Laplace operator, is the gradient regularization coefficient, is the Laplace regularization coefficient.

[0043] Optionally, the S5 specifically includes:

[0044] S51, the refined 3D model Input to the improved conditional diffusion inversion module, use the multi-step inverse diffusion update formula to correct the model noise, and adaptively adjust the local structure to generate an intermediate diffusion model ;

[0045] S52. Apply multi-scale discrete wavelet transform to the refined 3D model to extract the frequency domain detail features W, and modulate them using the adaptive frequency gating mechanism:

[0046] ;

[0047] in, represents the discrete wavelet transform operator, is the Fourier transform, is the Sigmoid function, is the frequency modulation coefficient, represents element-wise multiplication;

[0048] S53, construct a cross-domain fusion module, and transform the intermediate diffusion model Fuse with the frequency domain detail feature W to generate a fusion model :

[0049] ;

[0050] in, represents the topological features extracted from the frequency domain detail features through the graph convolutional network, and is the fusion weight coefficient, is the cosine function, represents the Frobenius norm;

[0051] S54, fusion model The input is sent to the differentiable rendering feedback correction module. The differentiable rendering feedback correction module uses the differentiable renderer to generate a multi-view rendering image for the fusion model, and compares it with the perspective features of the input image to calculate the deformation correction parameters:

[0052] ;

[0053] Where V is the number of viewing angles, represents the feature representation of the input image at the vth viewing angle, represents the image rendered by the fusion model at the vth viewing angle, is the correction function module, is the correction factor, is the calibrated 3D model;

[0054] S55. Using an alternating iterative optimization strategy, the improved conditional diffusion inversion module, multi-scale discrete wavelet transform module, cross-domain fusion module, and differentiable rendering feedback correction module are jointly updated until the output indicators of each module converge stably, and finally the optimized 3D model is output.

[0055] Optionally, the S51 specifically includes:

[0056] S511, the refined 3D model Perform normalization to obtain a normalized model ;

[0057] S512. A noise estimation mechanism is used on the normalized model to calculate the noise correction term through a noise prediction network, and an adaptive weight adjustment factor is introduced to modulate the noise:

[0058] ;

[0059] in, is the diffusion coefficient at time step t, is the cumulative product of the diffusion coefficients, is the noise prediction network’s estimate of the normalized model at time step t, represents the noise correction term, exp( ) is the exponential function, To adjust the parameters, is the local structure consistency index, Frobenius norm representing the local structure consistency index;

[0060] S513, normalized model Perform local structure analysis, use autocorrelation operation combined with second-order derivative information to extract local structure consistency indicators, and calculate local structure adjustment items:

[0061] ;

[0062] in, is the spectral norm, represents the Hessian norm, 、 and To adjust the parameters, ( ) is a logarithmic function, It is a local structural adjustment item;

[0063] S514, fusion noise correction term and local structural adjustment items Normalized model Perform multi-step inverse diffusion updates to generate a temporary intermediate diffusion model :

[0064] ;

[0065] in, represents element-wise multiplication, is the interaction modulation coefficient;

[0066] S515, temporary intermediate diffusion model As the final intermediate diffusion model output, generate the intermediate diffusion model .

[0067] Optionally, the S6 specifically includes:

[0068] S61. Establish a data input and preprocessing module to unify the formatting of single photos and preliminary 3D models and build a data flow pipeline;

[0069] S62: Build a deep feature extraction module to extract global and local features from the input photo, and pass the global and local features to the 3D model generation module and the 3D model refinement processing module;

[0070] S63, constructing a 3D model generation module, using a conditional generative adversarial network to generate a preliminary 3D model based on the depth features, and setting a discriminator to provide real-time feedback on the generation process;

[0071] S64. Build a 3D model refinement processing module. Based on implicit neural representation technology, convert the preliminary 3D model into a continuous function representation and continuously improve the model details and surface smoothness through a multi-scale refinement strategy.

[0072] S65. Build an end-to-end joint training system, integrating the deep feature extraction module, 3D model generation module, and 3D model refinement processing module into a unified training framework. The end-to-end joint training system continuously adjusts the parameters of the deep feature extraction module, 3D model generation module, and 3D model refinement processing module through multi-task control and global loss feedback.

[0073] The beneficial effects of the present invention are:

[0074] The present invention organically combines generative adversarial networks with implicit neural representation technology to directly generate printable 3D models from a single photo, achieving high-precision conversion from two-dimensional images to three-dimensional models. Compared with traditional 3D reconstruction methods that rely on multi-view data or laser scanning, the present invention significantly reduces data acquisition costs and equipment requirements, while breaking through the technical bottlenecks of local detail blurring, noise interference, and structural distortion caused by insufficient information in a single image. By performing deep feature extraction on the input image through a multi-resolution hybrid expert network, and using a conditional generative adversarial network to generate a preliminary 3D model, and then performing multi-scale refinement processing through an implicit neural representation network, the present invention achieves all-round reconstruction from global shape to local texture details, ensuring that the generated 3D model meets high standards in both geometric accuracy and surface smoothness.

[0075] On this basis, the present invention further introduces an improved conditional diffusion inversion module, which adaptively corrects the noise in the refined model through a multi-step inverse diffusion update formula, and compensates for it in combination with local structural consistency and second-order structural information, thereby retaining and enhancing local details while reducing noise. Combined with multi-scale discrete wavelet transform, cross-domain graph convolution fusion and differentiable rendering feedback correction technology, the intermediate diffusion model shows excellent performance in both global topological structure and local detail recovery. End-to-end joint training is achieved between modules through an alternating iterative optimization strategy, which further improves the robustness and stability of the model, ensuring that effective information transmission and accurate restoration of details can be achieved at each stage from the input photo to the final generated 3D model.

[0076] This method not only breaks through the traditional 3D reconstruction method's reliance on multi-view data, allowing a single image to generate a high-quality 3D model, but also improves the model's detail expression and structural stability through a number of innovative technologies, thereby significantly improving 3D printing adaptability. The final generated 3D model is realistic in visual effects and fine in geometric structure, and can meet the needs of industrial manufacturing, personalized customization, and cultural creativity for high-precision, high-detail 3D models. Through the method of the present invention, users can obtain high-quality 3D models at low cost and high efficiency, providing solid technical support for subsequent applications such as 3D printing, virtual reality, and digital twins, and has significant prospects for promotion and application and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0078] Figure 1 This is a flowchart of a method for generating a printable 3D model from a single photo based on a generative adversarial network proposed in the present invention;

[0079] Figure 2 This is a schematic diagram of the multi-scale refinement of the preliminary 3D model by the implicit neural representation network of the method for generating a printable 3D model from a single photo based on a generative adversarial network proposed in the present invention. DETAILED DESCRIPTION

[0080] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0081] refer to Figure 1 and Figure 2 A method for generating a printable 3D model from a single photo based on a generative adversarial network includes the following steps:

[0082] S1. Obtain a single input image of the target object and preprocess the input image;

[0083] S2. Input the preprocessed image into the encoder, extract the multi-scale features of the preprocessed image, and output the deep feature representation;

[0084] S3. The deep feature representation is input into the generative adversarial network. The generator generates a 3D model based on the deep feature representation. The discriminator compares the generated model with the real model and provides feedback to guide the generator to optimize the quality of the 3D model.

[0085] S4. Inputting the 3D model into the implicit neural representation network, using the implicit representation method of the signed distance function to convert the 3D model data into a continuous function representation, and performing multi-scale refinement processing on the 3D model through the implicit neural representation network;

[0086] S5. Input the refined 3D model into the improved conditional diffusion inversion module, which combines multi-scale discrete wavelet transform with differentiable rendering feedback correction technology to optimize the global structure and local details of the 3D model;

[0087] S6. Build an end-to-end joint training system to optimize feature extraction, 3D model generation, and refinement processes;

[0088] S7. Perform mesh smoothing, topology correction, and format conversion on the optimized 3D model to generate a file format that meets the 3D printer standard, perform printing path planning, and finally output a 3D model that meets the printing requirements.

[0089] In this embodiment, S3 specifically includes:

[0090] S31, the output deep feature representation O and the adaptive noise vector z are fused after linear transformation, and the fusion condition vector C is calculated:

[0091] ;

[0092] in, is the learnable transformation matrix of the noise vector, represents element-wise multiplication, is the smoothing parameter, and are exponential and logarithmic functions respectively;

[0093] S32, input the fusion condition vector C into the hierarchical decoding network, and use multi-scale upsampling and residual connection modules to generate a 3D model :

[0094] ;

[0095] Among them, L is the number of layers of residual modules in the decoding network, is the upsampling convolution kernel of the i-th layer, Represents the processing output of the fusion condition vector by the i-th layer residual module, is the fusion coefficient in the residual connection of the i-th layer, Represents the convolution operation;

[0096] S33, the generated 3D model Perform adaptive normalization processing to obtain a normalized model :

[0097] ;

[0098] in, is the mean value of the 3D model, is the variance, To prevent division by zero for small positive numbers, Represents the result of local area average pooling of the 3D model. and To balance the contribution of each part, represents the Frobenius norm;

[0099] S34, normalize the model Input the discriminator separately with the real 3D model and calculate the adversarial loss based on the mixed gradient difference The discriminator captures the differences between the normalized model and the real 3D model in local texture, edge and structural gradient information, and feeds the differences back to the generator to guide the generator to continuously improve the quality of the 3D model:

[0100] ;

[0101] in, and is the balance coefficient, is a real 3D model, tanh() is the hyperbolic tangent function, is the spatial gradient operator, It means taking the sum of the absolute values of the elements;

[0102] S35. Adopt an alternating optimization strategy to jointly update the parameters of the generator and discriminator and output a 3D model.

[0103] In this embodiment, the S4 specifically includes:

[0104] S41. Generating a 3D model In the , the three-dimensional point set P is extracted by uniform sampling method:

[0105] ;

[0106] in, represents the number of sampling points, represents the three-dimensional real space, Represents the i-th sampling point, where i is the index of the sampling point;

[0107] S42. Initialize the implicit neural representation network , the implicit neural representation network accepts three-dimensional coordinates as input and outputs the corresponding signed distance s;

[0108] S43, for each sampling point , using implicit networks to calculate the predicted signed distance ;

[0109] S44. Define multi-scale refinement loss function is the weighted average of the error between the predicted symbol distance and the target distance at each scale:

[0110] ;

[0111] Among them, S represents the number of scales, is the weight coefficient of the s-th scale, Representation based on preliminary 3D model The target distance calculated at scale s;

[0112] S45. Use gradient descent to minimize the multi-scale refinement loss function , optimize the parameters of the implicit neural representation network , and construct the refined 3D model by extracting the zero level set of the implicit function:

[0113] ;

[0114] in, represents the refined 3D model, Represents the gradient of the signed distance at point x, represents the 2-norm of the gradient, represents the Laplace operator, is the gradient regularization coefficient, is the Laplace regularization coefficient.

[0115] In this embodiment, the S5 specifically includes:

[0116] S51, the refined 3D model Input to the improved conditional diffusion inversion module, use the multi-step inverse diffusion update formula to correct the model noise, and adaptively adjust the local structure to generate an intermediate diffusion model ;

[0117] S52. Apply multi-scale discrete wavelet transform to the refined 3D model to extract the frequency domain detail features W, and modulate them using the adaptive frequency gating mechanism:

[0118] ;

[0119] in, represents the discrete wavelet transform operator, is the Fourier transform, is the Sigmoid function, is the frequency modulation coefficient, represents element-wise multiplication;

[0120] S53, construct a cross-domain fusion module, and transform the intermediate diffusion model Fuse with the frequency domain detail feature W to generate a fusion model :

[0121] ;

[0122] in, represents the topological features extracted from the frequency domain detail features through the graph convolutional network, and is the fusion weight coefficient, is the cosine function, represents the Frobenius norm;

[0123] S54, fusion model The input is sent to the differentiable rendering feedback correction module. The differentiable rendering feedback correction module uses the differentiable renderer to generate a multi-view rendering image for the fusion model, and compares it with the perspective features of the input image to calculate the deformation correction parameters:

[0124] ;

[0125] Where V is the number of viewing angles, represents the feature representation of the input image at the vth viewing angle, represents the image rendered by the fusion model at the vth viewing angle, is the correction function module, is the correction factor, is the calibrated 3D model;

[0126] S55. Using an alternating iterative optimization strategy, the improved conditional diffusion inversion module, multi-scale discrete wavelet transform module, cross-domain fusion module, and differentiable rendering feedback correction module are jointly updated until the output indicators of each module converge stably, and finally the optimized 3D model is output.

[0127] In this embodiment, the S51 specifically includes:

[0128] S511, the refined 3D model Perform normalization to obtain a normalized model ;

[0129] S512. A noise estimation mechanism is used on the normalized model to calculate the noise correction term through a noise prediction network, and an adaptive weight adjustment factor is introduced to modulate the noise:

[0130] ;

[0131] in, is the diffusion coefficient at time step t, is the cumulative product of the diffusion coefficients, is the noise prediction network’s estimate of the normalized model at time step t, represents the noise correction term, exp( ) is the exponential function, To adjust the parameters, is the local structure consistency index, Frobenius norm representing the local structure consistency index;

[0132] S513, normalized model Perform local structure analysis, use autocorrelation operation combined with second-order derivative information to extract local structure consistency indicators, and calculate local structure adjustment items:

[0133] ;

[0134] in, is the spectral norm, represents the Hessian norm, 、 and To adjust the parameters, ( ) is a logarithmic function, It is a local structural adjustment item;

[0135] S514, fusion noise correction term and local structural adjustment items Normalized model Perform multi-step inverse diffusion updates to generate a temporary intermediate diffusion model :

[0136] ;

[0137] in, represents element-wise multiplication, is the interaction modulation coefficient;

[0138] S515, temporary intermediate diffusion model As the final intermediate diffusion model output, generate the intermediate diffusion model .

[0139] In this embodiment, S6 specifically includes:

[0140] S61. Establish a data input and preprocessing module to unify the formatting of single photos and preliminary 3D models and build a data flow pipeline;

[0141] S62: Build a deep feature extraction module to extract global and local features from the input photo, and pass the global and local features to the 3D model generation module and the 3D model refinement processing module;

[0142] S63, constructing a 3D model generation module, using a conditional generative adversarial network to generate a preliminary 3D model based on the depth features, and setting a discriminator to provide real-time feedback on the generation process;

[0143] S64. Build a 3D model refinement processing module. Based on implicit neural representation technology, convert the preliminary 3D model into a continuous function representation and continuously improve the model details and surface smoothness through a multi-scale refinement strategy.

[0144] S65. Build an end-to-end joint training system, integrating the deep feature extraction module, 3D model generation module, and 3D model refinement processing module into a unified training framework. The end-to-end joint training system continuously adjusts the parameters of the deep feature extraction module, 3D model generation module, and 3D model refinement processing module through multi-task control and global loss feedback.

[0145] Example 1:

[0146] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the digital transformation of a large manufacturing enterprise. Traditional 3D reconstruction technology mainly relies on multi-view image acquisition or laser scanning equipment, which not only leads to high data acquisition costs, but also cumbersome operation process, and the generated 3D model often has deficiencies in detail recovery and printing adaptability. In order to solve this series of problems, the present invention adopts a method for generating a printable 3D model from a single photo based on a generative adversarial network, which can extract implicit three-dimensional information from an ordinary photo and realize the rapid generation of high-precision 3D models. The method automatically extracts global and local features in the photo through a deep neural network, and combines implicit neural representation technology to convert discrete 3D data into a continuous function representation, and then through a number of innovative technologies such as improved conditional diffusion inversion module, multi-scale discrete wavelet transform, cross-domain fusion and differentiable rendering feedback correction, the initially generated 3D model is multi-level refined and globally optimized, thereby generating a 3D model with rich details and high printing adaptability.

[0147] In practical applications, this manufacturing company faces the need to rapidly digitally model a number of complex parts for virtual simulation, process planning, and 3D printing. Traditional methods typically require the use of multiple cameras or high-precision scanners to capture image data from different angles. This process is time-consuming, complex, and prone to errors. However, using the method presented in this invention, a single image is captured with a standard camera. Through automated preprocessing, deep feature extraction, conditional generative adversarial network modeling, implicit neural representation refinement, and subsequent conditional diffusion inversion and cross-domain fusion optimization, a high-precision 3D model can be rapidly generated. Actual tests conducted in 2024 showed that generating a 3D model of a single part using the method presented in this invention took an average of only 3.2 minutes, compared to approximately 15 minutes using traditional multi-view scanning methods. Furthermore, the 3D models generated using this method excel in both global structure and local detail recovery, with a detail recovery rate of up to 92%, a 3D printing success rate of 96%, and an average user satisfaction rating of 9.2 out of 10. These data fully demonstrate that the method of the present invention has significant advantages in shortening generation time, improving reconstruction accuracy, and improving 3D printing adaptability.

[0148] In this embodiment, a single input photo is first preprocessed and deep features extracted using a multi-resolution hybrid expert network, achieving comprehensive extraction of multi-scale information from the image and outputting a high-dimensional deep feature representation. Next, a conditional generative adversarial network is used to generate a preliminary 3D model. During this process, adversarial training between the generator and the discriminator effectively improves the accuracy of the generated model's global structure and local details. Subsequently, the preliminary 3D model is input into an implicit neural representation network, where the signed distance function is used to convert discrete data into a continuous function representation. Multi-scale refinement further improves model accuracy. To further reduce noise interference and compensate for local structural distortion, the present invention introduces an improved conditional diffusion inversion module, which performs multi-step inverse diffusion updates on the refined 3D model to generate an intermediate diffusion model. Subsequently, the intermediate diffusion model is optimized globally and in local detail through cross-domain fusion and a differentiable rendering feedback correction module, ultimately achieving the generation of a high-quality 3D model. Finally, through an end-to-end joint training system, each module is collaboratively optimized, achieving a complete automated process from photo input to 3D model printing.

[0149] Table 1 Comparison of 3D model generation performance

[0150] ;

[0151] Table 1 shows that the average generation time for the proposed method across all test batches remained between 3.0 and 3.5 minutes, significantly outperforming the 14.5 to 15.2-minute average for traditional methods. This result demonstrates that the method for generating 3D models from a single photo offers significant advantages in data processing efficiency and model generation speed, particularly in scenarios requiring rapid response and immediate model output.

[0152] The method of the present invention also performs well in detail recovery, maintaining a detail recovery rate of over 90% in each batch of tests, reaching a maximum of 93%, while the detail recovery rate of traditional multi-view methods ranges from 75% to 80%. This is mainly due to the multi-resolution hybrid expert network, implicit neural representation method, and conditional diffusion inversion module introduced in the present invention. Through multi-scale refinement and adaptive regulation, it effectively captures the multi-scale features and detail information in the image, allowing the generated 3D model to present finer details in both visual effects and actual physical models.

[0153] The comparative data of 3D printing success rate further verified the practical application effect of the method of the present invention. The 3D printing success rate of the method of the present invention remained above 94% in each test batch, with the highest reaching 97%. In comparison, the 3D printing success rate of the traditional method was only 83% to 87%. This shows that through the optimization of cross-domain fusion and differentiable rendering feedback correction module, the 3D model generated by the method of the present invention has higher reliability in structural stability, topological consistency and printing adaptability, and can effectively avoid printing failure problems caused by model instability or insufficient data accuracy during the printing process.

[0154] User satisfaction scores, as a comprehensive evaluation metric, can intuitively reflect users' approval of model generation performance, user experience, and printed results. The user satisfaction scores for the method presented in this paper all exceeded 9.0 out of 10, with the highest reaching 9.4 points. In contrast, the scores for traditional methods ranged from 7.2 to 7.6 points. This data not only demonstrates the significant improvement in model generation performance achieved by the method presented in this paper, but also reflects the high level of user appreciation for its operational simplicity, generation speed, and model quality in practical applications.

[0155] In summary, by comparing the data in the analysis table, it can be seen that the method of the present invention is significantly superior to the traditional multi-view method in terms of 3D model generation time, detail recovery rate, printing success rate and user satisfaction. This is due to the multiple technical innovations of the present invention, including the combination of generative adversarial networks and implicit neural representations, the improvement of the conditional diffusion inversion module, the application of cross-domain fusion and differentiable rendering feedback correction, and the construction of an end-to-end joint training system. The comprehensive application of these technologies not only improves the efficiency and accuracy of model generation, but also significantly improves the success rate of 3D printing, making the present invention have extremely high application value and promotion prospects in the fields of actual production, digital design and 3D printing.

[0156] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for generating a printable 3D model from a single photo based on a generative adversarial network, characterized in that: The steps include: S1. Obtain a single input image of the target object and preprocess the input image; S2. Input the preprocessed image into the encoder, extract the multi-scale features of the preprocessed image, and output the deep feature representation; S3. The deep feature representation is input into the generative adversarial network. The generator generates a 3D model based on the deep feature representation. The discriminator compares the generated model with the real model and provides feedback to guide the generator to optimize the quality of the 3D model. S4. Inputting the 3D model into the implicit neural representation network, using the implicit representation method of the signed distance function to convert the 3D model data into a continuous function representation, and performing multi-scale refinement processing on the 3D model through the implicit neural representation network; S5. Input the refined 3D model into the improved conditional diffusion inversion module, which combines multi-scale discrete wavelet transform with differentiable rendering feedback correction technology to optimize the global structure and local details of the 3D model; S6. Build an end-to-end joint training system to optimize feature extraction, 3D model generation, and refinement processes; S7, performing mesh smoothing, topology correction, and format conversion on the optimized 3D model to generate a file format that meets the 3D printer standard, and performing printing path planning to ultimately output a 3D model that meets printing requirements; The S3 includes the following steps: S31, the output depth feature representation With adaptive noise vector After linear transformation, they are fused and the fusion condition vector is calculated. : ; in, is the learnable transformation matrix of the noise vector, represents element-wise multiplication, is the smoothing parameter, and are exponential and logarithmic functions respectively; S32, fusion condition vector Input hierarchical decoding network, use multi-scale upsampling and residual connection module to generate 3D model : ; in, is the number of layers of residual modules in the decoding network, For the Layer upsampling convolution kernel, Indicates the The layer residual module processes and outputs the fusion condition vector. For the The fusion coefficients in the layer residual connections, Represents the convolution operation; S33, the generated 3D model Perform adaptive normalization processing to obtain a normalized model : ; in, is the mean value of the 3D model, is the variance, To prevent division by zero for small positive numbers, Represents the result of local area average pooling of the 3D model. and To balance the contribution of each part, represents the Frobenius norm; S34, normalize the model Input the discriminator separately with the real 3D model and calculate the adversarial loss based on the mixed gradient difference The discriminator captures the differences between the normalized model and the real 3D model in local texture, edge and structural gradient information, and feeds the differences back to the generator to guide the generator to continuously improve the quality of the 3D model: ; in, and is the balance coefficient, For real 3D models, ( ) is the hyperbolic tangent function, is the spatial gradient operator, It means taking the sum of the absolute values of the elements; S35, using an alternating optimization strategy to jointly update the parameters of the generator and discriminator, and output a 3D model; The S4 comprises the following steps: S41. From the generated 3D model In the 3D point set, a uniform sampling method is used to extract the : ; in, represents the number of sampling points, represents the three-dimensional real space, Represents the i-th sampling point, where i is the index of the sampling point; S42. Initialize the implicit neural representation network , the implicit neural representation network accepts three-dimensional coordinates as input and outputs the corresponding signed distance ; S43, for each sampling point , using implicit networks to calculate the predicted signed distance ; S44. Define multi-scale refinement loss function is the weighted average of the error between the predicted symbol distance and the target distance at each scale: ; in, Indicates the number of scales, For the The weight coefficient of each scale, Representation based on preliminary 3D model In scale The target distance calculated below; S45. Use gradient descent to minimize the multi-scale refinement loss function , optimize the parameters of the implicit neural representation network , and construct the refined 3D model by extracting the zero level set of the implicit function: ; in, represents the refined 3D model, Indicates the symbol distance is The gradient of the point, represents the 2-norm of the gradient, represents the Laplace operator, is the gradient regularization coefficient, is the Laplace regularization coefficient; The S5 comprises the following steps: S51, the refined 3D model Input to the improved conditional diffusion inversion module, use the multi-step inverse diffusion update formula to correct the model noise, and adaptively adjust the local structure to generate an intermediate diffusion model ; S52. Apply multi-scale discrete wavelet transform to the refined 3D model to extract frequency domain detail features , and modulates using an adaptive frequency gating mechanism: ; in, represents the discrete wavelet transform operator, is the Fourier transform, is the Sigmoid function, is the frequency modulation coefficient, represents element-wise multiplication; S53, construct a cross-domain fusion module, and transform the intermediate diffusion model and frequency domain detail features Perform fusion and generate a fusion model : ; in, represents the topological features extracted from the frequency domain detail features through the graph convolutional network, and is the fusion weight coefficient, is the cosine function, represents the Frobenius norm; S54, fusion model The input is sent to the differentiable rendering feedback correction module. The differentiable rendering feedback correction module uses the differentiable renderer to generate a multi-view rendering image for the fusion model, and compares it with the perspective features of the input image to calculate the deformation correction parameters: ; in, is the number of viewing angles, Indicates the Feature representation of the input image under different viewing angles, Indicates the The image rendered by the fusion model under the viewing angle, is the correction function module, is the correction factor, is the calibrated 3D model; S55. Using an alternating iterative optimization strategy, jointly update the improved conditional diffusion inversion module, the multi-scale discrete wavelet transform module, the cross-domain fusion module, and the differentiable rendering feedback correction module until the output indicators of each module converge stably, and finally output the optimized 3D model; The S51 includes the following steps: S511, the refined 3D model Perform normalization to obtain a normalized model ; S512. A noise estimation mechanism is used on the normalized model to calculate the noise correction term through a noise prediction network, and an adaptive weight adjustment factor is introduced to modulate the noise: ; in, For the The diffusion coefficient of the time step, is the cumulative product of the diffusion coefficients, For the noise prediction network at time step Noise estimate for the normalized model, represents the noise correction term, ( ) is an exponential function, To adjust the parameters, is the local structure consistency index, Frobenius norm representing the local structure consistency index; S513, normalized model Perform local structure analysis, use autocorrelation operation combined with second-order derivative information to extract local structure consistency indicators, and calculate local structure adjustment items: ; in, is the spectral norm, represents the Hessian norm, 、 and To adjust the parameters, ( ) is a logarithmic function, It is a local structural adjustment item; S514, fusion noise correction term and local structural adjustment items Normalized model Perform multi-step inverse diffusion updates to generate a temporary intermediate diffusion model : ; in, represents element-wise multiplication, is the interaction modulation coefficient; S515, temporary intermediate diffusion model As the final intermediate diffusion model output, generate the intermediate diffusion model .

2. The method for generating a printable 3D model from a single photo based on a generative adversarial network according to claim 1, characterized in that: The S6 specifically includes: S61. Establish a data input and preprocessing module to unify the formatting of single photos and preliminary 3D models and build a data flow pipeline; S62: Build a deep feature extraction module to extract global and local features from the input photo, and pass the global and local features to the 3D model generation module and the 3D model refinement processing module; S63, constructing a 3D model generation module, using a conditional generative adversarial network to generate a preliminary 3D model based on the depth features, and setting a discriminator to provide real-time feedback on the generation process; S64. Build a 3D model refinement processing module. Based on implicit neural representation technology, convert the preliminary 3D model into a continuous function representation and continuously improve the model details and surface smoothness through a multi-scale refinement strategy. S65. Build an end-to-end joint training system, integrating the deep feature extraction module, 3D model generation module, and 3D model refinement processing module into a unified training framework. The end-to-end joint training system continuously adjusts the parameters of the deep feature extraction module, 3D model generation module, and 3D model refinement processing module through multi-task control and global loss feedback.

Citation Information

Patent Citations

  • Multi-view picture data acquisition method based on deep learning

    CN119904725A

  • Image super-resolution reconstruction method based on deep learning

    CN120125436A