A 3D model generation method based on text prompts

A neural network-based method generates three-dimensional models from text prompts, addressing the need for specialized software by automating the process and improving model realism.

CN119648948BActive Publication Date: 2025-07-15NANJING NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411708150.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-15
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing methods for constructing three-dimensional models require specialized software and significant time from professionals, limiting accessibility and efficiency.

Method used

A method utilizing a neural network to generate three-dimensional models from text prompts, involving steps like predicting a signed distance function for rendering, enhancing image quality with a pre-trained diffusion model, and generating point cloud data for three-dimensional models, followed by automatic printing using a local network.

Benefits of technology

Enables users to generate and print three-dimensional models directly from text descriptions without requiring specialized software knowledge, enhancing model realism and reducing the complexity of the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648948B_ABST
    Figure CN119648948B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a three-dimensional model based on text prompts. The method includes: through a constructed neural network, predicting a signed distance function SDF according to dynamically optimized randomly sampled camera poses and lighting vectors and performing volume rendering to obtain an initial rendered image; through a pre-trained diffusion model, improving the quality of the rendered image and the prediction accuracy of the signed distance function according to the encoded text prompt information; through a trained neural network, predicting the rendered color and the signed distance function according to known observation point coordinates, and obtaining the three-dimensional model surface point cloud data in an interpolation manner, and finally generating a three-dimensional model format file. This application effectively solves the problems such as the need for professionals to use professional software for a long time to construct a three-dimensional model. The three-dimensional model constructed by this method can be directly printed using a three-dimensional printing device, and finally a physical three-dimensional object associated with the text prompt can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and three - dimensional model reconstruction. Specifically, it relates to a method for generating three - dimensional models based on text prompts. Background Art

[0002] Artificial intelligence (AI) aims to simulate human intelligent behavior, including capabilities such as learning, reasoning, and decision - making. Since the 1950s, AI has undergone several evolutions. Especially in the context of the improvement of computing power and the explosion of data volume, machine learning and deep - learning technologies have made significant progress. Deep learning uses multi - layer neural networks and can automatically extract features from massive amounts of data, and is widely applied in fields such as image recognition, natural language processing, and speech recognition. Its high efficiency has had a profound impact on industries such as healthcare, finance, and autonomous driving, promoting the rapid development of technology.

[0003] Three - dimensional printing (3D Printing), also known as additive manufacturing, is a technology for creating three - dimensional objects by adding materials layer by layer. This technology originated in the 1980s but has developed rapidly in recent years with the progress of materials science. The process of three - dimensional printing includes three stages: design, printing, and post - processing. In the design stage, computer - aided design (CAD) software is used to create three - dimensional models. In the printing stage, different materials and technologies (such as fused deposition modeling and selective laser sintering) are used for construction. Post - processing can further improve the quality and appearance of the finished product.

[0004] Generating three - dimensional models from text prompts is an emerging technology that uses natural language processing and deep learning to convert text descriptions into three - dimensional visualization models. This technology usually combines image generation and 3D modeling algorithms. By analyzing the key concepts and relationships in the text, corresponding three - dimensional shapes are generated. This process enables users to quickly obtain personalized three - dimensional models through simple text descriptions, and is widely applied in fields such as game development, virtual reality, and education, providing great convenience for design and creation. With the continuous progress of technology, its accuracy and application scope are also constantly expanding. Summary of the Invention

[0005] The present invention provides a method for generating three - dimensional models based on text prompts to solve problems in related technologies such as the need for professionals to use professional software for a long time to construct three - dimensional models. The three - dimensional models constructed by the method of this application can be directly printed using three - dimensional printing equipment, and finally, physical three - dimensional objects associated with the text prompts can be obtained.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A method for generating three - dimensional models based on text prompts, characterized by including the following steps:

[0008] S1: Based on the constructed neural network, predict the signed distance function SDF according to the dynamically optimized random sampling camera pose and illumination vector, and perform volume rendering to obtain the initial rendered image;

[0009] S2: Through the pre-trained diffusion model, improve the quality of the rendered image and the prediction accuracy of the signed distance function according to the encoded text prompt information;

[0010] S3: Through the trained neural network, predict the rendering color and the signed distance function according to the known observation point coordinates, obtain the point cloud data on the surface of the three-dimensional model in an interpolated manner, and finally generate a three-dimensional model format file;

[0011] S4: Slice the generated three-dimensional model format file and store it in the gcode format, transfer it to the printer through the local area network, and automatically print the described physical object.

[0012] As a preferred technical solution of the present invention: In step S1, the method for predicting the signed distance function and performing volume rendering to create a rendered image includes:

[0013] S11. Randomly initialize the sampling coordinates

[0014] Sample the information of the light rays and camera parameters, and dynamically optimize the sampling points through the spatial resolution, importance sampling and feedback mechanism;

[0015] S12. Create a neural network

[0016] The neural network includes two encoders and three multi-layer perceptrons. Use one encoder to encode the sampling coordinates, and then use two multi-layer perceptrons respectively. One predicts the signed distance function SDF and calculates its gradient, and the other predicts the albedo. Use the other encoder to encode the sampling light ray rotation vector, and use the third multi-layer perceptron to predict the background color;

[0017] S13. Calculate the transparency according to the signed distance function, and then obtain the ray weight

[0018] Combine the albedo, ambient light, diffuse light and ray weight to obtain the foreground color, and combine the ray weight, foreground color and background color to obtain the rendered image to complete the volume rendering.

[0019] As a preferred technical solution of the present invention: In step S2, the method for improving the quality of the rendered image and the prediction accuracy of the signed distance function includes:

[0020] S21. Tokenize and encode the input text information. Duplicate the text information four times for four viewing directions. Use a pre-trained model to calculate the relevance of each text information to its corresponding viewing direction, and delete the parts of the text information that are weakly relevant to the current viewing direction.

[0021] S22. First, follow the training process of the diffusion model. Add gradually increasing random noise to the rendered image over time steps. During the training process, the time steps gradually decrease according to the law of square annealing. Determine the viewing direction based on the camera pose corresponding to the rendered image, extract the corresponding direction, and perform semantic debiasing and encoding on the text prompt information. Then, bring the text prompt information and the rendered image with added random noise into the pre-trained diffusion model to predict the noise.

[0022] S23. Truncate the differences between the processed random noise and the predicted noise, and the differences between the original image and the image after removing the predicted noise, respectively, in a score-debiasing manner. Combine the gradient of the predicted signed distance function and the variance of the z coordinate of the sampling points to finally form the loss function of the trained neural network, and train the neural network.

[0023] As a preferred technical solution of the present invention: In step S3, the method for generating a three-dimensional model format file includes:

[0024] S31. Read the vertex data file containing the viewing vertex coordinates and viewing vertex numbers. Use the viewing vertex coordinates and substitute them into the neural network model predicted by the trained signed distance function to obtain the signed distance function corresponding to the viewing vertex coordinates.

[0025] S32. According to the viewing vertex coordinates and the signed distance function, construct a tetrahedron that penetrates the object surface. For two vertex coordinates where the two endpoints of the line segment are inside and outside the geometry and the signed distance functions have different signs, calculate the three-dimensional surface point coordinates by interpolation to establish interpolation vertices, and record the interpolation vertex numbers that make up each triangular patch.

[0026] S33. First, bring the interpolation vertices into the trained neural network model for albedo prediction to obtain the albedo corresponding to the interpolation vertices. Then, rasterize the three-dimensional model and further interpolate the interpolation vertices according to the pixel height and width required for rendering to obtain interpolation vertices that conform to the pixels. Bring them into the trained neural network model for albedo prediction again to obtain the albedo of the interpolation vertices with pixel distribution.

[0027] S34. According to the interpolation vertices, the interpolation vertices with pixel distribution, the corresponding triangular patch numbers, the albedo corresponding to the interpolation vertices, and the albedo corresponding to the interpolation vertices with pixel distribution, follow the requirements of the three-dimensional model file storage format to finally obtain three-dimensional model files in different formats.

[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0029] 1. The threshold for obtaining 3D printing models is reduced:

[0030] Users no longer need to learn 3D modeling software and slicing software. They can directly obtain the required 3D model by only inputting the text description of the target object.

[0031] 2. High integration:

[0032] The whole solution of the present invention includes three steps: generating a 3D model from text based on a deep learning method, slicing the 3D model, and remotely sending the sliced file to the printer to start automatic printing. All three steps can be implemented using Python and encapsulated in the same functional package. Users only need to input the text description and some printing parameters to be modified to obtain the final physical object.

[0033] 3. The model is more vivid;

[0034] The 3D model generated by the deep learning-based method of the present invention aims to achieve photo-realism, and adopts smoother curves and techniques that integrate information such as colors and textures, so that more vivid effects may be obtained in some applications. This method can supplement the functions of traditional software such as SolidWorks and UG in specific scenarios, improving the realism and fineness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the overall method flow chart of this application;

[0036] Figure 2 is the method flow chart from text prompt to generating 3D model;

[0037] Figure 3 is the method flow chart from 3D model to slicing and printing;

[0038] Figure 4 is the calculation structure diagram from text prompt to generating 3D model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0040] As Figures 1-4 shown, the present invention proposes a method for generating a 3D model based on text prompts, including the following steps:

[0041] S1: Through the constructed neural network, predict the signed distance function SDF according to the dynamically optimized random sampling camera pose and lighting vector and perform volume rendering to obtain the initial rendered image;

[0042] A method for predicting a signed distance function and performing volume rendering to create a rendered image includes:

[0043] S11. Randomly initialize sampling coordinates

[0044] Sample information of rays and camera parameters, and dynamically optimize sampling points through spatial resolution, importance sampling, and a feedback mechanism;

[0045] S12. Create a neural network

[0046] The neural network includes two encoders and three multi-layer perceptrons. One encoder is used to encode the sampling coordinates, and then two multi-layer perceptrons are respectively used. One predicts the signed distance function SDF and calculates its gradient, and the other predicts the albedo. Another encoder is used to encode the sampling ray rotation vector, and the third multi-layer perceptron is used to predict the background color;

[0047] S13. Calculate the transparency according to the signed distance function, and then obtain the ray weight

[0048] Combine the albedo, ambient light, diffuse light, and ray weight to obtain the foreground color, and combine the ray weight, foreground color, and background color to obtain the rendered image to complete the volume rendering;

[0049] S2: Through a pre-trained diffusion model, improve the quality of the rendered image and the prediction accuracy of the signed distance function according to the encoded text prompt information;

[0050] Methods for improving the quality of the rendered image and the prediction accuracy of the signed distance function include:

[0051] S21. Tokenize and encode the input text information, copy the text information four times for four viewing directions, use a pre-trained model to calculate the relevance of each text information to its corresponding viewing direction, and delete the part of the text information that is weakly relevant to the current viewing direction;

[0052] S22. First, follow the training process of the diffusion model, add randomly increasing noise to the rendered image over time steps, and the time steps gradually decrease according to the law of square annealing during the training process. Determine the viewing direction according to the camera pose corresponding to the rendered image, extract the corresponding direction and perform semantic debiasing and encoded text prompt information. Then, bring the text prompt information and the rendered image with added random noise into the pre-trained diffusion model to predict the noise;

[0053] S23. Truncate the differences between the processed random noise and the predicted noise, and between the original image and the image after removing the predicted noise, respectively, in a score-debiased manner, and combine the gradient of the predicted signed distance function and the variance of the z-coordinate of the sampling points to finally form the loss function of the trained neural network, and train the neural network;

[0054] S3: Through the trained neural network, predict the rendering color and the signed distance function according to the known observation point coordinates, obtain the point cloud data on the surface of the 3D model by interpolation, and finally generate a 3D model format file;

[0055] The method for generating a 3D model format file includes:

[0056] S31. Read the vertex data file containing the observation vertex coordinates and the observation vertex numbers, and use the observation vertex coordinates to substitute into the neural network model predicted by the trained signed distance function to obtain the signed distance function corresponding to the observation vertex coordinates;

[0057] S32. According to the observation vertex coordinates and the signed distance function, construct a tetrahedron that penetrates the surface of the object, and for the two vertex coordinates where the two endpoints of the line segment are respectively inside and outside the geometry and the signed distance function has different signs, calculate the 3D surface point coordinates by interpolation to establish interpolation vertices, and record the interpolation vertex numbers that form each triangular patch;

[0058] S33. First, substitute the interpolation vertices into the trained neural network model for albedo prediction to obtain the albedo corresponding to the interpolation vertices. Then, rasterize the 3D model and further interpolate the interpolation vertices according to the pixel height and width required for rendering to obtain interpolation vertices that conform to the pixels. Substitute them into the trained neural network model for albedo prediction again to obtain the albedo of the interpolation vertices with pixel distribution;

[0059] S34. According to the interpolation vertices, the interpolation vertices with pixel distribution, the corresponding triangular patch numbers, the albedo corresponding to the interpolation vertices, and the albedo corresponding to the interpolation vertices with pixel distribution, and follow the requirements of the 3D model file storage format, finally obtain 3D model files in different formats, including 3D model files in multiple formats such as stl, obj, obj+mtl, etc.;

[0060] S4: Slice the generated 3D model format file and store it in gcode format, and transmit it to the printer through the local area network to automatically print the described physical object.

[0061] The part from text description to 3D model generation is mainly based on deep learning. Use the pre-trained transformers library and diffusers library in huggingface to encode the text and optimize the guiding model, and create a 3D model mesh based on the signed distance function SDF.

[0062] First, take the camera pose lighting vectors after normal distribution and adding perturbations, etc. as the model input batch, and use the nerfacc.OccGridEstimator.sampling function to dynamically optimize the lighting points.

[0063] Secondly, create a neural network. Use the tinycudann.Encoding encoder to encode the input lighting midpoint vector. Use two multi-layer perceptrons. One is used to predict the signed distance function SDF after receiving the output of the encoder, and the other is used to extract features after receiving the encoder output for subsequent volume rendering. For the predicted signed distance function SDF, interpolation and gradient calculation are performed through the finite difference method. On the one hand, the gradient of the signed distance function will be used as part of the neural network loss function. By reducing the gradient of the signed distance function, the generated model will be smoother and more realistic. On the other hand, based on the signed distance function SDF and its gradient, the cumulative density function CDF can be calculated, and then the transparency of this part of the 3D model can be obtained.

[0064] The transparency calculation formula is as follows:

[0065]

[0066] Where rad is the radius, sam is the number of sampling points, and std is the dynamic inverse standard deviation participating in the neural network training. After obtaining the transparency, the nerfacc.render_weight_from_alpha function can be used to convert it into ray weights and apply them to subsequent volume rendering.

[0067] Then perform volume rendering. Bring the previously extracted features into the activation function to obtain the albedo. The albedo fuses the hyperparameters of ambient light and diffuse light, and then combines with the ray weights converted from transparency. Based on the nerfacc.accumulate_along_rays function, the foreground color is accumulated along the rays. Use another set of neural networks, which are also the tinycudann.Encoding encoder, multi-layer perceptron, and activation function, to separately predict a background color. Similarly, follow the function nerfacc.accumulate_along_rays along the ray direction to accumulate the ray weights to obtain the rendering weights, and fuse the foreground color and background color according to the rendering weights to obtain the rendered image.

[0068] Next is the rendering image optimization. First, the input text needs to be processed. The T5Tokenizer and T5EncoderModel in the transformers library provided by huggingface are used to tokenize and encode the input text prompt. The text is copied four times and distributed to four perspectives. An off-the-shelf language model is used to identify the contradictions between the perspective prompts and the user prompts, and the four text description encodings that are strongly correlated with the four observation directions are retained. Then, the current direction can be determined according to the three-dimensional polar coordinates composed of the elevation angle, azimuth angle, and lighting distance of the lighting point, and the text description encoding with direction pertinence can be extracted. The rendering image optimization is mainly based on the diffusion model. The IFPipeline function and the pre-trained model DeepFloyd / IF-I-XL-v1.0 in the diffusers library provided by huggingface are used. During the forward process, random noise is added to the rendering image and continuously increases as the time step increases. During the reverse process, the pre-trained unet is used to predict the noise based on the rendered image with added noise and the text prompt encoding.

[0069] The original diffusion model loss function is based on the offset between the predicted noise and the random noise, following the formula:

[0070]

[0071] On this basis, this solution further introduces the supervision of the diffusion model's detected images. The original image is compared with the image after denoising according to the predicted noise and added to the loss function to obtain a new diffusion model loss function, following the formula:

[0072]

[0073] Thus, the supervision effect on the generated images is improved, and the model convergence is accelerated. Considering the multi-faceted Janus problem during the generation of the three-dimensional model, a score debiasing solution is introduced here, following the formula:

[0074]

[0075] where Clip(x,c) = max(min(x,c), -c), and the probability of artifacts appearing is reduced by truncating the larger part of the noise prediction loss in a timely manner.

[0076] In addition, as the model is optimized, the time step of the forward diffusion process is set to be inversely proportional to the square root of the number of training iterations, and the random time step size is gradually reduced in the way of square time step annealing, so as to effectively improve the model generation quality. The corresponding formula is written as:

[0077]

[0078] The last part of the loss function is the calculation of the variance of the z - coordinates of the ray sampling points. By reducing the variance of the z - coordinates of the sampling point distribution, the generated geometric shape can obtain a sharper surface, thus achieving a more vivid 3D model generation effect. The calculation method of the Z - variance loss follows the formula:

[0079]

[0080] where the z - variance is expressed as

[0081] Finally, following the UDP protocol, the gcode file is transmitted through the local area network to the host computer connected to the printer, and after decoding, the M24 instruction is automatically called to start printing.

[0082] Based on the above - mentioned technical solutions, the present invention can achieve:

[0083] 1. The threshold for obtaining a 3D printing model is reduced:

[0084] Users no longer need to learn 3D modeling software and slicing software. They only need to input a text description of the target object to directly obtain the required 3D model.

[0085] 2. High integration:

[0086] The whole set of solutions of the present invention includes three steps: generating a 3D model from text based on a deep - learning method, slicing the 3D model, and remotely sending the sliced file to the printer to start automatic printing. All three steps can be implemented using Python and are encapsulated in the same functional package. Users only need to input a text description and some printing parameters to be modified to obtain the final physical object.

[0087] 3. The model is more vivid;

[0088] The 3D model generated by the deep - learning - based method of the present invention aims to achieve photo - realistic quality. It adopts smoother curves and technologies that integrate information such as colors and textures, so that more vivid effects may be obtained in some applications. This method can complement the functions of traditional software such as SolidWorks and UG in specific scenarios, improving the realism and fineness of the model.

[0089] The above - mentioned are only the preferred embodiments of the present invention, and it is not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. A three-dimensional model generation method based on text prompts, characterized in that: It includes the following steps: S1: Through the constructed neural network, predict the signed distance function (SDF) based on the dynamically optimized random sampling camera pose and illumination vector and perform volume rendering to obtain the initial rendered image; In step S1, the method for predicting the signed distance function and performing volume rendering to create the rendered image includes: S11. Randomly initialize the sampling coordinates Sample the information of the ray and camera parameters, and dynamically optimize the sampling points through the spatial resolution, importance sampling, and feedback mechanism; S12. Create a neural network The neural network includes two encoders and three multi-layer perceptrons. Use one encoder to encode the sampling coordinates, and then use two multi-layer perceptrons respectively. One predicts the signed distance function SDF and calculates its gradient, and the other predicts the albedo. Use the other encoder to encode the sampling ray rotation vector, and use the third multi-layer perceptron to predict the background color; S13. Calculate the transparency based on the signed distance function, and then obtain the ray weight Combine the albedo, ambient light, diffuse light, and ray weight to obtain the foreground color, and combine the ray weight, foreground color, and background color to obtain the rendered image to complete the volume rendering; S2: Through the pre-trained diffusion model, improve the quality of the rendered image and the prediction accuracy of the signed distance function SDF according to the encoded text prompt information; S3: Through the trained neural network, predict the rendering color and signed distance function based on the known observation point coordinates, obtain the 3D model surface point cloud data in an interpolated manner, and finally generate a 3D model format file; S4: Slice the generated 3D model format file and store it in the gcode format, transmit it to the printer through the local area network, and automatically print the described physical object.

2. The 3D model generation method based on text prompts according to claim 1, wherein: In step S2, the method for improving the quality of the rendered image and the prediction accuracy of the signed distance function includes: S21. Tokenize and encode the input text information, copy the text information four times for the four observation directions, use the pre-trained model to calculate the relevance of each text information to its corresponding observation direction, and delete the part of the text information that is weakly related to the current observation direction; S22. First, follow the training process of the diffusion model, add randomly increasing noise to the rendered image with the time step gradually increasing, and the time step follows the law of square annealing and gradually decreases during the training process. Determine the observation direction according to the camera pose corresponding to the rendered image, extract the corresponding direction and perform semantic debiasing and encoded text prompt information. Then, bring the text prompt information and the rendered image with added random noise into the pre-trained diffusion model to predict the noise; S23. Truncate the differences between the processed random noise and the predicted noise, and the differences between the original image and the image after removing the predicted noise in a score debiasing manner, combine the gradient of the predicted signed distance function, and the variance of the z coordinate of the sampling points, and finally construct the loss function of the trained neural network and train the neural network.

3. A method for generating a three-dimensional model based on text prompts according to claim 1, characterized in that: In step S3, the method for generating the 3D model format file includes: S31. Read the vertex data file containing the observed vertex coordinates and the observed vertex numbers. Use the observed vertex coordinates and substitute them into the neural network model predicted by the trained signed distance function to obtain the signed distance function corresponding to the observed vertex coordinates. S32. According to the observed vertex coordinates and the signed distance function, construct a tetrahedron that penetrates the object surface. For the two vertex coordinates where the two endpoints of the line segment are respectively inside and outside the geometry and the signed distance function has different signs, calculate the three-dimensional surface point coordinates by interpolation to establish interpolation vertices, and record the interpolation vertex numbers that make up each triangular patch. S33. First, substitute the interpolation vertices into the trained neural network model for albedo prediction to obtain the albedo corresponding to the interpolation vertices. Then, rasterize the three-dimensional model and further interpolate the interpolation vertices according to the pixel height and width required for rendering to obtain interpolation vertices that conform to the pixels. Substitute them into the trained neural network model for albedo prediction again to obtain the albedo of the interpolation vertices with pixel distribution. S34. According to the interpolation vertices, the interpolation vertices with pixel distribution, the corresponding triangular patch numbers, the albedo corresponding to the interpolation vertices, and the albedo corresponding to the interpolation vertices with pixel distribution, follow the requirements of the three-dimensional model file storage format. Finally, obtain three-dimensional model files in different formats.

Citation Information

Patent Citations

  • Three-dimensional mesh generation using signed distance functions

    WO2024220914A1