Three-dimensional model generation method and device and electronic equipment
By adjusting the model parameters based on the degree of matching of rendered images based on the text and the initial three-dimensional model, the problem of inaccurate generation of three-dimensional models in the prior art is solved, and more efficient and accurate three-dimensional model generation is achieved.
Patent Information
- Application Number
- CN202311514762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
When the prior art generates a three-dimensional model based on text guidance, it is difficult to describe the structural characteristics of the three-dimensional model in a complete and accurate manner, resulting in the generated three-dimensional model being unreasonable and inefficient.
By obtaining the description text of the three-dimensional model to be generated and the initial three-dimensional deformable model, rendering is performed to obtain the initial rendered image, and adjusting the model parameters based on the degree of matching between the image and the text, a three-dimensional structural model that conforms to the description text is generated.
It improves the rationality and accuracy of the generated three-dimensional model, simplifies the generation process, reduces the need for manual editing of description texts, and reduces labor costs.
Smart Images

Figure CN120014149A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a three-dimensional model generation method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the development of Internet technology, 3D virtual models are increasingly used in various industries, such as virtual games, 3D printing, animation video production, virtual fitting, etc. The technology of generating 3D models guided by text can automatically and efficiently generate 3D models that conform to the text description through given text content, and is therefore increasingly widely used.
[0003] When generating a three-dimensional model based on text guidance, the related technology usually pre-trains a text-to-3D model mapping model based on a large number of sample three-dimensional models that have been pre-labeled with sample description texts. When generating the three-dimensional model, the description text is input into the mapping model to generate the corresponding three-dimensional model.
[0004] Since the three-dimensional model is a three-dimensional structure with a relatively complex model structure, the sample description text in the process of training the mapping model is usually difficult to fully and accurately describe the structural characteristics of each part of the three-dimensional model and the overall structural characteristics of the three-dimensional model. The description text corresponding to the three-dimensional model to be generated usually lacks a complete and accurate description of the three-dimensional model, which may result in the three-dimensional model generated by the mapping model being unreasonable and inaccurate. Summary of the invention
[0005] The present application provides a three-dimensional model generation method, device, electronic device and computer-readable storage medium, which can improve the rationality and accuracy of the generated three-dimensional model. The specific method is as follows.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a three-dimensional model, the method comprising:
[0007] Acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated;
[0008] Rendering the initial three-dimensional deformable model to obtain a first initial rendered image;
[0009] Based on the matching degree between the first initial rendered image and the description text, the model parameters of the initial three-dimensional deformable model are adjusted to generate a three-dimensional structure model that conforms to the description text.
[0010] In a second aspect, an embodiment of the present application further provides a three-dimensional model generation device, the device comprising:
[0011] An acquisition unit, used to acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated;
[0012] A rendering unit, configured to render the initial three-dimensional deformable model to obtain a first initial rendered image;
[0013] An adjustment unit is used to adjust the model parameters of the initial three-dimensional deformable model based on the matching degree between the first initial rendering image and the description text, so as to generate a three-dimensional structure model that conforms to the description text.
[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including:
[0015] Processor; and
[0016] The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method described in any one of the first aspects is executed.
[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing a data processing program, which is executed by a processor to perform the method as described in any one of the first aspects.
[0018] Compared with the prior art, this application has the following advantages:
[0019] The three-dimensional model generation method provided in the embodiment of the present application obtains a description text corresponding to the three-dimensional model to be generated and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated, renders the initial three-dimensional deformable model to obtain a first initial rendered image, and then adjusts the model parameters of the initial three-dimensional deformable model based on the degree of matching between the first initial rendered image and the description text. Since the description text is used to describe the three-dimensional model to be generated, the degree of matching between the first initial rendered image and the description text can reflect the degree of structural difference between the three-dimensional model corresponding to the first initial rendered image and the three-dimensional model to be generated. Therefore, the degree of matching between the first initial rendered image and the description text can guide the adjustment of the model parameters of the initial three-dimensional deformable model, so that the initial three-dimensional deformable model after adjusting the model parameters matches the three-dimensional model described by the description text, thereby generating a three-dimensional structural model that conforms to the description text.
[0020] When generating a three-dimensional model, the solution provided in the present application guides the initial three-dimensional deformable model to adjust model parameters based on the description text corresponding to the three-dimensional model to be generated, so as to generate a three-dimensional structural model that conforms to the description text. The initial three-dimensional deformable model corresponding to the three-dimensional model to be generated is usually a model that already conforms to the structural layout of the three-dimensional model. The three-dimensional structural model obtained by adjusting the initial deformable three-dimensional model according to the description text can satisfy the structure described in the description text, and can also make the generated three-dimensional structural model conform to the structural layout of the three-dimensional model, so that the structure of the generated three-dimensional model is more reasonable and the accuracy is higher.
[0021] In addition, the present application adjusts the model parameters of the initial three-dimensional deformable model based on the degree of matching between the first initial rendered image and the description text. Since the rendered image is a two-dimensional image, the feature description of the two-dimensional image is simpler and more convenient than that of the three-dimensional model. Therefore, the degree of matching between the two-dimensional image and the description text can be easily determined, making the process of generating the three-dimensional model simple and easy to implement.
[0022] Since most three-dimensional models do not have corresponding descriptive texts pre-set, the training data required for training mapping models in related technologies requires manual editing of description texts for a large number of three-dimensional models, which has high labor costs and a very cumbersome training process. Since the present application does not need to train the mapping model used to generate the three-dimensional model, there is no need to collect a large number of description texts for the three-dimensional models in advance, which can greatly reduce the manpower spent on the process of generating the three-dimensional model and make the three-dimensional model generation process simpler and more convenient. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flow chart of a three-dimensional model generation method provided in an embodiment of the present application;
[0024] Figure 2 is another example flow chart of the three-dimensional model generation method provided in the embodiments of the present application;
[0025] Figure 3 This is an example effect diagram of a three-dimensional model generation method provided in an embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of the process of training the second image-text diffusion model and the third image-text diffusion model in an embodiment of the present application;
[0027] Figure 5 This is a flow chart of an example of generating a three-dimensional structure model in the three-dimensional model generating method provided in the embodiment of the present application;
[0028] Figure 6 is another flow chart of the three-dimensional model generation method provided in the embodiments of the present application;
[0029] Figure 7 is a structural diagram of an example of a three-dimensional model generating device provided in an embodiment of the present application;
[0030] Figure 8 It is a schematic diagram of the logical structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0031] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0032] With the development of Internet technology, 3D virtual models are increasingly used in various industries, such as virtual games, 3D printing, animation video production, virtual fitting, etc. The technology of generating 3D models guided by text can automatically and efficiently generate 3D models that conform to the text description through given text content, and is therefore increasingly widely used.
[0033] When generating a three-dimensional model based on text guidance, the related technology usually pre-trains a text-to-3D model mapping model based on a large number of sample three-dimensional models that have been pre-labeled with sample description texts. When generating the three-dimensional model, the description text is input into the mapping model to generate the corresponding three-dimensional model.
[0034] Since the three-dimensional model is a three-dimensional structure with a relatively complex model structure, the sample description text in the process of training the mapping model is usually difficult to fully and accurately describe the structural characteristics of each part of the three-dimensional model and the overall structural characteristics of the three-dimensional model. The description text corresponding to the three-dimensional model to be generated usually lacks a complete and accurate description of the three-dimensional model, which may result in the three-dimensional model generated by the mapping model being unreasonable and inaccurate.
[0035] In the process of training the mapping model, it is necessary to collect a large amount of natural language description texts corresponding to the three-dimensional models in advance. However, most three-dimensional models usually do not have corresponding text descriptions. It is necessary to manually edit the description text according to the overall appearance of the three-dimensional model. Editing the description text corresponding to the three-dimensional model usually takes a long time, and the data of each dimension of the three-dimensional model is also relatively large. The accuracy of the manually edited description text is unstable, which makes the training process of the mapping model more cumbersome and the model prediction accuracy is not high, resulting in a large workload required for the generation of the three-dimensional model and a low generation accuracy rate.
[0036] In order to solve the above-mentioned problems in the related technology, an embodiment of the present application provides a three-dimensional model generation method. The executor of the solution provided in the present application can be an electronic device, which can be a desktop computer, a laptop computer, a mobile device, a smart watch, a game console, a smart TV, a tablet computer, a server, etc., or other devices with data processing and display functions.
[0037] like Figure 1 , Figure 3 As shown, the three-dimensional model generation method provided in the embodiment of the present application includes the following steps S110 to S130.
[0038] Step S110: obtaining a description text corresponding to the three-dimensional model to be generated and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated.
[0039] The three-dimensional model to be generated may be a three-dimensional model of a human face to be generated, a three-dimensional model of a human hand to be generated, a three-dimensional model of a human body to be generated, a three-dimensional model of a four-limbed reptile to be generated, etc., but is not limited thereto.
[0040] like Figure 3 As shown, taking the to-be-generated 3D model as a to-be-generated face model as an example, the description text corresponding to the to-be-generated 3D model can be a natural language text input by a user, and the initial 3D deformable model is as follows Figure 3 The initialized 3D face model is shown in .
[0041] The description text corresponding to the 3D model to be generated can be edited by the designer and input into the electronic device, and the electronic device can obtain the description text corresponding to the 3D model to be generated input by the designer. The description text is used to describe the structural features of the 3D model to be generated, and can also be used to describe the appearance shape, skin color and other structural features of the 3D model to be generated, that is, the description text includes the structural description text of the structural features of the 3D model to be generated, and can also include the appearance description text of the appearance features of the 3D model to be generated. The appearance description text may include skin color description text, expression description text, surface texture description text, etc., but is not limited to this.
[0042] In the embodiments of the present application, various types of initial three-dimensional deformable models can be pre-set, for example, an initial three-dimensional deformable model of a human face, an initial three-dimensional deformable model of a hand, an initial three-dimensional deformable model of a body, an initial three-dimensional deformable model of a four-limbed reptile, an initial three-dimensional deformable model of a human body, etc. Those skilled in the art can set various types of initial three-dimensional deformable models according to specific design requirements, which are not specifically limited in the present application. The model parameters of the initial three-dimensional deformable model can be set to the average model parameters of the existing three-dimensional models, so that the model can be adjusted to match the description text more quickly in the future.
[0043] The above-mentioned initial three-dimensional deformable model (3D Morphable Model, 3DMM) is provided with corresponding deformation bases. In the initial state, each deformation base corresponds to its own initial parameters. Each deformation base can be modified by inputting parameters. Each deformation base is linearly mixed to obtain the initial three-dimensional deformable model vertex deformation amount, which can be applied to the initial three-dimensional deformable model to deform and generate a new three-dimensional model. The parameters corresponding to each deformation base are the model parameters of the three-dimensional deformable model.
[0044] The initial three-dimensional deformable model corresponding to the three-dimensional model to be generated, that is, the three-dimensional model to be generated and the corresponding initial three-dimensional deformable model are of the same type. For example, if the three-dimensional model to be generated is a three-dimensional model of a human face, then the initial three-dimensional deformable model corresponding to the three-dimensional model to be generated is an initial three-dimensional deformable model of a human face; if the three-dimensional model to be generated is a three-dimensional model of a human body, then the initial three-dimensional deformable model corresponding to the three-dimensional model to be generated is an initial three-dimensional deformable model of the body main body; if the three-dimensional model to be generated is a three-dimensional model of a puppy, then the initial three-dimensional deformable model corresponding to the three-dimensional model to be generated is an initial three-dimensional deformable model of a four-limbed reptile.
[0045] Step S120: Rendering the initial three-dimensional deformable model to obtain a first initial rendered image.
[0046] Specifically, a rendering perspective may be determined, and the initial three-dimensional deformable model may be rendered based on the determined rendering perspective to obtain a first initial rendered image. The rendering perspective may be a randomly selected perspective, for example, a main perspective or a perspective offset by 20 degrees counterclockwise from the main perspective may be randomly selected, or the rendering perspective may be pre-set. The present application does not limit the specific direction of the rendering perspective.
[0047] Optionally, rendering the initial three-dimensional deformable model based on the determined rendering perspective can be implemented according to the following process: based on the determined rendering perspective and the preset texture, the initial three-dimensional deformable model is rendered. Specifically, the preset texture can be pasted on the initial three-dimensional deformable model, and the initial three-dimensional deformable model pasted with the preset texture is rendered. The preset texture can be a gray texture, such as Figure 3 As shown, the first initial rendering image obtained by rendering with a gray map is Figure 3 Rendering result using gray map I grey shown.
[0048] In step S120, the initial three-dimensional deformable model may be subjected to differentiable rendering to obtain a first initial rendered image with better rendering effect.
[0049] Step S130: Based on the matching degree between the first initial rendering image and the description text, adjusting the model parameters of the initial three-dimensional deformable model to generate a three-dimensional structure model that conforms to the description text.
[0050] In this step, the model parameters of the initial three-dimensional deformable model can be adjusted based on the principle of increasing the degree of matching between the first initial rendered image and the description text. Specifically, the model parameters of the initial three-dimensional deformable model can be adjusted based on the principle of making the degree of matching between the first initial rendered image and the description text greater than a first set threshold. In the specific process of adjusting the model parameters, the model parameters of the initial three-dimensional deformable model can be adjusted using a gradient descent method until the adjusted three-dimensional model obtained by adjusting the initial three-dimensional deformable model with the adjusted model parameters satisfies the following condition: the degree of matching between the rendered image corresponding to the adjusted three-dimensional model and the description text is greater than the first set threshold.
[0051] Among them, the gradient descent method is an optimization algorithm that iteratively updates parameters and gradually adjusts parameter values along the gradient direction of the objective function to achieve a solution that meets the conditions.
[0052] In an embodiment of the present application, an estimated image matching the description text can be generated based on the description text, and the degree of match between the first initial rendered image and the description text can be determined based on the difference between the estimated image and the first initial rendered image. The greater the difference between the estimated image and the first initial rendered image, the lower the degree of match between the first initial rendered image and the description text. The degree of match between the first initial rendered image and the description text can also be represented by other methods, which will be explained later. When generating an estimated image matching the description text based on the description text, a text-to-image mapping function can be pre-trained, and the description text can be input into the text-to-image mapping function to output an estimated image matching the description text. The text-to-image mapping function can be trained based on a large number of images marked with description text. The specific training process can refer to the supervised learning algorithm in the relevant technology, which will not be described in detail here.
[0053] It is understandable that the three-dimensional model includes a three-dimensional structure model and a model map pasted on the three-dimensional structure model, wherein the three-dimensional structure model is a three-dimensional model used to reflect the structure shape in the three-dimensional model, and the map is used to show the color, texture and other map appearance of the model. Step S130 can determine the three-dimensional structure model that matches the description text, and the determined three-dimensional structure model can be used for mapping to generate a three-dimensional model that conforms to the actual scene. The determined three-dimensional structure model can also be directly used in virtual makeup and virtual modeling design, so that the user can independently apply the map.
[0054] The three-dimensional model generation method provided in the embodiment of the present application obtains the description text corresponding to the three-dimensional model to be generated and the initial three-dimensional deformable model corresponding to the three-dimensional model to be generated, renders the initial three-dimensional deformable model to obtain a first initial rendered image, and then adjusts the model parameters of the initial three-dimensional deformable model based on the degree of matching between the first initial rendered image and the description text. Since the description text is used to describe the three-dimensional model to be generated, the degree of matching between the first initial rendered image and the description text can reflect the degree of structural difference between the three-dimensional model corresponding to the first initial rendered image and the three-dimensional model to be generated. Therefore, the degree of matching between the first initial rendered image and the description text can guide the adjustment of the model parameters of the initial three-dimensional deformable model, so that the initial three-dimensional deformable model after adjusting the model parameters matches the three-dimensional model described by the description text, thereby generating a three-dimensional structural model that conforms to the description text. Taking the face model as an example, the generated three-dimensional structural model that conforms to the description text is as follows: Figure 3 The optimized face model S is shown in FIG.
[0055] When generating a three-dimensional model, the solution provided in the present application guides the initial three-dimensional deformable model to adjust model parameters based on the description text corresponding to the three-dimensional model to be generated, so as to generate a three-dimensional structural model that conforms to the description text. The initial three-dimensional deformable model corresponding to the three-dimensional model to be generated is usually a model that already conforms to the structural layout of the three-dimensional model. The three-dimensional structural model obtained by adjusting the initial deformable three-dimensional model according to the description text can satisfy the structure described in the description text, and can also make the generated three-dimensional structural model conform to the structural layout of the three-dimensional model, so that the structure of the generated three-dimensional model is more reasonable and the accuracy is higher.
[0056] In addition, the present application adjusts the model parameters of the initial three-dimensional deformable model based on the degree of matching between the first initial rendered image and the description text. Since the rendered image is a two-dimensional image, the feature description of the two-dimensional image is simpler and more convenient than that of the three-dimensional model. Therefore, the degree of matching between the two-dimensional image and the description text can be easily determined, making the process of generating the three-dimensional model simple and easy to implement.
[0057] Since most three-dimensional models do not have corresponding descriptive texts pre-set, the training data required for training mapping models in related technologies requires manual editing of description texts for a large number of three-dimensional models, which has high labor costs and a very cumbersome training process. Since the present application does not need to train the mapping model used to generate the three-dimensional model, there is no need to collect a large number of description texts for the three-dimensional models in advance, which can greatly reduce the manpower spent on the process of generating the three-dimensional model and make the three-dimensional model generation process simpler and more convenient.
[0058] In one implementation, before the above step S130, the following steps S120a to S120b may also be included.
[0059] S120a: Acquire a first initial noise, and use the first initial noise to add noise to the first initial rendered image to obtain a first initial noisy image.
[0060] In this step, the first initial noise intensity may be determined first, and the first initial noise may be determined based on the first initial noise intensity, or a preset noise may be determined as the first initial noise.
[0061] The above-mentioned first initial noise intensity can also be represented by the time step t of the noise. In the embodiment of the present application, given any time step t (between 1 and 0), a Gaussian noise corresponding to the given time step can be obtained, and the closer t is to 1, the greater the noise intensity, and the closer t is to 0, the smaller the noise intensity. The first initial noise intensity can be set randomly. By determining the first initial noise by the noise intensity, the noise corresponding to different rendered images can be made more random, the comprehensiveness of the adjustment of the model parameters can be improved, and the accuracy of the adjusted three-dimensional structure model can be made higher.
[0062] In an embodiment of the present application, a pre-trained diffusion model can be used to add noise to the first initial rendered image. The diffusion model models the diffusion process in which an image is gradually transformed into a noisy image, and the image can be denoised through the diffusion model. Optionally, the diffusion model can have two stages: forward denoising and reverse denoising. In the forward denoising stage, the diffusion model gradually adds noise to the image through a noise function related to the time step to obtain a noisy image; in the reverse denoising stage, the diffusion model can predict the noise in the noisy image at the current time step and denoise it. The training process of the diffusion model can be obtained by conventional training methods, and the training data in the training process can include sample images, and the sample images are provided with corresponding noisy images and the noise of the noisy images.
[0063] S120b: Using the description text corresponding to the to-be-generated three-dimensional model as a guide condition, predicting the noise of the first initial noisy image to obtain a first predicted noise.
[0064] Optionally, when step S120a determines the first initial noise based on the first initial noise intensity, step S120b may predict the noise of the first initial noisy image based on the description text of the 3D model to be generated and the first initial noise intensity to obtain the first predicted noise.
[0065] In this embodiment, a noise prediction model for predicting noise under the guidance of a guide text can be pre-trained. The noise prediction model can be a text-image diffusion model obtained after training the text guidance function on the basis of the above-mentioned diffusion model, wherein the text-image diffusion model is used to predict noise for images under the guidance of text. The training data in the process of training the text-image diffusion model can include a sample description text corresponding to the sample image, and the sample image is provided with a corresponding sample noise image and a corresponding sample noise. The sample description text and the sample noise image are input into the model to be trained, and the sample predicted noise is output. The parameters of the model to be trained are adjusted according to the difference between the sample predicted noise and the actual sample noise to obtain the text-image diffusion model. The specific training process of the text-image diffusion model is not described in detail in this application. The specific training process can refer to the method of model training in related technologies.
[0066] In a specific embodiment, step S120b can be implemented according to the following steps S121b-1 to S121b-2.
[0067] Step S121b-1: obtaining a pre-trained first image-text diffusion model, where the first image-text diffusion model is used to perform noise prediction on a feature image that conforms to an object in an actual scene under the guidance of a guide text.
[0068] The above-mentioned images that conform to the presentation characteristics of objects in actual scenes may be, for example, actually taken photos, actually rendered three-dimensional models, etc.
[0069] The method of training the first image-text diffusion model can refer to the method of training the image-text diffusion model above, which will not be repeated here. The sample images in the training data used to train the first image-text diffusion model are images that conform to the characteristics of objects in actual scenes.
[0070] Step S121b-2: inputting the description text corresponding to the to-be-generated three-dimensional model, the first initial noise intensity and the first initial noisy image into the first image-text diffusion model to obtain the first predicted noise of the first initial noisy image.
[0071] This embodiment can conveniently determine the first predicted noise corresponding to the first initial noisy image through the pre-trained first image-text diffusion model.
[0072] Optionally, before step S121b-2, the first initial noisy image can be digitally encoded to obtain a first initial noisy image code. In step S121b-2, the first initial noisy image code can replace the first initial noisy image, that is, step S121b-2 can be implemented according to the following steps: the description text corresponding to the three-dimensional model to be generated, the above-mentioned first initial noise intensity and the above-mentioned first initial noisy image code are input into the first image-text diffusion model to obtain the first predicted noise of the first initial noisy image.
[0073] The above-mentioned numerical encoding can use a pre-trained image encoder, and the specific training process will not be described in detail.
[0074] Step S130 may adjust the model parameters of the initial three-dimensional deformable model according to the following step S131.
[0075] Step S131: Based on the difference between the first predicted noise and the first initial noise, a gradient descent method is used to adjust the model parameters of the initial three-dimensional deformable model to generate a three-dimensional structure model that conforms to the description text.
[0076] When the difference between the first predicted noise and the first initial noise is relatively large, it indicates that the matching degree between the first initial rendered image and the description text corresponding to the three-dimensional model to be generated is low, and step S131 indicates the matching degree between the first initial rendered image and the description information corresponding to the three-dimensional model to be generated by the difference between the first predicted noise and the first initial noise. Step S131 can specifically adjust the model parameters of the initial three-dimensional deformable model by using a gradient descent method based on the principle that the difference between the first predicted noise and the first initial noise is less than or equal to a first preset threshold value to generate a three-dimensional structure model that conforms to the description text.
[0077] Specifically, based on the principle of making the difference between the first predicted noise and the first initial noise less than or equal to a first preset threshold, the gradient descent method is used to adjust the model parameters of the initial three-dimensional deformable model to generate a three-dimensional structure model that conforms to the above description text. This can be achieved by following steps A to G.
[0078] Step A: Let i=1.
[0079] Step A assigns i a value of 1.
[0080] Step B: When it is detected that the difference between the i-th predicted noise and the i-th input noise is greater than a first preset threshold, based on the principle of making the difference between the i-th predicted noise and the i-th input noise smaller than the first preset threshold, the model parameters of the three-dimensional deformable model updated for the i-1th time are adjusted to obtain the i-th adjustment parameters.
[0081] Among them, the first predicted noise is the first predicted noise, the first input noise is the first initial noise, and the three-dimensional variability model updated for the 0th time is the initial three-dimensional deformable model.
[0082] When i is 1, step B, that is, when it is detected that the difference between the first predicted noise and the first initial noise is greater than the first preset threshold, the model parameters of the initial three-dimensional deformable model are adjusted based on the principle of making the difference between the first predicted noise and the first initial noise less than the first preset threshold, to obtain the first adjusted parameters.
[0083] When adjusting the model parameters of the three-dimensional deformable model updated for the i-1th time, the model parameters of the three-dimensional deformable model updated for the i-1th time can be determined based on a batch gradient descent method, a stochastic gradient descent method, etc. The specific model parameter tuning method will not be described in detail here.
[0084] Step C: using the i-th adjustment parameters to adjust the three-dimensional deformable model updated for the i-1th time to obtain the i-th updated three-dimensional deformable model, and rendering the i-th updated three-dimensional deformable model to obtain the i-th updated rendering image.
[0085] When i is 1, step C uses the first adjustment parameters to adjust the initial three-dimensional deformable model to obtain a first updated three-dimensional deformable model, and renders the first updated three-dimensional deformable model to obtain a first updated rendered image.
[0086] Step D: Obtain the i+1th input noise, and use the i+1th input noise to add noise to the i-th updated rendering image to obtain the i-th updated noisy image.
[0087] When i is 1, step D is to obtain the second input noise. The method for obtaining the second input noise is similar to the method for obtaining the first initial noise. The second input noise can be determined by randomly setting the time step and according to the time step. The subsequent methods for obtaining the third, fourth, ... input noises are similar to the method for obtaining the second input noise.
[0088] Step E: using the description text corresponding to the to-be-generated three-dimensional model as a guide condition, predicting the noise of the i-th updated noisy image to obtain the (i+1)-th predicted noise.
[0089] When i is 1, step E uses the description text corresponding to the to-be-generated three-dimensional model as a guide condition to predict the noise of the first updated noisy image to obtain the second predicted noise.
[0090] Step F: set i=i+1, and execute steps B to F until it is detected that the difference between the i-th predicted noise and the i-th input noise is less than or equal to a first preset threshold, and the three-dimensional deformable model updated for the i-1th time is determined as a three-dimensional structural model that conforms to the description text.
[0091] When i in step E is 1, step F will assign i a value of 2 and execute steps B to F. If it is detected that the difference between the second predicted noise and the second input noise is less than or equal to the first preset threshold, the three-dimensional deformable model updated for the first time is determined as a three-dimensional structural model that conforms to the above description text.
[0092] The above embodiment introduces in detail the process of adjusting the model parameters of the initial three-dimensional deformable model by the gradient descent method. Through the specific process introduced in this embodiment, a three-dimensional structure model that conforms to the description text can be easily generated.
[0093] This embodiment reflects the degree of matching between the rendered image and the description text by measuring the difference between the predicted noise of the noisy image corresponding to the rendered image under the guidance of the description text and the real noise used in the noisy process of the noisy image corresponding to the rendered image. The degree of matching between the rendered image and the description text can be conveniently, accurately and efficiently determined.
[0094] The execution process of the above steps S110 to S130 can refer to Figure 5 Steps 31 to 39 in are implemented.
[0095] In one embodiment, if Figure 2 As shown, the above three-dimensional model generation method may further include the following steps S140 to S160.
[0096] Step S140: obtaining an initial model map corresponding to the three-dimensional model to be generated.
[0097] In the embodiments of the present application, various initial model maps can be pre-set, for example, an initial model map for a face, an initial model map for a hand, an initial model map for a body, an initial model map for a four-limbed reptile, an initial model map for a human body, etc. Those skilled in the art can set various initial model maps according to specific design requirements, which is not specifically limited in the present application.
[0098] The initial model texture corresponding to the 3D model to be generated, that is, the 3D model to be generated is of the same type as the corresponding initial model. For example, if the 3D model to be generated is a face 3D model, then the initial model texture corresponding to the 3D model to be generated is a face initial model texture; if the 3D model to be generated is a body 3D model, then the initial model texture corresponding to the 3D model to be generated is a body initial model texture. Figure 3As shown, when the 3D model to be generated is a face model, Figure 3 The initialized face color map d in rgb That is, an example image of the initial model map corresponding to the three-dimensional model to be generated.
[0099] Step S150: Rendering the three-dimensional structure model based on the initial model map to obtain a second initial rendering image.
[0100] Specifically, the initial model map can be pasted on the three-dimensional structure model generated in step S130, and the three-dimensional structure model pasted with the initial model map can be rendered at a second rendering perspective to obtain a second initial rendered image. The second rendering perspective can be the same as or different from the rendering perspective mentioned in step S120 above. The second rendering perspective can be a randomly determined perspective or a manually pre-set perspective, which is not specifically limited in this application. Taking a face model as an example, Figure 3 Use of d rgb Rendering Result I D That is, an example image of the obtained second initial rendering image.
[0101] Step S160: adjusting the initial model texture based on the texture adjustment information to generate a three-dimensional model texture that matches the description text, wherein the texture adjustment information includes a matching degree between the second initial rendering image and the description text.
[0102] The three-dimensional model map generated in step S160 is used to be pasted on the above-mentioned three-dimensional structural model to obtain a three-dimensional model that conforms to the above-mentioned description text. Step S160 is to adjust the initial model map based on the degree of matching between the second initial rendering image and the description text corresponding to the three-dimensional model to be generated. Specifically, the initial model map can be adjusted based on the principle of increasing the degree of matching between the second initial rendering image and the description text corresponding to the three-dimensional model to be generated. For example, the initial model map can be adjusted based on the principle that the degree of matching between the second initial rendering image and the description text corresponding to the three-dimensional model to be generated is greater than a second set threshold. In the process of adjusting the initial model map, the gradient descent method can be used to adjust the initial model map until the adjusted model map is pasted on the three-dimensional structural model and the following conditions are met: the degree of matching between the rendered image corresponding to the three-dimensional structural model pasted with the adjusted model map and the description text is greater than the second set threshold.
[0103] The manner in which the degree of match between the second initial rendered image and the description text is determined in step S160 and the manner in which the initial model map is adjusted using the gradient descent method can refer to the manner in which the degree of match between the first initial rendered image and the description text is determined in the above step S130, and the repeated parts will not be described in detail here.
[0104] In this embodiment, the initial model map is guided to adjust the image based on the description text corresponding to the three-dimensional model to be generated, so as to generate a model map that conforms to the description text. This embodiment divides the structure model generation stage and the model map generation stage into two stages to generate the corresponding three-dimensional structure model and three-dimensional model map respectively. Compared with generating the entire three-dimensional model in one go, a more accurate three-dimensional model can be generated. This embodiment generates a model map that matches the description text, and the generated model map can be pasted on the three-dimensional structure model to generate a complete three-dimensional model that matches the description text as a whole. Figure 3 As shown, Figure 3 The optimized face color map d rgb That is, the generated three-dimensional model map conforming to the description text, Figure 3 The optimized rendering result is the rendering result of the three-dimensional model obtained by pasting the three-dimensional model map on the above three-dimensional structure model.
[0105] In one implementation, before step S160, the three-dimensional model generation method may further include the following steps S160a to S160b.
[0106] Step S160a: obtaining a second initial noise, and using the second initial noise to add noise to the second initial rendered image to obtain a second initial noisy image.
[0107] Step S160b: using the description text corresponding to the to-be-generated three-dimensional model as a guide condition, predicting the noise of the second initial noisy image to obtain a second predicted noise.
[0108] Optionally, before step S160a, the following steps may be further included: determining a second initial noise intensity; determining a second initial noise based on the second initial noise intensity. Step S160b may be implemented by the following steps: using the description text and the second initial noise intensity as guide conditions, predicting the noise of the second initial noisy image to obtain a second predicted noise.
[0109] Specifically, step S160b can be implemented according to the following steps S160b-1 to S160b-2.
[0110] Step S160b-1: obtaining a pre-trained first image-text diffusion model, where the first image-text diffusion model is used to perform noise prediction on a feature image that conforms to an object in an actual scene under the guidance of a guide text.
[0111] Step S160b-2: input the description text, the second initial noise intensity, and the second initial noisy image into the first image-text diffusion model to obtain a second predicted noise of the second initial noisy image.
[0112] The first image-text diffusion model in steps S160b-1 and S160b-2 may be the same model as the first image-text diffusion model in steps S121b-1 and S121b-2. The execution process of steps S160b-1 and S160b-2 may refer to steps S121b-1 and S121b-2, and the similarities will not be described in detail.
[0113] Correspondingly, the degree of matching between the second initial rendered image and the description text corresponding to the three-dimensional model to be generated includes the difference between the second predicted noise and the second initial noise. That is to say, the degree of matching between the second initial rendered image and the description text corresponding to the three-dimensional model to be generated can be represented by the difference between the second predicted noise and the second initial noise. The greater the difference between the second predicted noise and the second initial noise, the lower the degree of matching between the second initial rendered image and the description text corresponding to the three-dimensional model to be generated.
[0114] The process of predicting the second predicted noise in steps S160a and S160b may refer to the process of predicting the first predicted noise in steps S120a and S120b, and the repeated parts will not be described in detail.
[0115] In one embodiment, before step S160, the three-dimensional model generation method may further include the following steps S160c:
[0116] Step S160c: Acquire the texture features of the three-dimensional model.
[0117] Correspondingly, the texture adjustment information may also include: the degree of matching between the initial model texture and the texture feature. That is, when adjusting the initial model texture in step S160, the matching degree between the second initial rendering image and the description text corresponding to the 3D model to be generated is based on the matching degree between the initial model texture and the texture feature.
[0118] The mapping of a three-dimensional model can be understood as a planar image formed by unfolding the mapping of the surface of the three-dimensional model. The mapping of a three-dimensional model is usually different from a conventional image. In the process of adjusting the initial model mapping, in order to make the adjusted model mapping more consistent with the mapping features, the mapping features of the three-dimensional model can be obtained first, so that when the initial model mapping is subsequently adjusted, the adjusted model mapping can be matched with the mapping features and the adjustment basis, so that the adjusted three-dimensional model mapping is more consistent with the characteristics of the three-dimensional model mapping, which can better avoid the occurrence of breakage, wrinkles and the like after overlaying the mapping, thereby improving the accuracy of the generated three-dimensional model.
[0119] The mapping features of the three-dimensional model may be the mapping features of the three-dimensional model edited manually in advance, or may be the mapping features determined by the electronic device based on a large number of existing model mappings.
[0120] In a specific embodiment, before step S160, the three-dimensional model generation method may further include the following steps S160d to S160e.
[0121] Step S160d: Obtain a third initial noise, and use the third initial noise to add noise to the initial model map to obtain a third initial noisy image.
[0122] Step S160e: using the description text corresponding to the texture feature as a guide condition, predicting the noise of the third initial noisy image to obtain a third predicted noise.
[0123] The description text corresponding to the texture feature is the description text used to describe the texture feature of the three-dimensional model. The process of predicting the third predicted noise in steps S160d to S160e can refer to the process of predicting the second predicted noise in steps S160a to S160b, and the similarities will not be described in detail.
[0124] Correspondingly, the degree of matching between the initial model map and the map feature may include: the difference between the third predicted noise and the third initial noise, that is, the degree of matching between the initial model map and the map feature may be represented by the difference between the third predicted noise and the third initial noise, and the greater the difference between the third predicted noise and the third initial noise, the lower the degree of matching between the initial model map and the map feature.
[0125] In a specific embodiment, before step S160e, the 3D model generation method may further include: obtaining a pre-trained second image-text diffusion model, where the second image-text diffusion model is used to perform noise prediction on the 3D model map under the guidance of the guide text.
[0126] Step S160e can be implemented by the following steps: inputting the description text corresponding to the map feature and the third initial noisy image into the second image-text diffusion model to obtain a third predicted noise.
[0127] The training method of the second image-text diffusion model is similar to that of the above-mentioned first image-text diffusion model. The difference is that the training data used in the training process of the first image-text diffusion model are sample images that conform to the presentation characteristics of objects in actual scenes (such as photos, video screenshots, three-dimensional model rendering images, etc.), while the training data used in the training process of the second image-text diffusion model are sample images that conform to the characteristics of three-dimensional model maps (such as maps of existing three-dimensional models, etc.), thereby making the second image-text diffusion model more suitable for predicting the noise in the maps of three-dimensional models.
[0128] Specifically, the above step S160e can first digitally encode the third initial noisy image to obtain the third initial noisy image code. Specifically, step S160e can input the description text corresponding to the map feature and the third initial noisy image code into the second image-text diffusion model to obtain the third predicted noise.
[0129] Optionally, the second image-text diffusion model can be trained through the following steps a to b.
[0130] Step a: Obtain a first image-text diffusion model and training samples, wherein the first image-text diffusion model is used to predict noise for images that conform to the characteristics of objects in actual scenes under the guidance of guide text, and the training samples include sample three-dimensional model maps.
[0131] The first image-text diffusion model in step a may be the first image-text diffusion model mentioned above. The sample three-dimensional model map is the model map corresponding to each existing three-dimensional model.
[0132] Step b: taking the description text corresponding to the texture feature as a guide condition, using the sample three-dimensional model texture to adjust the parameters of the first texture diffusion model to obtain a second texture diffusion model.
[0133] The execution process of steps a and b is Figure 4 Step 22, Step 24, Step 26, Step 27 generate process, Figure 4 d in rgb That is, any sample 3D model map, That is the second image and text diffusion model.
[0134] In this embodiment, when the third predicted noise is predicted by the second image-text diffusion model, the description text corresponding to the map feature input into the second image-text diffusion model in step S160e when predicting the third predicted noise can be the same as the description text corresponding to the map feature used to train the second image-text diffusion model. This is because, if a certain description text is used to represent the map feature during the training of the second image-text diffusion model, the second image-text diffusion model will establish a correspondence between the description text used during training and each correct sample model map, that is, the second image-text diffusion model will consider that the description text used during training is the correct description of each correct sample model map. Therefore, when the second image-text diffusion model is subsequently used, the description text used during training is used as the description text of the map feature, and the second image-text diffusion model can be guided to predict noise with the guide text that conforms to the map feature.
[0135] The description text corresponding to the map feature used in training the second image-text diffusion model can be set arbitrarily, for example, it can be the text of "rgb face texture" or the text of "map feature", which is not specifically limited in this application.
[0136] Specifically, Figure 4 As shown in step 21 in FIG. 1 , the description text corresponding to the texture feature (e.g., “rgb face texture”) can be first encoded using a text-image embedding model (CLIP) to obtain the text encoding y corresponding to the texture feature after encoding. rgb-tex , and then y rgb-tex As a guide condition, the sample three-dimensional model map is used to adjust the parameters of the first image-text diffusion model to obtain a second image-text diffusion model.
[0137] In an embodiment of the present application, the above-mentioned sample three-dimensional model map is provided with a corresponding sample noise map. In step b, when adjusting the parameters of the first image-text diffusion model, the sample noise map and the description text corresponding to the map features can be input into the first image-text diffusion model whose parameters are to be adjusted, and a sample denoised image is output. According to the gap between the sample denoised image and the corresponding sample three-dimensional model map, the parameters of the first image-text diffusion model are adjusted based on the principle that the gap is smaller than a preset gap. When the number of adjustment iterations reaches a preset number, a second image-text diffusion model is obtained.
[0138] This embodiment makes adjustments based on the existing first image-text diffusion model to obtain the above-mentioned second image-text diffusion model, and can quickly obtain an intelligent model that can accurately predict the model map noise.
[0139] In one implementation, before step S160, the three-dimensional model generation method may further include the following steps S160f to S160g.
[0140] Step S160f: converting the above initial model map into an initial brightness format map in brightness format.
[0141] The above-mentioned initial model texture can be a model texture in a color format such as an RGB format texture, a CMYK format texture, etc. The initial model texture can reflect the appearance of the texture, and the brightness format can be a YUV format. The brightness format texture can reflect the different brightness of different areas in the texture, thereby simulating the different brightness conditions of different areas reflected by light irradiating the surface of the three-dimensional model. Figure 3 As shown, Figure 3 The texture d converted to yuv space yuv This is an example of an initial brightness format map.
[0142] Step S160g: Obtain brightness format map features of the three-dimensional model.
[0143] The brightness format map features of the three-dimensional model may be features pre-edited manually, or may be map features determined by the electronic device based on brightness format maps corresponding to a large number of existing model maps.
[0144] Correspondingly, the texture adjustment information may also include: the degree of matching between the initial brightness format texture and the brightness format texture feature. That is, when adjusting the initial model texture in step S160, the matching degree between the second initial rendering image and the description text corresponding to the 3D model to be generated is based on the matching degree between the initial model texture and the texture feature, and the matching degree between the initial brightness format texture and the brightness format texture feature.
[0145] When adjusting the initial model texture, this real-time method can make the adjusted model texture more consistent with the brightness expression mode of the three-dimensional model texture based on the brightness format texture feature, thereby making the generated three-dimensional model more realistic and accurate.
[0146] In a specific embodiment, before step S160, the three-dimensional model generation method may further include the following steps S160h to S160i.
[0147] Step S160h: Obtain a fourth initial noise, and use the fourth initial noise to add noise to the initial brightness format map to obtain an initial noisy brightness image.
[0148] Step S160i: using the description text corresponding to the brightness format map feature as a guide condition, predicting the noise of the initial noisy brightness image to obtain a fourth predicted noise.
[0149] The process of predicting the fourth predicted noise in step S160h to step S160i is similar to the process of predicting the third predicted noise in step S160d to step S160e, and will not be described in detail here.
[0150] Correspondingly, the matching degree between the initial brightness format map and the brightness format map features includes: the difference between the fourth predicted noise and the initial noisy brightness image. That is, the matching degree between the initial brightness format map and the brightness format map features can be represented by the difference between the fourth predicted noise and the fourth initial noise, and the greater the difference between the fourth predicted noise and the fourth initial noise, the lower the matching degree between the initial brightness format map and the brightness format map features.
[0151] In a specific embodiment, before step S160i, the method may further include: obtaining a pre-trained third image-text diffusion model, wherein the third image-text diffusion model is used to perform noise prediction on the three-dimensional model map in brightness format under the guidance of the guide text.
[0152] Step S160i can be implemented by the following steps: inputting the description text corresponding to the brightness format map feature and the initial noisy brightness image into the third image-text diffusion model to obtain a fourth predicted noise.
[0153] The third image-text diffusion model can be trained in the following manner: converting the sample three-dimensional model map into a sample brightness map in a brightness format; using the description text corresponding to the brightness format map feature as a guide condition, and using the sample brightness map to adjust the parameters of the first image-text diffusion model to obtain the third image-text diffusion model.
[0154] The training method of the third image-text diffusion model is similar to that of the above-mentioned second image-text diffusion model. The difference is that the third image-text diffusion model is obtained by training based on sample brightness maps converted into brightness format, and the description text corresponding to the brightness format map features in the training process of the third image-text diffusion model is different from that of the second image-text diffusion model.
[0155] In this embodiment, when the fourth predicted noise is predicted by the third image-text diffusion model, the description text corresponding to the brightness format map feature input into the third image-text diffusion model in step S160i when predicting the fourth predicted noise can be the same as the description text corresponding to the brightness format map feature used to train the third image-text diffusion model.
[0156] The description text corresponding to the brightness format map feature used in training the third image-text diffusion model can be set arbitrarily, for example, it can be the text of "yuv face texture" or the text of "brightness map feature", which is not specifically limited in this application.
[0157] Specifically, Figure 4 As shown in step 21 in , the description text corresponding to the brightness format map feature (e.g., "yuv face texture") can be first encoded using the image-text embedding model (CLIP) to obtain the text encoding y corresponding to the brightness format map feature after encoding yuv-tex , and then y yuv-tex As a guide condition, the sample brightness map is used to adjust parameters of the first image-text diffusion model to obtain a third image-text diffusion model.
[0158] Specifically, when adjusting the initial model map, the initial model map may be adjusted using the following loss function as the target loss.
[0159] L=L D +L d-rgb +L d-yuv
[0160] Among them, L is the target loss, LD is the difference between the second predicted noise and the second initial noise, L d-rgb is the difference between the third predicted noise and the third initial noise, L d-yuv The difference between the fourth predicted noise and the fourth initial noise.
[0161] During the optimization process, L is optimized by the gradient descent method to obtain a model map that conforms to the text description.
[0162] Figure 4 Step 23, Step 25, Step 26, Step 27 in The process is the process of optimizing and generating the third image-text diffusion model. Figure 4 The optimization convergence condition is whether the number of iterations reaches the maximum number of iterations.
[0163] The execution process of the above steps S140 to S160 can refer to Figure 6 Steps 41 to 412 in are implemented.
[0164] Corresponding to the three-dimensional model generation method provided in the embodiment of the present application, the embodiment of the present application also provides a three-dimensional model generation device, such as Figure 7 As shown, the three-dimensional model generation device provided in this embodiment includes:
[0165] An acquisition unit, used to acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated;
[0166] A rendering unit, configured to render the initial three-dimensional deformable model to obtain a first initial rendered image;
[0167] An adjustment unit is used to adjust the model parameters of the initial three-dimensional deformable model based on the matching degree between the first initial rendering image and the description text, so as to generate a three-dimensional structure model that conforms to the description text.
[0168] Optionally, the device further comprises:
[0169] a noise adding unit, configured to obtain a first initial noise, and use the first initial noise to add noise to the first initial rendered image to obtain a first initial noisy image;
[0170] A prediction unit, configured to predict the noise of the first initial noisy image based on the description text as a guide condition, to obtain a first predicted noise;
[0171] The adjustment unit is specifically used for:
[0172] Based on the difference between the first predicted noise and the first initial noise, a gradient descent method is used to adjust the model parameters of the initial three-dimensional deformable model to generate a three-dimensional structure model that conforms to the description text.
[0173] Optionally, the device further comprises:
[0174] a determining unit, configured to determine a first initial noise intensity, and determine a first initial noise based on the first initial noise intensity;
[0175] The prediction unit is specifically used for:
[0176] The description text and the first initial noise intensity are used as guiding conditions to predict the noise of the first initial noisy image to obtain a first predicted noise.
[0177] Optionally, the adjustment unit is specifically used for:
[0178] Let i = 1;
[0179] When it is detected that the difference between the i-th predicted noise and the i-th input noise is greater than a first preset threshold, based on the principle of making the difference between the i-th predicted noise and the i-th input noise smaller than the first preset threshold, the model parameters of the three-dimensional deformable model updated for the i-1th time are adjusted to obtain the i-th adjustment parameters, wherein the first predicted noise is the first predicted noise, the first input noise is the first initial noise, and the three-dimensional deformable model updated for the 0th time is the initial three-dimensional deformable model;
[0180] Using the i-th adjustment parameter to adjust the three-dimensional deformable model updated for the (i-1)th time to obtain the i-th updated three-dimensional deformable model, and rendering the i-th updated three-dimensional deformable model to obtain the i-th updated rendered image;
[0181] Obtaining the (i+1)th input noise, and using the (i+1)th input noise to add noise to the (i)th updated rendered image, to obtain the (i)th updated noisy image;
[0182] Using the description text as a guide condition, predicting the noise of the i-th updated noisy image to obtain the (i+1)-th predicted noise;
[0183] Let i=i+1, and execute the step of adjusting the model parameters of the three-dimensional deformable model updated for the i-1th time based on the principle of making the difference between the i-th predicted noise and the i-th input noise smaller than the first preset threshold when it is detected that the difference between the i-th predicted noise and the i-th input noise is greater than the first preset threshold, until it is detected that the difference between the i-th predicted noise and the i-th input noise is less than or equal to the first preset threshold, and determining the three-dimensional deformable model updated for the i-1th time as a three-dimensional structural model that conforms to the description text.
[0184] Optionally, the prediction unit is specifically used for:
[0185] Acquire a pre-trained first image-text diffusion model, where the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of an object in an actual scene under the guidance of a guide text;
[0186] The description text, the first initial noise intensity, and the first initial noisy image are input into the first image-text diffusion model to obtain a first predicted noise of the first initial noisy image.
[0187] Optionally, the acquisition unit is further used to: acquire an initial model map corresponding to the three-dimensional model to be generated;
[0188] The rendering unit is further used to: render the three-dimensional structure model based on the initial model map to obtain a second initial rendering image;
[0189] The adjustment unit is also used to: adjust the initial model texture based on the texture adjustment information to generate a three-dimensional model texture that conforms to the description text, the three-dimensional model texture is used to be pasted on the three-dimensional structure model to obtain a three-dimensional model that conforms to the description text, and the texture adjustment information includes the degree of matching between the second initial rendered image and the description text.
[0190] Optionally, the noise adding unit is further used to: obtain a second initial noise, and use the second initial noise to add noise to the second initial rendered image to obtain a second initial noisy image;
[0191] The prediction unit is further used to: predict the noise of the second initial noisy image based on the description text as a guide condition to obtain a second predicted noise;
[0192] The degree of matching between the second initial rendered image and the description text includes a difference between the second predicted noise and the second initial noise.
[0193] Optionally, the determining unit is further configured to: determine a second initial noise intensity; determine a second initial noise based on the second initial noise intensity;
[0194] The prediction unit is specifically used to: obtain a pre-trained first image-text diffusion model, the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of objects in an actual scene under the guidance of a guide text; input the description text, the second initial noise intensity and the second initial noisy image into the first image-text diffusion model to obtain a second predicted noise of the second initial noisy image.
[0195] Optionally, the acquisition unit is further used to: acquire a mapping feature of the three-dimensional model;
[0196] The texture adjustment information also includes: the matching degree between the initial model texture and the texture feature.
[0197] Optionally, the noise adding unit is further used to: obtain a third initial noise, and use the third initial noise to add noise to the initial model map to obtain a third initial noisy image;
[0198] The prediction unit is further used to: predict the noise of the third initial noisy image based on the description text corresponding to the map feature as a guide condition to obtain a third predicted noise;
[0199] The degree of matching between the initial model map and the map feature includes: a difference between the third predicted noise and the third initial noise.
[0200] Optionally, the acquisition unit is further used to: acquire a pre-trained second image-text diffusion model, where the second image-text diffusion model is used to perform noise prediction on the three-dimensional model map under the guidance of the guide text;
[0201] The prediction unit is specifically used to: input the description text corresponding to the map feature and the third initial noisy image into the second image-text diffusion model to obtain a third predicted noise.
[0202] Optionally, the second image-text diffusion model is trained in the following manner:
[0203] Acquire a first image-text diffusion model and a training sample, wherein the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of an object in an actual scene under the guidance of a guide text, and the training sample includes a sample three-dimensional model map;
[0204] The description text corresponding to the texture feature is used as a guide condition, and the sample three-dimensional model texture is used to adjust the parameters of the first texture diffusion model to obtain a second texture diffusion model.
[0205] Optionally, the device further comprises:
[0206] A conversion unit, used for converting the initial model map into an initial brightness format map in a brightness format;
[0207] The acquisition unit is also used to: acquire brightness format map features of the three-dimensional model;
[0208] The texture adjustment information also includes: a matching degree between the initial brightness format texture and the brightness format texture features.
[0209] Optionally, the noise adding unit is further used to: obtain a fourth initial noise, and use the fourth initial noise to add noise to the initial brightness format map to obtain an initial noisy brightness image;
[0210] The prediction unit is further used to: predict the noise of the initial noisy brightness image using the description text corresponding to the brightness format map feature as a guide condition to obtain a fourth predicted noise;
[0211] The degree of matching between the initial luminance format map and the luminance format map features comprises: a difference between the fourth predicted noise and the initial noisy luminance image.
[0212] Optionally, the acquisition unit is further used to: acquire a pre-trained third image-text diffusion model, wherein the third image-text diffusion model is used to perform noise prediction on the three-dimensional model map in brightness format under the guidance of the guide text;
[0213] The prediction unit is specifically used to: input the description text corresponding to the brightness format map feature and the initial noisy brightness image into the third image-text diffusion model to obtain a fourth predicted noise.
[0214] Optionally, the third image-text diffusion model is trained in the following manner:
[0215] Converting the sample three-dimensional model map into a sample brightness map in a brightness format;
[0216] The description text corresponding to the brightness format map feature is used as a guide condition, and the sample brightness map is used to adjust the parameters of the first image-text diffusion model to obtain a third image-text diffusion model.
[0217] Corresponding to the method for generating a three-dimensional model provided in the embodiment of the present application, the embodiment of the present application also provides an electronic device. Figure 8 As shown, the electronic device includes: a processor 401; and a memory 402, which is used to store a program of a method for generating a three-dimensional model. After the electronic device is powered on and the program storing the method for generating a three-dimensional model is run by the processor, the following steps are performed:
[0218] Acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated;
[0219] Rendering the initial three-dimensional deformable model to obtain a first initial rendered image;
[0220] Based on the matching degree between the first initial rendered image and the description text, the model parameters of the initial three-dimensional deformable model are adjusted to generate a three-dimensional structure model that conforms to the description text.
[0221] Corresponding to the method for generating a three-dimensional model provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium storing a program for the three-dimensional model generating method, the program is executed by a processor to perform the following steps:
[0222] Acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated;
[0223] Rendering the initial three-dimensional deformable model to obtain a first initial rendered image;
[0224] Based on the matching degree between the first initial rendered image and the description text, the model parameters of the initial three-dimensional deformable model are adjusted to generate a three-dimensional structure model that conforms to the description text.
[0225] It should be noted that for the detailed description of the device, electronic device and computer-readable storage medium provided in the embodiments of the present application, reference can be made to the relevant description of the method in the first embodiment of the present application, which will not be repeated here.
[0226] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0227] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0228] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0229] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory (RAM) of other attributes, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage media or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0230] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0231] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A three-dimensional model generation method, characterized in that: The method comprises: Acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated; Rendering the initial three-dimensional deformable model to obtain a first initial rendered image; Based on the matching degree between the first initial rendered image and the description text, the model parameters of the initial three-dimensional deformable model are adjusted to generate a three-dimensional structure model that conforms to the description text.
2. The method according to claim 1, characterized in that: Before adjusting the model parameters of the initial three-dimensional deformable model based on the matching degree between the first initial rendered image and the description information, the method further includes: Acquire a first initial noise, and use the first initial noise to add noise to the first initial rendered image to obtain a first initial noisy image; Using the description text as a guide condition, predicting the noise of the first initial noisy image to obtain a first predicted noise; The adjusting the model parameters of the initial three-dimensional deformable model based on the matching degree between the first initial rendered image and the description information to generate a three-dimensional structure model that conforms to the description text includes: Based on the difference between the first predicted noise and the first initial noise, a gradient descent method is used to adjust the model parameters of the initial three-dimensional deformable model to generate a three-dimensional structure model that conforms to the description text.
3. The method according to claim 2, characterized in that Before obtaining the first initial noise, the method further includes: determining a first initial noise intensity; determining a first initial noise based on the first initial noise intensity; The step of predicting the noise of the first initial noisy image based on the description text as a guide condition to obtain a first predicted noise includes: The description text and the first initial noise intensity are used as guiding conditions to predict the noise of the first initial noisy image to obtain a first predicted noise.
4. The method according to claim 2, characterized in that: The step of adjusting the model parameters of the initial three-dimensional deformable model by using a gradient descent method based on the difference between the first predicted noise and the first initial noise to generate a three-dimensional structure model that conforms to the description text includes: Let i = 1; When it is detected that the difference between the i-th predicted noise and the i-th input noise is greater than a first preset threshold, based on the principle of making the difference between the i-th predicted noise and the i-th input noise smaller than the first preset threshold, the model parameters of the three-dimensional deformable model updated for the i-1th time are adjusted to obtain the i-th adjustment parameters, wherein the first predicted noise is the first predicted noise, the first input noise is the first initial noise, and the three-dimensional deformable model updated for the 0th time is the initial three-dimensional deformable model; Using the i-th adjustment parameter to adjust the three-dimensional deformable model updated for the (i-1)th time to obtain the i-th updated three-dimensional deformable model, and rendering the i-th updated three-dimensional deformable model to obtain the i-th updated rendered image; Obtaining the (i+1)th input noise, and using the (i+1)th input noise to add noise to the (i)th updated rendered image, to obtain the (i)th updated noisy image; Using the description text as a guide condition, predicting the noise of the i-th updated noisy image to obtain the (i+1)-th predicted noise; Let i=i+1, and execute the step of adjusting the model parameters of the three-dimensional deformable model updated for the i-1th time based on the principle of making the difference between the i-th predicted noise and the i-th input noise smaller than the first preset threshold when it is detected that the difference between the i-th predicted noise and the i-th input noise is greater than the first preset threshold, until it is detected that the difference between the i-th predicted noise and the i-th input noise is less than or equal to the first preset threshold, and determining the three-dimensional deformable model updated for the i-1th time as a three-dimensional structural model that conforms to the description text.
5. The method according to claim 3, characterized in that: The step of predicting the noise of the first initial noisy image based on the description text and the first initial noise intensity to obtain a first predicted noise includes: Acquire a pre-trained first image-text diffusion model, where the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of an object in an actual scene under the guidance of a guide text; The description text, the first initial noise intensity, and the first initial noisy image are input into the first image-text diffusion model to obtain a first predicted noise of the first initial noisy image.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining an initial model map corresponding to the three-dimensional model to be generated; Based on the initial model map, the three-dimensional structure model is rendered to obtain a second initial rendering image; The initial model texture is adjusted based on the texture adjustment information to generate a three-dimensional model texture that conforms to the description text. The three-dimensional model texture is used to be pasted on the three-dimensional structural model to obtain a three-dimensional model that conforms to the description text. The texture adjustment information includes the degree of matching between the second initial rendered image and the description text.
7. The method according to claim 6, characterized in that Before adjusting the initial model texture based on the texture adjustment information, the method further includes: Acquire a second initial noise, and use the second initial noise to add noise to the second initial rendered image to obtain a second initial noisy image; Using the description text as a guide condition, predicting the noise of the second initial noisy image to obtain a second predicted noise; The degree of matching between the second initial rendered image and the description text includes a difference between the second predicted noise and the second initial noise.
8. The method according to claim 7, characterized in that Before obtaining the second initial noise, the method further includes: determining a second initial noise intensity; determining a second initial noise based on the second initial noise intensity; The step of predicting the noise of the second initial noisy image based on the description text as a guide condition to obtain a second predicted noise includes: Acquire a pre-trained first image-text diffusion model, where the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of an object in an actual scene under the guidance of a guide text; The description text, the second initial noise intensity, and the second initial noisy image are input into the first image-text diffusion model to obtain a second predicted noise of the second initial noisy image.
9. The method according to claim 6, characterized in that Before adjusting the initial model texture based on the texture adjustment information, the method further includes: Obtaining texture features of a three-dimensional model; The texture adjustment information also includes: the matching degree between the initial model texture and the texture feature.
10. The method according to claim 9, characterized in that Before adjusting the initial model texture based on the texture adjustment information, the method further includes: Acquire a third initial noise, and use the third initial noise to add noise to the initial model map to obtain a third initial noisy image; Using the description text corresponding to the texture feature as a guide condition, predicting the noise of the third initial noisy image to obtain a third predicted noise; The degree of matching between the initial model map and the map feature includes: a difference between the third predicted noise and the third initial noise.
11. The method according to claim 10, characterized in that Before predicting the noise of the third initial noisy image based on the description text corresponding to the map feature as a guide condition to obtain a third predicted noise, the method further includes: Obtaining a pre-trained second image-text diffusion model, where the second image-text diffusion model is used to perform noise prediction on a three-dimensional model map under the guidance of a guide text; The step of predicting the noise of the third initial noisy image based on the description text corresponding to the map feature as a guide condition to obtain a third predicted noise includes: The description text corresponding to the map feature and the third initial noisy image are input into the second image-text diffusion model to obtain a third predicted noise.
12. The method according to claim 11, characterized in that The second image-text diffusion model is trained by the following method: Acquire a first image-text diffusion model and a training sample, wherein the first image-text diffusion model is used to perform noise prediction on an image that conforms to the characteristics of an object in an actual scene under the guidance of a guide text, and the training sample includes a sample three-dimensional model map; The description text corresponding to the texture feature is used as a guide condition, and the sample three-dimensional model texture is used to adjust the parameters of the first texture diffusion model to obtain a second texture diffusion model.
13. The method according to claim 12, characterized in that Before adjusting the initial model texture based on the texture adjustment information, the method further includes: Converting the initial model map into an initial brightness format map in brightness format; Obtain the brightness format map characteristics of the three-dimensional model; The texture adjustment information also includes: a matching degree between the initial brightness format texture and the brightness format texture features.
14. The method according to claim 13, characterized in that Before adjusting the initial model texture based on the texture adjustment information, the method further includes: Acquire a fourth initial noise, and use the fourth initial noise to add noise to the initial brightness format map to obtain an initial noisy brightness image; Using the description text corresponding to the brightness format map feature as a guide condition, predicting the noise of the initial noisy brightness image to obtain a fourth predicted noise; The degree of matching between the initial luminance format map and the luminance format map features comprises: a difference between the fourth predicted noise and the initial noisy luminance image.
15. The method according to claim 14, characterized in that Before predicting the noise of the initial noisy brightness image using the description text corresponding to the brightness format map feature as a guide condition to obtain a fourth predicted noise, the method further includes: Acquire a pre-trained third image-text diffusion model, where the third image-text diffusion model is used to perform noise prediction on a three-dimensional model map in a brightness format under the guidance of a guide text; The step of predicting the noise of the initial noisy brightness image using the description text corresponding to the brightness format map feature as a guide condition to obtain a fourth predicted noise includes: The description text corresponding to the brightness format map feature and the initial noisy brightness image are input into the third image-text diffusion model to obtain a fourth predicted noise.
16. The method according to claim 15, characterized in that The third image-text diffusion model is trained in the following way: Converting the sample three-dimensional model map into a sample brightness map in a brightness format; The description text corresponding to the brightness format map feature is used as a guide condition, and the sample brightness map is used to adjust the parameters of the first image-text diffusion model to obtain a third image-text diffusion model.
17. A three-dimensional model generating device, characterized in that: The device comprises: An acquisition unit, used to acquire a description text corresponding to the three-dimensional model to be generated, and an initial three-dimensional deformable model corresponding to the three-dimensional model to be generated; A rendering unit, configured to render the initial three-dimensional deformable model to obtain a first initial rendered image; An adjustment unit is used to adjust the model parameters of the initial three-dimensional deformable model based on the matching degree between the first initial rendering image and the description text, so as to generate a three-dimensional structure model that conforms to the description text.
18. An electronic device, characterized in that: include: processor; as well as The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method according to any one of claims 1 to 16 is executed.
19. A computer-readable storage medium, characterized in that: A data processing program is stored, and the program is run by a processor to execute the method according to any one of claims 1 to 16.