Three-dimensional image generation method and apparatus, electronic device, and medium

Through the data-free three-dimensional image generation method, the prior knowledge of the pre-trained model and FLAME model is used to explicitly represent geometric shape and texture information, solving the problems of high demand for training data and slow generation speed in the prior art, and achieving high-quality and rapid generation of three-dimensional images suitable for animation.

WO2025148920A1PCT designated stage expired Publication Date: 2025-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/071244
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-08
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The prior art requires a large number of training data sets when generating three-dimensional images, which are costly and slow to generate, making it difficult to apply to animation, especially in controlling bone structure and facial expressions.

Method used

Using data-free training method, three-dimensional images are generated through shape generators and texture generators, pre-trained models and average texture word characters are used, combined with the prior knowledge of the FLAME model, explicitly represent geometric shapes and texture information, reducing training data needs and improving generation speed.

Benefits of technology

It realizes high-quality and rapid generation of 3D images suitable for animation, reduces training costs, and can effectively control bone structure and facial expressions. It is suitable for 3D modeling of games, audio and video and wearable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071244_17072025_PF_FP_ABST
    Figure CN2025071244_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a three-dimensional image generation method and apparatus, an electronic device, and a medium. The three-dimensional image generation method comprises: receiving a text prompt for describing an object; generating a geometric shape model of the object from the text prompt by means of a shape generator; incorporating an average texture token into the text prompt to obtain a modified text prompt, wherein the average texture token represents common texture information of the category of the object; generating a texture parameter of the object from the modified text prompt by means of a texture generator; and generating a three-dimensional image of the object on the basis of the geometric shape model and the texture parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional image generation method, device, electronic device and medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 12, 2024, with application number 202410051143.9 and application name “Three-dimensional image generation method and device”. The entire contents of the application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of artificial intelligence, and more specifically, to a three-dimensional image generation method, device, electronic device, and medium. Background Art

[0003] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks in various AI fields. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0004] Computer vision (CV) technology attempts to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has revolutionized the development of computer vision technology. Pre-trained models in the field of vision, such as swin-transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely adapt to specific downstream tasks. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional (3D) object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0005] With the development of digital entertainment industries like gaming and audiovisual, computer vision technologies like virtual reality and augmented reality are finding increasing application. Their implementation relies heavily on 3D data resources, particularly 3D portraits. Therefore, generating high-quality 3D images has become a research hotspot in computer vision technology in recent years. Summary of the Invention

[0006] The present disclosure provides a three-dimensional image generation method, device, electronic device, medium, and computer program product.

[0007] According to one aspect of an embodiment of the present disclosure, a three-dimensional image generation method is proposed, comprising: receiving a text prompt for describing an object; generating a geometric shape model of the object from the text prompt through a shape generator; merging an average texture token into the text prompt to obtain a modified text prompt, wherein the average texture token represents common texture information of a category of the object; generating texture parameters of the object from the modified text prompt through a texture generator; and generating a three-dimensional image of the object based on the geometric shape model and the texture parameters.

[0008] According to another aspect of an embodiment of the present disclosure, a three-dimensional image generation device is provided, the device comprising: an input unit configured to receive a text prompt for describing an object; a shape generation unit configured to generate a geometric shape model of the object from the text prompt through a shape generator; a texture generation unit configured to: merge an average texture token into the text prompt to obtain a modified text prompt, the average texture token representing common texture information of the category of the object; generate texture parameters of the object from the modified text prompt through a texture generator; and an output unit configured to generate a three-dimensional image of the object based on the geometric shape model and the texture parameters.

[0009] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: one or more processors; and one or more memories, wherein the memories store computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the one or more processors execute the methods described in the above aspects.

[0010] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the processor executes a method as described in any one of the above aspects of the present disclosure.

[0011] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, which includes computer-readable instructions. When the computer-readable instructions are executed by a processor, the processor executes the method as described in any one of the above aspects of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other purposes, features, and advantages of the embodiments of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] FIG1 shows an exemplary scene diagram of a 3D image generation system according to an embodiment of the present disclosure.

[0014] FIG2 shows a flowchart of a method for generating a three-dimensional image according to an embodiment of the present disclosure.

[0015] FIG3 shows a flowchart of a method for training a shape generator according to an embodiment of the present disclosure.

[0016] FIG4 shows a flowchart of a method for training a texture generator according to an embodiment of the present disclosure.

[0017] FIG5 shows a training data generation process according to an example of an embodiment of the present disclosure.

[0018] FIG6 illustrates a training process of a shape generator and a texture generator according to an embodiment of the present disclosure.

[0019] FIG7 shows a schematic processing process of a three-dimensional image generation model according to an example of an embodiment of the present disclosure.

[0020] FIG8 shows a schematic structural diagram of a three-dimensional image generating device according to an embodiment of the present disclosure.

[0021] FIG9 shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0023] As shown in the embodiments and claims of the present disclosure, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not specifically refer to the singular, but may also include the plural. "First", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0024] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0025] In addition, flowcharts are used in this disclosure to illustrate the operations performed by the system according to the embodiments of the present disclosure. It should be understood that the preceding or following operations do not necessarily need to be performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more operations can be removed from these processes.

[0026] In recent years, computer vision technology has developed rapidly, with 3D modeling technology, particularly AI-generated content (AIGC), being a research hotspot. AIGC is typically implemented using large models, which refer to neural network models with extremely large parameters (typically over a billion). These models possess powerful expressive and learning capabilities, but also require vast amounts of data and computing resources for training. This has led to the development of pre-trained models. Pre-trained models are deep neural networks (DNNs) with large parameters. These models are trained on massive amounts of unlabeled data to learn universal feature representations. These models can be adapted for various downstream tasks using techniques such as fine tuning, parameter-efficient fine tuning (PEFT), and prompt tuning. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. Pre-trained models are an important tool for outputting AIGC and can also serve as a universal interface for connecting multiple task-specific models.

[0027] Today's AIGC is mainly based on a diffusion model that generates images from text. Generally speaking, a diffusion model can include a forward process and a reverse process, where the forward process adds random noise to the image, and the reverse process restores the image from the noisy image. In order to solve the speed bottleneck of the diffusion model, the latent diffusion model (LDM) converts image processing in pixel space to a latent space with smaller dimensions, thereby greatly improving training efficiency. A typical example is the Stable Diffusion model. In 2022, Google released the DreamFusion model that can realize text to 3D objects, which introduced a training strategy called Score Distillation Sampling (SDS), which greatly promoted the design of text to 3D object models and provided a good foundation for a large number of subsequent scientific research work.

[0028] Related text-to-3D object models often use SDS techniques to independently optimize the shape and texture information of 3D objects based on pre-trained diffusion models. Although these SDS-based methods can produce good 3D static objects, they have the following drawbacks: (1) a large amount of training datasets are required for model training, which is costly; (2) current methods require a large number of iterations to optimize each text prompt, which consumes a lot of time when generating 3D objects; (3) current methods mainly rely on implicit representations, which cannot directly reflect the geometric shape information of 3D objects and lack the necessary control over elements such as bone structure and facial expressions, making them unsuitable for animation.

[0029] In response to the above problems, the present disclosure proposes a three-dimensional image generation model and a three-dimensional image generation method based on the model, which adopts a data-free training method, can directly generate training data without collecting large-scale training data sets, and can generate three-dimensional images suitable for animation at a faster speed.

[0030] FIG1 shows an exemplary scene diagram of a 3D image generation system according to an embodiment of the present disclosure. As shown in FIG1 , the 3D image generation system 100 may include a user terminal 110 , a network 120 , a server 130 , and a database 140 .

[0031] User terminal 110 may be, for example, computer 110-1 or mobile phone 110-2 shown in FIG1 . It is understood that, in fact, user terminal 110 may be any other type of electronic device capable of performing data processing, including but not limited to fixed terminals such as desktop computers and smart TVs, mobile terminals such as smart phones, tablet computers, portable computers, handheld devices, or any combination thereof, and the present disclosure does not impose specific limitations on this.

[0032] According to an embodiment of the present disclosure, the user terminal 110 can be used to receive a text prompt and use the 3D image generation method provided by the present disclosure to generate a 3D image of an object based on the text prompt. In some embodiments, the 3D image generation method provided by the present disclosure can be executed by a processing unit of the user terminal 110. In some implementations, the user terminal 110 can use an application built into the user terminal to execute the 3D image generation method provided by the present disclosure. In other implementations, the user terminal 110 can execute the 3D image generation method provided by the present disclosure by calling an application stored externally to the user terminal.

[0033] In other embodiments, user terminal 110 transmits the received text prompt to be processed to server 130 via network 120, and server 130 executes the 3D image generation method. In some implementations, server 130 may execute the 3D image generation method using a built-in application program. In other implementations, server 130 may execute the 3D image generation method by invoking an application program stored externally to the server.

[0034] The network 120 may be a single network or a combination of at least two different networks. For example, the network 120 may include, but is not limited to, a local area network, a wide area network, a public network, a private network, or a combination of one or more of the following. The server 130 may be an independent server, or a server cluster or distributed system consisting of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, positioning services, and big data and artificial intelligence platforms. The embodiments of the present disclosure do not impose specific limitations on this.

[0035] The database 140 can generally refer to a device with a storage function. The database 140 is mainly used to store various data used, generated and output during the operation of the user terminal 110 and the server 130. The database 140 can be local or remote. The database 140 can include various memories, such as random access memory (RAM), read-only memory (ROM), etc. The storage devices mentioned above are just some examples, and the storage devices that can be used in the system are not limited to these. The database 140 can be interconnected or communicated with the server 130 or a part thereof via the network 120, or directly interconnected or communicated with the server 130, or a combination of the above two methods.

[0036] The following describes a three-dimensional image generation method according to an embodiment of the present disclosure with reference to FIG2 . FIG2 shows a flow chart of a three-dimensional image generation method 200 according to an embodiment of the present disclosure. As described above, the three-dimensional image generation method 200 can be executed by a user terminal or a server, and the present disclosure does not impose any specific limitations on this. The three-dimensional image generation model used to implement the three-dimensional image generation method according to an embodiment of the present disclosure may include a shape generator and a texture generator, as described in further detail below.

[0037] In step S210, a text prompt for describing the object is received. In the embodiment of the present disclosure, a three-dimensional representation of any object can be generated, such as a three-dimensional avatar of a specified person or animal, and the embodiment of the present disclosure does not impose specific restrictions on this. The text prompt is a description of the content, style, etc. of the object for which the three-dimensional representation is to be generated, and the embodiment of the present disclosure does not impose specific restrictions on this. For example, if the object for which the three-dimensional representation is to be generated is the head of a specified person, the text prompt may include one or more of the person's name, age, gender, appearance, etc. In one example, if the user wishes to generate a three-dimensional avatar of person A, he or she may enter the text prompt "person A", or further enter the text prompt "person A, old age", etc. In another example, if the user wishes to generate a three-dimensional avatar of animal B, he or she may enter the text prompt "animal B", or further enter the text prompt "animal B, yellow", etc.

[0038] In step S220, a geometric shape model of the object is generated from the text prompt via a shape generator. Here, a geometric shape model refers to a three-dimensional representation that includes only geometric structure and does not have texture information. In some embodiments, step S220 includes: generating shape parameters from the text prompt via a shape generator; and generating the geometric shape model of the object from the shape parameters. For example, the shape generator generates shape parameters based on the text prompt, and generates the geometric shape model of the object under the control of the shape parameters. Specifically, the shape generator according to an embodiment of the present disclosure may include a text encoder that can encode the text prompt to generate a text embedding vector, and then generate shape parameters based on the text embedding vector. In an embodiment of the present disclosure, the text encoder can be constructed based on a relevant pre-trained model, for example, based on a contrastive language-image pre-training (CLIP) model. The CLIP model is pre-trained using large-scale text-image pairs (i.e., an image and its corresponding text description) to learn the matching relationship between text and image. The text embedding vector extracted from the text prompt by the CLIP model can be expressed as l = CLIP(T), where T represents the text prompt. The text encoder is described above using CLIP as an example, but the embodiments of the present disclosure are not limited thereto, and other appropriate pre-training models may also be used to construct the text encoder.

[0039] In some embodiments, the shape parameters generated by the shape generator can represent the geometric shape of the object. In some embodiments, unlike the implicit representation commonly used in related neural network models, the shape parameters generated by the shape generator according to the embodiment of the present disclosure can be used to explicitly represent the geometric shape of the object based on point clouds, grids or voxels, etc., so that they can be directly used to generate a geometric shape model of the object, and then can be directly animated. Moreover, when generating a geometric shape model based on shape parameters, a more accurate geometric shape model can also be generated in combination with prior knowledge of the object. Taking the object as an example of a human or animal head, a three-dimensional deformation model with prior knowledge of the facial shape, expression and posture of a human or animal can be used to generate a geometric shape model under the control of shape parameters. In some embodiments, the shape generator may include a three-dimensional deformation model.

[0040] For example, the FLAME model, a 3D morphable model, is a statistical model for 3D face reconstruction. FLAME is a linear shape space trained on 3,800 head scans. It incorporates shape, expression, and pose variations, processing these parameters into a head mesh, outputting 5,023 mesh vertices. For a 3D human head, the FLAME model can be used to generate a geometric model of the head under the control of shape parameters. Because the FLAME model incorporates extensive prior knowledge of the human face, it effectively controls elements such as skeletal structure and facial expression, resulting in a high-fidelity 3D head model.

[0041] Typically, the text embedding vector generated by the CLIP model has a higher dimension, while the shape parameters generated by the shape generator according to an embodiment of the present disclosure have a lower dimension in order to be suitable for being driven by a priori models such as the FLAME model. To this end, the shape generator according to an embodiment of the present disclosure may further include a multi-layer perception (MLP) block that reduces the dimensionality of the text embedding vector, and is used to project the text embedding vector into the shape parameter space to generate predicted shape parameters. Taking the CLIP model as an example, the predicted shape parameter S can be expressed as S = f(CLIP(T)), where f represents the mapping function of the multi-layer perception block.

[0042] In step S230, the average texture token is merged into the text prompt to obtain a modified text prompt. The average texture token represents the common texture information of the category of the object. More specifically, the average texture token represents the common texture information of the category or style to which the object belongs. For example, when the object is a human or animal head, the average texture token can represent one or more facial features, facial topology, facial texture coordinate mapping relationships, etc. shared between humans or animals. The facial texture coordinate mapping relationship represents the mapping relationship between the texture coordinates of the face and the vertex coordinates of the geometric shape model of the head. According to the example of the embodiment of the present disclosure, the average texture token can be appended before or after the text prompt to generate a modified text prompt. The embodiment of the present disclosure does not specifically limit the specific modification method.

[0043] In step S240, a texture generator is used to generate texture parameters of the object from the modified text prompt, wherein the texture parameters represent texture information of the object.

[0044] In step S250, a three-dimensional image of the object is generated based on the geometric shape model and the texture parameters. In some embodiments, by mapping the texture parameters to the geometric shape model generated based on the shape parameters, a three-dimensional image with texture information can be generated. For example, the texture parameters can be texture values ​​represented by texture coordinates (such as UV coordinates), each of which corresponds one-to-one to each vertex in the geometric shape model. Texture mapping is performed based on the correspondence, and a three-dimensional image with texture information can be generated. Here, the three-dimensional image is a representation of the three-dimensional model of the object, which can also be called a three-dimensional representation or three-dimensional data. According to an embodiment of the present disclosure, the texture generator can be constructed based on a related pre-trained model. For example, the texture generator can include a stable diffusion model, but the embodiment of the present disclosure is not limited to this, and other appropriate pre-trained models can also be used.

[0045] In the embodiment of the present disclosure, in addition to the text prompt pointing to a specific object, an average texture token is introduced to represent the average texture information of the category to which the object belongs, such as the average facial features of a human head. The trained texture generator is able to learn the average texture information represented by the average texture token, as will be described in further detail below, and is thus able to generate more accurate texture parameters based on the text prompt modified using the average texture token. In the embodiment of the present disclosure, the average texture token may be a rare token, that is, a token composed of rare characters other than natural language characters, such as "T^*", "# / ", "¥%", etc., which is not detailed in the embodiment of the present disclosure. Compared with the use of natural language characters, the use of rare tokens can avoid ambiguity caused by its strong prior knowledge of natural language when training the texture generator.

[0046] In an embodiment of the present disclosure, a data-free training method for a three-dimensional image generation model is proposed. Instead of manually collecting a large-scale training data set, a training data set can be directly generated based on a pre-trained model. Specifically, a shape training data set for training a shape generator and a texture training data set for training a texture generator can be generated by training a pre-trained model. The following describes the training method for the shape generator and the texture generator according to an embodiment of the present disclosure with reference to Figures 3 to 5. Figure 3 shows a flowchart of a training method 300 for a shape generator according to an embodiment of the present disclosure, Figure 4 shows a flowchart of a training method 400 for a texture generator according to an embodiment of the present disclosure, and Figure 5 shows an example training data generation process according to an embodiment of the present disclosure.

[0047] When generating training data, a candidate text prompt set including multiple candidate text prompts can be prepared. As shown in FIG3 , for each candidate text prompt in the candidate text prompt set, in step S310, an initial shape parameter can be set. For example, the initial shape parameter can be set to a sequence of all zeros. This is not specifically limited in the present embodiment. In the example of FIG5 , the generation of a three-dimensional human portrait is used as an example for illustration, wherein FIG5 shows an example text prompt "Person A".

[0048] In step S320, a geometric shape model is generated based on the initial shape parameters, and a first two-dimensional image at a predetermined perspective is generated from the geometric shape model. For example, the geometric shape model can be generated using a model that has prior knowledge of the object. In the example of Figure 5 , FLAME is used as an example. The initial shape parameters are input into the FLAME model to generate the geometric shape model. Thereafter, a first two-dimensional image at a predetermined perspective is generated from the geometric shape model. For example, a two-dimensional image at any perspective is captured from the generated geometric shape model. For ease of distinction, this is referred to as the first two-dimensional image, as shown in Figure 5 .

[0049] In step S330, the first two-dimensional image and the candidate text prompt are input into the first pre-trained model to calculate the first loss of the first pre-trained model. In step S340, it is determined whether the first loss satisfies the first predetermined condition. When it is determined that the first loss does not meet the first predetermined condition, the initial shape parameters are updated based on the first loss, and the above steps S320-S340 are continued to be executed until the first predetermined condition is met to generate optimized shape parameters corresponding to the candidate text prompt. For example, the first predetermined condition is that the first loss is minimized or reaches a predetermined value, or reaches a predetermined number of iterations. When the first predetermined condition is met, the iterative execution of steps S320-S340 is stopped, and the optimized shape parameters corresponding to the candidate text prompt are output. In the example of Figure 5, the first pre-trained model can adopt a stable diffusion model, then the first loss can be a score distillation sampling (SDS) loss L SDS1, that is, the loss between the random noise added to the image during the diffusion process and the noise predicted from the noisy image. Specifically, the first loss L can be calculated according to the following equation (1): SDS1 :

[0050] in, represents the function of generating images from text; x represents the input two-dimensional image sample; z is the potential feature map of the two-dimensional image sample in the latent space, z t It represents the version with added noise; y represents the text embedding vector; t represents the time step; ∈ represents the added random noise; represents the predicted noise; w(t) represents the weight; θ is the three-dimensional volume parameter; Indicates L SDS The gradient of the learnable parameters, Indicates expected value.

[0051] In the above formula (1), the predicted noise It can be calculated according to the following equation (2):

[0052] Among them, ∈ φ is the pre-trained denoising function, w e The scale parameter w is introduced to improve the sample fidelity and balance the diversity of the generated samples. e Set to a higher value to enhance the text control ability and enhance the sample fidelity, but at the expense of sample diversity. In the embodiment of the present disclosure, since the strong prior knowledge provided by the FLAME model is used in the above process, the ratio parameter w can be set to e Set it to a lower value to ensure sample fidelity without sacrificing sample diversity.

[0053] By the method shown in FIG3 , optimized shape parameters corresponding to each candidate text prompt in the candidate text prompt set can be generated, and these text prompt-optimized shape parameter pairs can be used to train the shape generator. For example, a shape training data set for training the shape generator may include multiple training text prompts selected from the candidate text prompt set and corresponding multiple optimized shape parameters. It should be noted that although the process of generating the shape training data set is described using the stable diffusion model as an example in FIG5 , the present disclosure is not limited thereto, and any other appropriate pre-training model may also be used to generate the shape training data set.

[0054] Based on the shape training dataset including optimized shape parameters generated according to the method of FIG. 3 , a texture training dataset can be further generated using the method shown in FIG. 4 . As shown in FIG. 4 , for each candidate text prompt in the candidate text prompt set, initial texture parameters can be set in step S410 . Here, the initial texture parameters can, for example, be the average texture parameters extracted from multiple real 3D representations. For example, when the object is a human head, the average texture parameters can be extracted from multiple 3D representations of real heads and used as the initial texture parameters. Alternatively, the initial texture parameters can be directly derived from generic texture parameters from existing 3D morphable models such as FLAME, for example, a generic UV map as shown in the example of FIG. 5 . In other words, the initial texture parameters can be expressed as ψ = ψ0 + Δψ, where ψ0 represents the average texture parameters extracted from the real 3D representation or generic texture parameters provided by FLAME, etc., and Δψ represents the object-specific texture details. Its initial value is set to 0, and it is expected that Δψ will be continuously optimized through the training process to enable it to represent personalized texture information.

[0055] In step S420, an initial 3D model is generated based on the initial texture parameters and the optimized shape parameters corresponding to the current candidate text, and a second 2D image at a predetermined perspective is generated from the initial 3D model. For example, a geometric model can be first generated based on the optimized shape parameters, and then the initial texture parameters are mapped onto the geometric model to generate the initial 3D model. Subsequently, a 2D image at a predetermined perspective is generated from the generated initial 3D model. For ease of distinction, this is referred to herein as the second 2D image, as shown in FIG5 .

[0056] In step S430, the second two-dimensional image and the candidate text prompt are input into the second pre-trained model to calculate the second loss of the second pre-trained model. In step S440, it is determined whether the second loss satisfies the second predetermined condition. When the second predetermined condition is not met, the initial texture parameters are updated based on the second loss, and the above steps S420-S440 are continued until the second predetermined condition is met. The second predetermined condition is, for example, that the second loss is minimized or reaches a predetermined value, or reaches a predetermined number of iterations. When the second predetermined condition is reached, the iteration is stopped and the optimized texture parameters corresponding to the candidate text prompt are output. In the example of Figure 5, the second pre-trained model can also adopt a stable diffusion model, then the second loss can be an SDS loss L SDS2 , which is the loss between the random noise added to the image during the diffusion process and the noise predicted from the noisy image. For example, the second loss L can be calculated according to the above equation (1): SDS2 .

[0057] By the method shown in FIG5 , optimized texture parameters corresponding to each candidate text prompt in the candidate text prompt set can be generated, and these text prompt-optimized texture parameter pairs can be used to train a texture generator. For example, a texture training data set for training a texture generator can include the multiple training text prompts selected above and the corresponding multiple optimized texture parameters. It should be noted that although the process of generating a texture training data set is described using a stable diffusion model as an example in FIG5 , the present disclosure is not limited thereto, and any other appropriate pre-training model can also be used to generate a texture training data set.

[0058] The above describes the process of generating shape and texture training datasets by training a pre-trained model, such as a stable diffusion model. This process can generate any amount of training data, eliminating the need to collect large-scale training datasets during training of a 3D image generation model according to embodiments of the present disclosure. This reduces the data cost required for training and simplifies the training process.

[0059] The following describes the training process of the shape generator and texture generator according to an embodiment of the present disclosure with reference to FIG6 . FIG6 illustrates the training process of the shape generator and texture generator according to an embodiment of the present disclosure, wherein the shape generator includes a text encoder and an MLP block constructed based on a pre-trained model, and the texture generator can also be constructed based on a pre-trained model. Here, to distinguish it from the first and second pre-trained models used to generate training data above, the pre-trained model of the texture generator is referred to as a third pre-trained model, which can be, for example, a stable diffusion model.

[0060] When training the shape generator, each training text prompt in the shape training data set and its corresponding optimized shape parameters can be used as the ground truth (Ground Truth), and the shape generator can be trained in a model fine-tuning manner. In an embodiment of the present disclosure, the text encoder of the shape generator is constructed based on a pre-trained model. For example, the text encoder may include a CLIP model. Since pre-trained models such as CLIP have been pre-trained using large-scale text-image pairs (i.e., images and their corresponding text descriptions), the shape generator can be fine-tuned to make it suitable for generating shape parameters according to an embodiment of the present disclosure. In order to achieve fine-tuning of the shape generator, as shown in Figure 6, the shape generator according to an embodiment of the present disclosure may also include a shape adaptation block for converting the training of a high-dimensional parameter matrix of a text encoder such as CLIP into fine-tuning through a shape adaptation block.

[0061] According to the examples of the embodiments of the present disclosure, the shape generator can be fine-tuned by using the low-rank adaptation (LoRA) method of a large language model, but this is only an example and not a limitation. Other appropriate fine-tuning methods can also be used, such as efficient parameter fine-tuning, prompt fine-tuning, etc. The basic principle of LoRA is to add a "side branch" to the original model, which decomposes the high-dimensional matrix into two low-rank matrices. During training, only these two low-rank matrices are trained, thereby greatly reducing the amount of training parameters and increasing the training speed. In the case of fine-tuning using LoRA, assuming that the dimension of the text embedding vector output by the text encoder is d*d, the d*d dimensional matrix can be decomposed into a d*r matrix and an r*d matrix through the shape adaptation block, where r is much smaller than d. During training, only the d*r matrix and the r*d matrix are trained, thereby achieving low-parameter fine-tuning of the large model.

[0062] The shape generator trained in a fine-tuning manner can generate shape parameters with desired characteristics based on the input text prompts. It can explicitly represent the geometric shape of the object through point clouds, meshes or voxels, and can be directly used to generate geometric shape models. For example, the FLAME model can be directly used to generate geometric shape models. Compared with the implicit representation that requires further processing by neural networks to represent geometric shapes, it is more suitable for animation purposes.

[0063] The training method for the texture generator is similar to that of the shape generator, except that for each training text prompt in the texture training dataset, the average texture token is merged into the text prompt to obtain a modified text prompt, and then the modified training text prompt and the corresponding optimized texture parameters are used as the ground truth for training the texture generator. As mentioned above, the purpose of the average texture token is to represent the average texture information of the category to which the object belongs. Therefore, it is expected that through training, the texture generator can learn the expected average texture information contained in the average texture token. In the embodiment of the present disclosure, the average texture token can be a rare token composed of rare characters, such as "T*".

[0064] Similarly, the training of the texture generator also adopts the model fine-tuning method. In the embodiment of the present disclosure, the third pre-trained model of the texture generator can be, for example, a stable diffusion model. Since the third pre-trained model such as the stable diffusion model has been pre-trained using large-scale text-image pairs (i.e., images and their corresponding text descriptions), the texture generator can be fine-tuned to make it suitable for generating texture parameters according to the embodiment of the present disclosure. In order to achieve fine-tuning of the texture generator, as shown in Figure 6, the texture generator according to the embodiment of the present disclosure may further include a texture adaptation block for converting the training of the high-dimensional parameter matrix of the third pre-trained model such as the stable diffusion model into fine-tuning through the texture adaptation block. According to the example of the embodiment of the present disclosure, the texture generator can be fine-tuned in a LoRA manner, but this is only an example and not a limitation. Other appropriate fine-tuning methods may also be used, such as efficient parameter fine-tuning, prompt fine-tuning, and the like.

[0065] The texture generator trained in a fine-tuning manner can learn the average texture information included in the average texture token, so that it can generate high-quality texture parameters based on text prompts pointing to a specific object and the average texture token, which can not only reflect the average texture information of the category to which the object belongs, such as facial features shared by humans, facial topology, facial UV mapping relationship, etc., but also represent the personalized texture characteristics of the specific object.

[0066] The above describes the training process of the shape generator and texture generator of the 3D image generation model according to an embodiment of the present disclosure. The trained 3D image generation model is capable of generating a 3D image of an object based on an input text prompt, such as outputting a 3D head model of a person based on an input name. Figure 7 illustrates a schematic processing process of the 3D image generation model according to an example embodiment of the present disclosure. In this example, the generation of a 3D head portrait of a specified person is used as an example. As shown in Figure 7, for example, the text prompt "Person A" can be input into the 3D image generation model. The shape generator encodes this text prompt using a text encoder to generate a text embedding vector. The text embedding vector is then subjected to dimensionality reduction processing using an MLP block to generate shape parameters. These shape parameters can then be used to generate a geometric shape model using a model with prior knowledge, such as the FLAME model. Alternatively, the text prompt is modified using an average texture token, for example, by simply adding the average texture token to the front of the text prompt. The modified text prompt is then processed by a third pre-trained model of the texture generator to output texture parameters. By combining the geometric shape model and texture parameters, for example, by mapping the texture parameters to the geometric shape model, a 3D image can be generated, as shown in Figure 7.

[0067] Using the 3D image generation method according to an embodiment of the present disclosure, a data-free training method is used to train the shape generator and texture generator of the 3D image generation model, reducing the data cost required for training and simplifying the training process. During the training process of the shape generator according to an embodiment of the present disclosure, a 3D deformable model with prior knowledge of 3D objects, such as the FLAME model, is utilized. The trained shape generator can generate shape parameters that can explicitly represent geometric shapes. These shape parameters can be directly driven by a 3D deformable model such as FLAME, enabling more effective control of elements such as skeletal structure and facial expressions. Compared with traditional methods that rely on implicit representations, they are more suitable for animation purposes. Furthermore, because the shape generator according to an embodiment of the present disclosure uses a generalized shape generation model rather than an identity-specific model, it can efficiently generate 3D images of various objects without the need for individual optimization of each object during testing, significantly improving the model inference speed. Furthermore, the present embodiment introduces an average texture token to supplement the text prompt. This token can encapsulate important average texture information of 3D objects (such as facial texture features shared by humans), thereby facilitating the generation of higher-quality 3D images.

[0068] Compared with related models, the three-dimensional image generation method according to the embodiment of the present disclosure can stably and quickly generate higher-quality three-dimensional images, and the generated three-dimensional avatar can be conveniently used to generate continuous and natural animations. The method is particularly suitable for generating three-dimensional avatars of humans or animals and can be widely used in three-dimensional modeling and animation production in the fields of games, audio and video, wearable devices, etc. Taking the shape generator according to the embodiment of the present disclosure using the CLIP encoder as an example, the quantitative results of the three-dimensional image generation model proposed in the present disclosure and some related models under the same parameter setting conditions are compared, and the results are shown in Table 1 below. Among them, the CLIP score is used to indicate the fidelity of the generated image, and the higher the CLIP score, the better the image quality. As can be seen from Table 1, the three-dimensional image generation model according to the embodiment of the present disclosure has the highest CLIP score and the highest inference speed, which fully demonstrates the superiority of the three-dimensional image generation model proposed in the present disclosure.

[0069] Table 1 Quantitative comparison of the 3D image generation model of the present disclosure with some related models

[0070] The following describes a three-dimensional image generating device according to an embodiment of the present disclosure with reference to FIG8 . FIG8 shows a schematic structural diagram of a three-dimensional image generating device 800 according to an embodiment of the present disclosure. As shown in FIG8 , the three-dimensional image generating device 800 includes an input unit 810, a shape generating unit 820, a texture generating unit 830, and an output unit 840. In addition to these four units, the device 800 may also include other related components, but since these components are not relevant to the content of the present disclosure, a detailed description of their specific contents is omitted here. In addition, since the details of some functions of the device 800 are similar to the details of the steps of the method 200 described with reference to FIG2 , for the sake of brevity, a repeated description of some contents is omitted here. The device 800 according to an embodiment of the present disclosure can be implemented as a terminal or a server, as described above with reference to FIG1 .

[0071] Input unit 810 is configured to receive a text prompt describing an object. In the embodiment of the present disclosure, the target three-dimensional image can be a three-dimensional image of any object that the user wishes to generate, such as a three-dimensional portrait of a specified person or animal, and the embodiment of the present disclosure does not impose specific restrictions on this. The text prompt is a description of the content, style, etc. of the target three-dimensional image, and the embodiment of the present disclosure does not impose specific restrictions on this. For example, if the target three-dimensional object is a three-dimensional image of a specified person, the text prompt may include one or more of the person's name, age, gender, appearance, etc. In one example, if the user wishes to generate a three-dimensional portrait of person A, they can enter the text prompt "Person A" through input unit 810, or further enter the text prompt "Person A, elderly" or the like. In another example, if the user wishes to generate a three-dimensional portrait of animal B, they can enter the text prompt "Animal B" through input unit 810, or further enter the text prompt "Animal B, yellow" or the like.

[0072] The shape generation unit 820 is configured to generate a geometric shape model of the object from the text prompt through a shape generator. Here, the geometric shape model refers to a three-dimensional representation that only includes geometric shape information but does not have texture information. Specifically, the shape generator according to an embodiment of the present disclosure may include a text encoder, which can encode the text prompt to generate a text embedding vector, and then generate shape parameters based on the text embedding vector. In an embodiment of the present disclosure, the text encoder can be constructed based on a related pre-training model, for example, it can be constructed based on a contrastive language-image pre-training (CLIP) model. The CLIP model uses large-scale text-image pairs (i.e., images and their corresponding text descriptions) for pre-training to learn the matching relationship between text and images. The text embedding vector extracted by the CLIP model from the text prompt can be expressed as l=CLIP(T), for example, where T represents the text prompt. The text encoder is described above using CLIP as an example, but the embodiment of the present disclosure is not limited to this, and other appropriate pre-training models can also be used to construct the text encoder.

[0073] Unlike the implicit representation commonly used in related neural network models, the shape parameters generated by the shape generator according to the embodiment of the present disclosure can be used to explicitly represent the geometric shape of the object based on point clouds, meshes, or voxels, so that they can be directly used to generate a geometric shape model, and then can be directly animated. Moreover, when generating a geometric shape model based on shape parameters, it is also possible to combine prior knowledge of the object to generate a more accurate geometric shape model. Taking the object as an example of a human or animal head, a three-dimensional deformation model with prior knowledge of the facial shape, expression, and posture of the human or animal can be used to generate a geometric shape model of the target three-dimensional image under the control of shape parameters.

[0074] For example, the FLAME model is a statistical model for 3D face reconstruction. It uses a linear shape space trained on 3,800 head scans. It incorporates shape, expression, and pose variations, processing these parameters into a head mesh, outputting 5,023 mesh vertices. When the target 3D image is a 3D human head, the FLAME model can be used to generate a geometric model of the avatar under the control of shape parameters. Because the FLAME model incorporates extensive prior knowledge of the human face, it effectively controls elements such as skeletal structure and facial expression, resulting in a highly accurate 3D avatar.

[0075] Typically, the text embedding vector generated by the CLIP model has a higher dimension, while the shape parameters generated by the shape generator according to an embodiment of the present disclosure have a lower dimension in order to be suitable for being driven by a priori models such as the FLAME model. To this end, the shape generator according to an embodiment of the present disclosure may further include a multi-layer perception (MLP) block that reduces the dimensionality of the text embedding vector, and is used to project the text embedding vector into the shape parameter space to generate predicted shape parameters. Taking the CLIP model as an example, the predicted shape parameter S can be expressed as S = f(CLIP(T)), where f represents the mapping function of the multi-layer perception block.

[0076] The texture generation unit 830 is configured to merge the average texture token into the text prompt to obtain a modified text prompt, wherein the average texture token represents the common texture information of the category of the object, and more specifically, represents the average texture information of the category to which the object belongs. For example, when the category of the object is human or animal, the average texture token can represent one or more of facial features, facial topology, facial texture coordinate mapping relationships, etc. shared between humans or animals. According to an example of an embodiment of the present disclosure, the average texture token can be appended before or after the text prompt to generate a modified text prompt. The embodiment of the present disclosure does not specifically limit the specific modification method.

[0077] Afterwards, the texture generation unit 830 generates texture parameters of the object from the modified text prompt through a texture generator, wherein the texture parameters represent texture information of the target three-dimensional image. The output unit 840 is configured to generate a three-dimensional image of the object based on the geometric shape model and the texture parameters. For example, a three-dimensional image with texture information can be generated by mapping the texture parameters to a geometric shape model generated based on shape parameters, and the three-dimensional image can be output. For example, the texture parameters can be texture values ​​represented by texture coordinates (such as UV coordinates), each of which has a one-to-one correspondence with each vertex in the geometric shape model. Texture mapping is performed based on this correspondence, and a three-dimensional image with texture information can be generated. According to an embodiment of the present disclosure, the texture generator can be constructed based on a relevant pre-trained model, for example, it can be constructed based on a stable diffusion model, but the embodiment of the present disclosure is not limited thereto, and other appropriate pre-trained models can also be used.

[0078] In the embodiment of the present disclosure, in addition to the text prompt pointing to a specific object, an average texture token is introduced to represent the average texture information of the category to which the object belongs, such as the average facial features of a human head. The trained texture generator is able to learn the average texture information represented by the average texture token, as described in detail above, and is thus able to generate more accurate texture parameters based on the text prompt modified using the average texture token. In the embodiment of the present disclosure, the average texture token may be a rare token, that is, a token composed of rare characters other than natural language characters, such as "T^*", "# / ", "¥%", etc., which is not detailed in the embodiment of the present disclosure. Compared with the use of natural language characters, the use of rare tokens can avoid ambiguity caused by its strong prior knowledge of natural language when training the texture generator.

[0079] The three-dimensional image generation device 800 according to an embodiment of the present disclosure proposes a data-free training method. This method can generate a shape training dataset for training a shape generator and a texture training dataset for training a texture generator based on training an existing pre-trained model. The details are as described above with reference to Figures 3 to 5 and will not be repeated here. After obtaining the shape training dataset and the texture training dataset, the shape generator and the texture generator can be trained using the method described above with reference to Figure 6 by fine-tuning the large model. This method will not be repeated here.

[0080] Using a 3D image generation device according to an embodiment of the present disclosure, a data-free training method is used to train the shape generator and texture generator of a 3D image generation model, reducing the data cost required for training and simplifying the training process. During the training process of the shape generator according to an embodiment of the present disclosure, a 3D deformable model with prior knowledge of 3D objects, such as the FLAME model, is utilized. The trained shape generator can generate shape parameters that explicitly represent geometric shape information. This can be directly driven by a 3D deformable model such as FLAME, enabling more effective control of elements such as skeletal structure and facial expressions. Compared to traditional methods that rely on implicit representations, it is more suitable for animation purposes. Furthermore, because the shape generator according to an embodiment of the present disclosure uses a generalized shape generation model rather than an identity-specific model, it can efficiently generate 3D images of various objects without the need for individual optimization of each object during testing, significantly improving model inference speed. Furthermore, an average texture token is introduced to supplement textual prompts. This token can encapsulate important average texture information of 3D objects (such as facial texture features shared by humans), thereby facilitating the generation of higher-quality 3D images.

[0081] In addition, devices according to embodiments of the present disclosure (e.g., a three-dimensional image generation device, etc.) can also be implemented with the aid of the architecture of the exemplary computing device shown in FIG9 . FIG9 shows a schematic diagram of the architecture of an exemplary computing device according to embodiments of the present disclosure. As shown in FIG9 , computing device 900 may include a bus 910, one or more CPUs 920, a read-only memory (ROM) 930, a random access memory (RAM) 940, a communication port 950 connected to a network, an input / output component 960, a hard disk 970, and the like. Storage devices in computing device 900, such as ROM 930 or hard disk 970, can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU. Computing device 900 may also include a user interface 980. Of course, the architecture shown in FIG9 is merely exemplary, and when implementing different devices, one or more components of the computing device shown in FIG9 may be omitted according to actual needs. Devices according to embodiments of the present disclosure may be configured to execute the three-dimensional image generation method according to the various embodiments described above, or to implement the three-dimensional image generation apparatus according to the various embodiments described above.

[0082] The embodiments of the present disclosure may also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on a computer-readable storage medium according to an embodiment of the present disclosure. When the computer-readable instructions are executed by a processor, the three-dimensional image generation method according to the embodiment of the present disclosure described with reference to the above figures may be executed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0083] According to an embodiment of the present disclosure, a computer program product or computer program is also provided. The computer program product or computer program includes computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer-readable instructions from the computer-readable storage medium and execute the computer-readable instructions, causing the computer device to perform the three-dimensional image generation method described in each of the above embodiments.

[0084] The program portion of the technology can be considered a "product" or "article of manufacture" in the form of executable code and / or related data, implemented or implemented through computer-readable media. Tangible, permanent storage media can include any memory or storage used by a computer, processor, or similar device or related module. For example, various semiconductor memories, tape drives, disk drives, or any similar device that can provide storage for software.

[0085] All or part of the software may sometimes be communicated over a network, such as the Internet or other communications network. Such communications can load the software from one computer device or processor to another. Therefore, another medium capable of transmitting software elements may also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., which are transmitted through cables, optical cables or air. Physical media used to carry data, such as cables, wireless connections or optical cables and the like, can also be considered as the medium that carries the software. As used herein, unless limited to tangible "storage" media, other terms referring to computer or machine "readable media" refer to media that participate in the process of executing any instructions by the processor.

[0086] This application uses specific terms to describe the embodiments of this application. For example, "first / second embodiment", "one embodiment", and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or multiple times in different places in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.

[0087] In addition, it will be understood by those skilled in the art that various aspects of the present application can be illustrated and described by a number of patentable categories or situations, including any new and useful process, machine, product or combination of substances, or any new and useful improvements thereto. Accordingly, various aspects of the present application can be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may all be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". In addition, various aspects of the present application may be represented as a computer product located in one or more computer-readable media, which includes computer-readable program code.

[0088] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology and should not be interpreted in an idealized or highly formal sense, unless expressly defined as such herein.

[0089] The above is an illustration of the present invention and should not be considered as limiting thereof. Although several exemplary embodiments of the present invention have been described, it will be readily understood by those skilled in the art that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present invention. Therefore, all such modifications are intended to be included within the scope of the present invention as defined by the claims. It should be understood that the above is an illustration of the present invention and should not be considered as being limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present invention is defined by the claims and their equivalents.

Claims

1. A method for generating a three-dimensional image, which is executed in an electronic device, and the method includes: Receiving a text prompt for describing an object; Generating a geometric shape model of the object from the text prompt through a shape generator; Merging average texture tokens into the text prompt to obtain a modified text prompt, where the average texture tokens represent common texture information of the category of the object; Generating texture parameters of the object from the modified text prompt through a texture generator; and Generating a three-dimensional image of the object based on the geometric shape model and the texture parameters.

2. The method according to claim 1, wherein The object is a human head or an animal head, and the average texture tokens represent one or more of facial features, facial topologies, and facial texture coordinate mapping relationships shared among humans or animals.

3. The method according to claim 1 or 2, wherein generating the geometric shape model of the object from the text prompt through the shape generator includes: Generating shape parameters from the text prompt through a shape generator; Generating a geometric shape model of the object from the shape parameters.

4. The method according to claim 3, wherein, Generating shape parameters from the text prompt through a shape generator includes: Performing text encoding on the text prompt to generate a text embedding vector; and Performing dimensionality reduction processing on the text embedding vector to generate the shape parameters.

5. The method according to claim 3 or 4, wherein The shape parameters represent the geometric shape of the object.

6. The method according to any one of claims 1-5, wherein, The shape training dataset for training the shape generator and the texture training dataset for training the texture generator are generated by training a pre-trained model.

7. The method according to claim 6, wherein, The shape training dataset for training the shape generator is generated in the following manner: For each candidate text prompt in a candidate text prompt set: Setting initial shape parameters; Generating a geometric shape model based on the initial shape parameters, and generating a first two-dimensional image of a predetermined perspective from the geometric shape model; Inputting the first two-dimensional image and the candidate text prompt into a first pre-trained model to calculate a first loss of the first pre-trained model; Updating the initial shape parameters based on the first loss to generate optimized shape parameters corresponding to the candidate text prompt; Wherein, the shape training dataset includes a plurality of training text prompts selected from the candidate text prompt set and corresponding plurality of optimized shape parameters.

8. The method according to claim 7, wherein, The texture training dataset for training the texture generator is generated in the following manner: For each candidate text prompt in the candidate text prompt set: Setting initial texture parameters; Generating an initial three-dimensional model based on the initial texture parameters and the optimized shape parameters corresponding to the candidate text prompt; Generating a second two-dimensional image of a predetermined perspective from the initial three-dimensional model; Inputting the second two-dimensional image and the candidate text prompt into a second pre-trained model to calculate a second loss of the second pre-trained model; Updating the initial texture parameters based on the second loss to generate optimized texture parameters corresponding to the candidate text prompt; Wherein, the texture training dataset includes the plurality of training text prompts and corresponding plurality of optimized texture parameters.

9. The method according to any one of claims 1-8, wherein, The shape generator is trained in the following manner: based on each training text prompt in the shape training dataset and its corresponding optimized shape parameters, the shape generator is fine-tuned; The texture generator is trained in the following manner: For each training text prompt in the texture training dataset, the average texture token is merged into the training text prompt to generate a modified training text prompt; Based on each modified training text prompt in the texture training dataset and its corresponding optimized texture parameters, the texture generator is fine-tuned.

10. The method according to any one of claims 1-9, wherein, The shape generator includes a text encoder, a multi-layer perceptron block for dimensionality reduction, and a shape adaptation block for model fine-tuning; The texture generator includes a third pre-trained model and a texture adaptation block for model fine-tuning.

11. According to the method according to any one of claims 1-10, wherein, The average texture token is a token composed of rare characters other than natural language characters.

12. The method according to any one of claims 3-11, wherein, The object is a human head or an animal head; Generating the geometric shape model of the object from the shape parameters includes: Generating the geometric shape model under the control of the shape parameters using a three-dimensional deformation model with prior knowledge of the facial shape, expression, and pose of a human or an animal.

13. A three-dimensional image generation device, the device includes: An input unit configured to receive a text prompt for describing an object; A shape generation unit configured to generate the geometric shape model of the object from the text prompt through a shape generator; A texture generation unit configured to: merge the average texture token into the text prompt to obtain a modified text prompt, where the average texture token represents the common texture information of the category of the object; Generate the texture parameters of the object from the modified text prompt through a texture generator; And An output unit configured to generate a three-dimensional image of the object based on the geometric shape model and the texture parameters.

14. An electronic device, including: One or more processors; And One or more memories, where computer-readable instructions are stored in the memories, and when the computer-readable instructions are run by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1-12.

15. A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-12.

16. A computer program product, which includes computer-readable instructions, and when the computer-readable instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Virtual image generation method and device, equipment and storage medium

    CN115908657A

  • High-fidelity three-dimensional face model generation method based on natural text description

    CN115984485A

  • Text generation 3D printing model method based on big data deep learning

    CN116580156A

  • Three-dimensional model generation method and device and electronic equipment

    CN116843833A

  • Digital human generation method and device, electronic equipment and storage medium

    CN116993875A

Cited By

  • Data-driven UV mapping generation method, electronic device and program product

    CN120976397A