Text guidance-based two-stage tandem type three-dimensional model generation method
By constructing a secondary tandem three-dimensional model generation method based on text guidance, the three-plane self-attention and cross-word cross-attention module are used to solve the problem of loss of text information in the three-dimensional model generation, and efficient text information injection and three-dimensional model generation are achieved.
Patent Information
- Application Number
- CN202510318695.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, in the generation of text-guided three-dimensional model based on text guidance, there is a problem of loss of detailed information of text information and three-dimensional shape and texture modules, and the unsupervised training method cannot effectively transmit text information to the three-dimensional model.
Using a secondary tandem three-dimensional model generation method based on text guidance, the text information injection step is optimized by constructing a three-plane self-attention and cross-word cross-talk attention module, the text information injection step is extracted, and text features are extracted using the CLIP model, and combined with the retention mechanism to enhance the spatial continuity of the feature, the generator generates a fine-grained three-dimensional model.
It improves the efficiency of text information transmission in the three-dimensional model, enhances the parallel capability and operation speed of the deep learning network, and realizes efficient text information injection and three-dimensional model generation.
Smart Images

Figure BDA0005316772700000036 
Figure BDA0005316772700000038 
Figure BDA0005316772700000042
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and computer graphics, especially the field of text-guided 3D reconstruction and generation research. In the academic field, it can be used to construct a deep learning model for text-guided 3D reconstruction and generation, and can also be applied in the industrial field that requires customized 3D model generation. Background Art
[0002] In recent years, 3D model generation has become a booming field in computer vision research. Especially with the increasing popularity of AR, VR technologies, video game development, movie visual effects, and robot simulation, the demand for 3D models from concept to reality has doubled. In the pursuit of automated 3D model creation, many researchers have strived to develop methods for generating high-quality 3D assets. Early methods for 3D model generation mainly focused on how to enable the model to learn an efficient and effective representation of 3D models. However, the inherent unconditional nature of these methods not only hinders the customization of generated shapes based on specific preferences or requirements but also increases the difficulty of subsequent operations on the resulting models. Inspired by the latest achievements of text-to-image generation models, some research has attempted to achieve the goal by conditioning on specific texts. In 2018, Chen Jing et al. proposed Text2Shape, introducing the first large-scale 3D furniture object natural language description dataset and combining conditional WGAN with 3DCNN to achieve supervised training for text-guided 3D object generation. This 3D dataset with manually annotated captions encourages various text-guided 3D generation methods, especially implicit 3D representation methods. This method directly learns the mapping between text and the corresponding 3D shape, enabling the generation of objects that closely align with the specific details mentioned in the input text. Although the inclusion of manually annotated text supervision enhances the correlation between text prompts and generated shapes, the scarcity of manually annotated 3D datasets for various objects limits the applicability of this method to specific object classes. To reduce the dependence on manually annotated datasets and achieve unsupervised text-to-3D generation, subsequent methods utilize pre-trained text-driven 2D image generation networks or large vision and language models to address the inherent modality differences between text and vision. In 2023, Huang Tianyu et al. proposed the Textfield3D method to train a latent generator using rendered image embeddings and pseudo-text embeddings encoded by CLIP to solve the unsupervised training problem. Leveraging the visual and text alignment latent space of CLIP, this method can generate correct latent shapes based on text prompt embeddings. Although this method simplifies the necessity of pairing 3D models and texts, since the global features of the input text are used as guidance for generating 3D objects, mainly generating shapes and textures at the general class and color levels, this means that detailed information in the text prompt will be lost, resulting in the inability to fully convey text information to the 3D model.
[0003] Although significant progress has been made in text-guided 3D model generation technology, there are still deficiencies in constructing the text information and 3D shape and texture modules. Therefore, the present invention proposes a text-guided two-stage cascaded 3D model generation method for use in the generation and construction process of 3D models. Summary of the Invention
[0004] Object of the Invention: Aiming at the problem of detailed information loss in the construction of text information and 3D shape and texture by a text-guided 3D model generation network, a text-guided two-stage cascaded 3D model generation method is invented. By this method, the steps of injecting text information into the 3D model are optimized to reduce the loss rate of text information, and the continuity of the word-level feature space is optimized through an attention retention module, while improving the parallel ability and running speed of the deep learning network.
[0005] Technical Solution:
[0006] 1. A text-guided two-stage cascaded 3D model generation method generally includes the following steps:
[0007] Step 1.1: Construct a 3D model representation method based on three planes, use three matrices to represent the XY, XZ, and YZ planes respectively, and perform spatial interpolation calculations through these three matrices to obtain the feature representation of the 3D space to generate a 3D model;
[0008] Step 1.2: Construct a text feature extraction module, use the encoder of the text part in the Contrastive Language-Image Pretraining (CLIP) model to convert the input text into a mixed text feature t, which is divided into sentence-level feature t s and word-level feature t w ; fuse t s into the shape feature w geo and texture feature w tex output by the mapping network in the generator, and use t w as detailed information to input into a three-plane self-attention module based on a retention mechanism;
[0009] Step 1.3: Construct a three-plane self-attention module based on a retention mechanism, use the plane feature map retention mechanism to improve the continuity of plane features, and respectively construct a geometric three-plane and a texture three-plane with the latent shape feature and texture feature generated by the generator;
[0010] Step 1.4: Construct a cross-word cross-attention module, input the global text information into the three-plane features output by the three-plane self-attention module based on the retention mechanism, perform word-level refinement on the shape and texture three-plane features through cross-word attention, and the finally generated three-plane representation will include three-dimensional spatial information and word-level information;
[0011] Step 1.5: Connect the three-plane self-attention module based on the retention mechanism with the cross-word cross-attention module to construct a secondary cascaded text injection module. Utilize plane attention and cross-plane attention with the retention mechanism to refine the shape and texture three-planes at each layer of the generator, thereby correspondingly generating word-level geometric three-planes and texture three-planes;
[0012] Step 1.6: Construct a three-plane rendering and extraction module, perform feature decoding on the inferred three-plane features, use the volume rendering method to render the three-planes into RGB images for training, and the model training calculates the overall loss function including two parts: the generative adversarial loss and the CLIP loss; after the training is completed, extract the mesh object in traditional computer graphics through the isosurface extraction method (Marching Cubes, MC) in computer graphics.
[0013] 2. The method of the three-plane self-attention module based on the retention mechanism in the above Step 1.3 is as follows:
[0014] Step 2.1: Select the output features of the previous layer of the generator (where f represents features, the superscript g is for geometry, t is for texture, and the subscript i - 1 represents the previous layer), and input them into the three-plane attention module based on the retention mechanism;
[0015] Step 2.2: Synthesize the output features of the previous layer of the generator with the three-plane features of the current layer to form the features of the current layer When calculating the features of the texture layer, the features of the shape layer are also added to the calculation to generate textures that match the corresponding geometric shapes. The calculation process is expressed as:
[0016]
[0017] where the subscript i,xy represents the xy plane of the current layer. Similarly, i,yz and i,xz represent the yz plane and the xz plane, and these three planes are jointly used to represent the three-dimensional model in the network model;
[0018] Step 2.3: Calculate the plane self-attention for the three-plane features respectively. The calculation process is as follows:
[0019]
[0020] Q = W q X
[0021] K = W k X
[0022] V = W v X
[0023] Among them, Q, K, and V are three matrices in the attention mechanism, calculated from X, and W q,k,v are three weight matrices that the attention module needs to learn. Softmax is the normalization exponential function, ⊙ represents element-wise multiplication of matrices, and D 2d As the weight decay of the decay matrix module of the retention mechanism, mn is used to represent the coordinates of points, and the Manhattan distance is used to represent the two-dimensional coordinates of plane features;
[0024] Calculate the fused three-plane features in the same way
[0025]
[0026] Among them, concat means concatenating the tensor parameters horizontally to form a new matrix, and by default, it is extended on the second axis;
[0027] Step 2.4: Use the three-plane self-attention features calculated by the module in Step 2.3 as V and K, as Q, and calculate the cross-plane attention as follows:
[0028]
[0029] Q = W q Y
[0030] K = W k X
[0031] V = W v X
[0032] Among them, represents the dimension of the query set;
[0033] Connect the cross-plane attention features calculated separately into and send them to the next module, specifically as follows:
[0034]
[0035] where M2.4 is Module2.4.
[0036] 3. The method of the cross-word cross-attention module in the said Step 1.4 is as follows:
[0037] Step 3.1: Select the word-level feature t w and the cross-plane cross-attention feature of the previous module as the input of this module;
[0038] Step 3.2: By calculating to obtain Q, t w calculating to obtain K and V, calculating cross-word attention, and dynamically refining t w into the three-dimensional expression of the model to obtain the three-plane feature after word-level refinement The specific method is as follows:
[0039]
[0040] where, W q and W k are the matrix features to be learned by this module.
[0041] 4. The loss function for model training in the said step 1.6 is as follows:
[0042]
[0043] where, the loss function is divided into three terms, which are respectively the pixel discriminator loss with the sentence-level text feature t s added and the mask discriminator loss Finally, in order to guide the model to learn text information, the CLIP loss is added, and λ is a hyperparameter.
[0044] Advantages of the present invention:
[0045] 1. In the field of academic research, applying this method to construct a text-guided three-dimensional model generative deep learning network can realize injecting the text information input by the user into the three-dimensional model in a global and local manner, endowing the deep neural network with the ability to generate corresponding three-dimensional models based on text information.
[0046] 2. In the field of engineering applications, applying this method to the fields that require lightweight three-dimensional model editing functions can meet the needs of users to edit three-dimensional models through natural language, reduce industrial costs, and thus improve industrial efficiency. Brief description of the drawings
[0047] Figure 1 Overall model structure schematic diagram;
[0048] Figure 2 Method step schematic diagram;
[0049] Figure 3Method Structure and Technical Details Diagram;
[0050] Figure 4 Schematic Diagram of GET3D Depth Network Model Specific Embodiment
[0051] The present invention will be further described below with reference to the accompanying drawings.
[0052] The overall structural schematic diagram of the text-guided secondary cascaded three-dimensional model generation method is as Figure 1 shown, and generally includes the following steps:
[0053] (1) Construct a text feature extraction module, and use a pre-trained CLIP text encoder to convert the input text into sentence-level feature t s and word-level feature t w , and fuse the sentence-level text feature t s into the shape feature w geo and texture feature w tex output by the mapping network in the generator;
[0054] (2) Construct a text injection module based on a retention mechanism, perform word-level refinement through a tri-planar self-attention module and a cross-word cross-attention module, and inject the word-level feature t w into the three-dimensional expression of the deep learning network model to generate a three-dimensional model using the text. The method steps of this module are as Figure 2 shown, and the method structure and technical details are as Figure 3 shown;
[0055] (3) Put the text injection module based on the retention mechanism into a pre-trained three-dimensional model generation deep neural network model to inject text information between layers of the generator;
[0056] (4) Use the pre-trained CLIP loss and adversarial loss to train the modified model to construct the function of generating a three-dimensional model according to the text.
[0057] We transform the GET3D network model, as Figure 4 shown. This network consists of 6 stages, namely a mapping module, a shape generation module, a texture generation module, a three-dimensional model extraction module, a differentiable rendering module, and an image discriminator and contour discriminator for training.
[0058] Apply the proposed text feature extraction module to the mapping module part, as Figure 1 shown. The input text information is represented as S, and by using the text feature extraction module, S is converted into sentence-level feature t s and word-level feature t w , and the sentence-level text feature t sIntegrate the shape feature w of the mapping network output into the generator geo and the texture feature w tex .
[0059] In the stage of injecting text information into the network, we add the proposed text injection module based on the retention mechanism between each layer of the original network generator part for text information injection, as Figure 1 shown. The purpose of this step is to improve the coupling degree of text information and the network and strengthen the connection between features. The detailed method steps of the text injection module based on the retention mechanism are as Figure 2 shown, and the specific method structure and details are as Figure 3 shown. The generator combined with this module will finally infer f (g,t) .
[0060] Furthermore, use the 3D model extraction module and the differentiable rendering module in the original model to convert the three planes into pictures and contour masks and send them into the discriminator for judgment, and use the loss function combined with the corresponding text part to train the model:
[0061]
[0062]
[0063] Among them, the loss function is divided into three terms, namely the pixel discriminator loss s with the sentence-level text feature t and the mask discriminator loss Finally, in order to guide the model to learn text information, the CLIP loss is added, and λ is a hyperparameter.
[0064] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation modes of the present invention, and they are not used to limit the protection scope of the present invention. Any equivalent implementation modes or changes made without departing from the technology of the present invention should be included in the protection scope of the present invention.
Claims
1. A text-guided two-stage cascaded 3D model generation method generally includes the following steps: Step 1.1: Construct a 3D model representation method based on three planes. Use three matrices to represent the XY, XZ, and YZ planes respectively. Through spatial interpolation calculations with these three matrices, obtain the feature representation of the 3D space to generate a 3D model; Step 1.2: Construct a text feature extraction module, and use the encoder of the text part in the Contrastive Language-Image Pretraining (CLIP) model to convert the input text into a mixed text feature t, which is divided into a sentence-level feature t s and a word-level feature t w ; fuse t s into the shape feature w geo and the texture feature w tex output by the mapping network in the generator. t w is input as detailed information into the construction of a tri-planar self-attention module based on the retention mechanism; Step 1.3: Construct a three-plane self-attention module based on a retention mechanism. Utilize the plane feature map retention mechanism to improve the continuity of plane features. Construct geometric three-planes and texture three-planes respectively with the latent shape features and texture features generated by the generator; Step 1.4: Construct a cross-word cross-attention module. Input the global text information into the three-plane features output by the three-plane self-attention module based on the retention mechanism. Through cross-word attention, refine the shape and texture three-plane features at the word level. The finally generated three-plane expression will include 3D space information and word-level information; Step 1.5: Connect the three-plane self-attention module based on the retention mechanism and the cross-word cross-attention module to construct a two-stage cascaded text injection module. Utilize plane attention and cross-plane attention with a retention mechanism to refine the shape and texture three-planes at each layer of the generator, and accordingly generate word-level geometric three-planes and texture three-planes; Step 1.6: Construct a three-plane rendering and extraction module to decode the three-plane features obtained from the inference, and use the volume rendering method to render the three planes into RGB images for training. The model training is carried out by calculating the overall loss function which includes two parts: the generative adversarial loss and the CLIP loss. After the training is completed, the mesh object in traditional computer graphics is extracted by the isosurface extraction method (Marching Cubes, MC) in computer graphics.
2. The method of the three-plane self-attention module based on the retention mechanism in step 1.3 is as follows: Step 2.1: Select the output features of the layer above the generator (where f represents features, with superscripts g for geometry and t for texture, and subscript i - 1 representing the previous layer), and input them into the tri - plane attention module based on the retention mechanism; Step 2.2: The output features of the upper layer of the generator are combined with the three-plane features of the current layer to synthesize the feature f of the current layer i (g,t) , when calculating the features of the texture layer, the features of the shape layer are also added to the calculation to generate textures matching the corresponding geometric shapes. The calculation process is expressed as: Among them, The subscript i,xy represents the xy plane of the current layer. Similarly, i,yz and i,xz represent the yz plane and xz plane. These three planes are jointly used to represent the 3D model in the network model; Step 2.3: For the three-plane feature Calculate the plane self-attention respectively. The calculation process is as follows: Q = W q X K = W k X V = W v X Among them, Q, K, and V are three matrices in the attention mechanism, which are calculated from X, and W q,k,v are three weight matrices that the attention module needs to learn. Softmax is the normalized exponential function, and ⊙ represents element-wise multiplication of matrices. D 2d is the weight decay of the attenuation matrix module of the retention mechanism adaptively. mn is used to represent the coordinates of a point, and the Manhattan distance is used to represent the two-dimensional coordinates of the planar features; Calculate the fused three-plane features in the same way Among them, concat means to horizontally arrange the tensor parameters and splice them into a new matrix, and by default, expand on the second axis; Step 2.4: Use the three-plane self-attention features calculated by the module in Step 2.3 as V and K, and as Q, calculate the cross-plane attention as follows: Q = W q Y K = W k X V = W v X Among them, represents the dimension of the query set; The cross-plane attention features calculated separately are concatenated as and fed into the next module as follows: Among them, M2.4 is Module2.
4.
3. The method of the cross-word cross-attention module in step 1.4 is as follows: Step 3.1: Select the word-level feature t w and the cross-plane cross-attention feature of the previous module as the input of this module; Step 3.2: By calculating to obtain Q and t w calculating to obtain K and V, calculating cross-word attention, and dynamically refining t w into the three-dimensional representation of the model to obtain the three-plane features f refined at the word level i (g,t) , the specific method is as follows: K = V = W k t w Among them, W q With W k is the matrix feature that the module needs to learn.
4. The loss function used for model training in step 1.6 is as follows: Among them, The loss function is divided into three terms, namely the pixel discriminator loss incorporating the sentence-level text feature t s and the mask discriminator loss Finally, to guide the model to learn text information, a CLIP loss is added, where λ is a hyperparameter.