Method, system, and storage medium for generating three-dimensional target textures based on text
The multi-stage generative training process with iterative updating and visual angle preprocessing addresses semantic inaccuracies in 3D texture generation, enhancing accuracy and versatility.
Patent Information
- Application Number
- JP2025528812
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-31
- Filing Date
- 2024-03-08
- Publication Date
- 2025-10-30
- Estimated Expiration
- 2044-03-08
AI Technical Summary
Existing methods for generating 3D target textures based on text exhibit large semantic differences between generated content and input text, leading to inaccuracies.
A multi-stage generative training process involving 3D model data and descriptive text data, with preprocessing to include visual angle information, and iterative updating using similarity and update weight scores to refine the texture image, ensuring consistency and accuracy.
The method enhances the accuracy and continuity of generated 3D target textures by minimizing semantic discrepancies and reducing the need for extensive training data, while improving generalization and versatility.
Smart Images

Figure 2025536103000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a method, system, and storage medium for generating a three-dimensional target texture based on text. [Background technology]
[0002] AI technology has been deeply integrated into various industries and fields, injecting innovative power into traditional industries and achieving breakthroughs, while also driving the development of AI technology itself to a higher level. At present, AI-driven content generation technology has become a key force in the leap forward of the AI era. Through continuous and in-depth research and exploration, AI-driven content generation technology has achieved great results in many fields. This AI technology can creatively generate text, images, audio, and other content. Among these, generating image content based on text guidance has become a hot research and application direction.
[0003] Generating images based on text guidance involves inputting a personalized description and using artificial intelligence (AI) techniques to generate an image that matches the description. Most prior art relies on pre-trained large-scale models or generative adversarial networks to achieve text-guided image generation. With the development of text-guided image generation technology, many AI-based methods have been proposed to generate 3D target textures. Among these, methods for generating images based on text guidance stand out and can generate more creative 3D target textures.
[0004] However, existing methods for generating 3D target textures based on text are not very accurate due to large semantic differences between the generated content and the input text. Summary of the Invention [Problem to be solved by the invention]
[0005] The main objective of the present invention is to provide a method, system, and storage medium for generating a three-dimensional target texture based on text, in order to solve the technical problems existing in existing methods for generating a three-dimensional target texture based on text, which have a large semantic difference between the generated content and the input text and are not very accurate. [Means for solving the problem]
[0006] In order to achieve the above object, the present invention provides: 3D model data and corresponding descriptive text data a step of acquiring the descriptive text data JPEG2025536103000002.jpg413; the step in which JPEG2025536103000003.jpg413 contains visual angle information; obtaining a JPEG2025536103000004.jpg29148 image set; conducting multi-stage generative training, JPEG2025536103000005.jpg37148 projecting the image to obtain a two-dimensional image; 2D images and composite images Calculating a similarity and an update weight score with JPEG2025536103000006.jpg523, and outputting a texture image of the current stage based on the update weight score; Multi-stage iterative update is performed to obtain the repaired texture image. and obtaining a three-dimensional target texture image based on JPEG2025536103000007.jpg521 and the texture image obtained by performing multiple stages of iterative updating.
[0007] Optionally, 3D model data and descriptive text data JPEG2025536103000008.jpg413 has undergone preprocessing, and the preprocessing is performing spatial position processing on the initial 3D model data to obtain a mesh rendering image thereof in a 2D texture coordinate system; The initial three-dimensional model data and the initial description text data are combined according to a plurality of preset texture synthesis viewing angles to obtain three-dimensional model data, and a plurality of description text data including viewing angle information are generated. and outputting JPEG2025536103000009.jpg413, The number of stages in the multi-stage generative training is equal to the number of preset texture synthesis viewing angles.
[0008] Optionally, one composite image from the composite image set Specifically, the step of selecting JPEG2025536103000010.jpg523 is as follows: The first stage of multi-stage generative training involves selecting one synthetic image from the synthetic image set. JPEG2025536103000011.jpg523 is randomly selected, and in the next step, one composite image is selected from the composite image set based on the composite image selected in the previous step. The best option is to select JPEG2025536103000012.jpg523.
[0009] JPEG2025536103000013.jpg69148
[0010] JPEG2025536103000014.jpg29148 represents taking the Z value of the normal vector of the image. If the difference between the Z values corresponding to the two steps is greater than 0.3, the calculated pixel dot is the area that needs to be updated. JPEG2025536103000015.jpg27148 represents the interval, if the pixel dot interval is greater than 0.7, the corresponding pixel area is the area that needs to be updated.
[0011] Optionally, in the first stage of multi-stage generative training: JPEG2025536103000016.jpg518 is blank, JPEG2025536103000017.jpg518 is the entire texture area.
[0012] Optionally, the update weight score is calculated according to the following formula: JPEG2025536103000018.jpg13128, where source represents the update weight score, whose value range is [0,1], M represents the target plane region, and J represents a point within the target plane region. JPEG2025536103000019.jpg421 represents normalization, JPEG2025536103000020.jpg534 represents a rendering of 3D model data, JPEG2025536103000021.jpg543 represents a projection transformation for 3D model data.
[0013] Optionally, the texture image of the current stage is obtained by the formula: JPEG2025536103000022.jpg5142 where, JPEG2025536103000023.jpg412 represents the texture image at the current stage, N represents the region of the texture image where the target exists, and r is the update weight value.
[0014] Corresponding to the method for generating a three-dimensional target texture based on the text, the present invention comprises: 3D model data and corresponding descriptive text data A data acquisition module for acquiring JPEG2025536103000024.jpg413, VIEW a data acquisition module, the data acquisition module including the visual angle information; a composite image set generation module for obtaining a JPEG2025536103000025.jpg29148 image set; One composite image from the composite image set An image group difference resolution selection module that selects JPEG2025536103000026.jpg523, a rendering projection module that projects the image onto the JPEG2025536103000027.jpg52148 to obtain a two-dimensional image; 2D images and composite images an update weight score calculation module for calculating a similarity with JPEG2025536103000028.jpg523 and an update weight score; a current texture image generation module that outputs a texture image of the current stage based on the updated weight scores; Repair texture image and a 3D target texture image generation module that obtains a 3D target texture image based on JPEG2025536103000029.jpg520 and a texture image obtained by performing multiple stages of iterative updating.
[0015] Furthermore, in order to achieve the above object, the present invention further provides a computer-readable storage medium having stored thereon a program for generating a three-dimensional target texture based on text, the program performing the steps of the method for generating a three-dimensional target texture based on text when executed by a processor. [Effects of the Invention]
[0016] The beneficial effects of the present invention are as follows: (1) Compared with the prior art, the present invention solves the problem of large differences in sampling sources during texture generation by selecting a synthetic image, and can ensure the continuity of the final generated 3D target texture image. Furthermore, the entire texture generation process adopts a step-by-step method, and uses the calculation of the normal image to update the generated texture with large discontinuous areas. This strategy makes the generated texture image more fine and delicate. By combining it with the depth image as the input of the generative model, the generated texture can be more adapted to the target shape. In addition, the 2D image and the synthetic image can be easily combined. The similarity with JPEG2025536103000030.jpg523 and the update weight score are calculated, and the texture image of the current stage is output based on the update weight score, thereby improving the consistency between the texture image of the current stage and the synthesized image, which is equivalent to matching the semantics of the input text (i.e., describing the text data), and can make the accuracy of the generated three-dimensional target texture image more accurate. (2) Compared with the prior art, the present invention provides a method for preprocessing a plurality of written text data including visual angle information. Output JPEG2025536103000031.jpg413, 3D model data and corresponding descriptive text data Since JPEG2025536103000032.jpg413 is used as training data, a large amount of training data is not required. (3) Compared with the prior art, the present invention can eliminate the influence of differences in depth images in a depth image diffusion model by selecting a synthetic image. Differences in the input depth images tend to cause inconsistencies in shallow features such as color and pattern in the image output by the depth image diffusion model, which significantly affects the continuity of the generated texture. Therefore, by selecting a synthetic image, the influence of these differences can be avoided. (4) Compared with the prior art, the present invention has a higher accuracy in restoring texture images. Render and project the 3D model data based on JPEG2025536103000033.jpg520 to obtain a 2D image, and then combine the 2D image and the composite image. Calculate the similarity and update weight score with JPEG2025536103000034.jpg523, and output the current stage texture image based on the update weight score. Improve the matching accuracy of JPEG2025536103000035.jpg413. (5) Compared with the prior art, the present invention can significantly improve the generalization of content generation capabilities by using a pre-trained depth image diffusion model, and can quickly realize the method for generating 3D target textures based on text described in the present invention. The method for generating 3D target textures based on text described in the present invention not only has great advancement in effectiveness, but also has high versatility and practical application significance. [Brief explanation of the drawings]
[0017] The drawings described herein are intended to provide a further understanding of the present invention and constitute a part of the present invention, and the schematic examples of the present invention and the description thereof are for purposes of illustrating the present invention and are not intended to unduly limit the present invention. [Figure 1] 3 is a schematic flow chart of an embodiment of a method for generating a three-dimensional target texture based on text according to the present invention; [Figure 2] 2 is a schematic diagram of one embodiment of a method for generating a three-dimensional target texture based on text according to the present invention; FIG. [Figure 3] 1 is a block diagram of one embodiment of a system for generating a three-dimensional target texture based on text according to the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0018] In order to clarify the objectives, technical solutions and advantages of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. It is clear that the described embodiments are only a part of the embodiments of the present invention, and are not all of the embodiments. It should be understood that the specific embodiments described in this specification are for the purpose of explaining the present invention, and do not limit the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without any creative work fall within the protection scope of the present invention.
[0019] As shown in FIGS. 1 and 2, the method for generating a 3D target texture based on text according to the present invention is to generate a 3D model data and a corresponding descriptive text data. a step of acquiring the descriptive text data JPEG2025536103000036.jpg413; JPEG2025536103000037.jpg413 contains visual angle information. Rendering and projecting the three-dimensional model data based on JPEG2025536103000038.jpg104148JPEG2025536103000039.jpg520 to obtain a two-dimensional image; and and a step of calculating a similarity and an update weight score with JPEG2025536103000040.jpg523, and outputting a texture image of the current stage based on the update weight score. and obtaining a three-dimensional target texture image based on JPEG2025536103000041.jpg520 and the texture image obtained by performing multiple stages of iterative updating.
[0020] Preferably, a depth image diffusion model JPEG2025536103000042.jpg513 is pre-trained and uses a stable-diffusion-2-depth model.
[0021] Preferably, an image inpainting diffusion model JPEG2025536103000043.jpg513 uses the stable-diffusion-inpainting model.
[0022] By selecting a synthetic image, the present invention can solve the problem of large differences in sampling sources during texture generation and ensure the continuity of the final generated 3D target texture image. Furthermore, the entire texture generation process adopts a step-by-step method, and uses the calculation of the normal image to update the generated texture with large discontinuous areas. This strategy makes the generated texture image more detailed and delicate. By combining it with the depth image as the input of the generative model, the generated texture can be more closely matched to the target shape. In addition, the 2D image and the synthetic image The similarity with JPEG2025536103000044.jpg523 and the update weight score are calculated, and the texture image of the current stage is output based on the update weight score, thereby improving the consistency between the texture image of the current stage and the synthesized image, which is equivalent to matching the semantics of the input text (i.e., describing the text data), and making the accuracy of the generated three-dimensional target texture image more accurate.
[0023] In this embodiment, the three-dimensional model data and the descriptive text data JPEG2025536103000045.jpg413 has undergone preprocessing, which includes performing spatial position processing on the initial 3D model data to obtain a mesh rendering image thereof in a 2D texture coordinate system, and converting the target texture in the 3D space into a texture on a 2D plane through the spatial position processing; combining the initial 3D model data and the initial description text data according to a plurality of preset texture synthesis visual angles to obtain 3D model data; and combining the initial 3D model data and the initial description text data including visual angle information. and outputting JPEG2025536103000046.jpg413.
[0024] Preferably, the three-dimensional model data is obj model data. The plurality of preset texture synthesis viewing angles are two or more of six viewing angles: front, rear, left, right, top, and bottom.
[0025] The present invention provides a method for processing a plurality of written text data including visual angle information by preprocessing. Output JPEG2025536103000047.jpg413, 3D model data and corresponding descriptive text data Since JPEG2025536103000048.jpg413 is used as training data, a large amount of training data is not required.
[0026] Preferably, the number of stages of the multi-stage generative training is the same as the number of preset texture synthesis viewing angles. For example, if the number of preset texture synthesis viewing angles is four, the number of stages of the multi-stage generative training is also four.
[0027] In this embodiment, one composite image is selected from the composite image set. Specifically, the step of selecting JPEG2025536103000049.jpg523 involves selecting one synthetic image from the synthetic image set when performing the first stage of multi-stage generative training. JPEG2025536103000050.jpg523 is randomly selected, and in the next step, one composite image is selected from the composite image set based on the composite image selected in the previous step. The best option is to select JPEG2025536103000051.jpg523.
[0028] Preferably, in the next step, a composite image is selected from the composite image set based on the composite image selected in the previous step. The step of selecting JPEG2025536103000052.jpg523 specifically involves selecting the composite image that is complete, matches the depth image, and has the most similar image features to the composite image selected in the previous step. The best option is to select JPEG2025536103000053.jpg523.
[0029] In this embodiment, the selection of the composite image by the image difference resolution module may be specifically expressed by the following formula: JPEG2025536103000054.jpg982, JPEG2025536103000055.jpg40148
[0030] The present invention can eliminate the influence of differences in depth images in a depth image diffusion model by selecting a synthetic image. Differences in input depth images are likely to cause inconsistencies in shallow features such as color and pattern in the image output by the depth image diffusion model, which significantly affects the continuity of the generated texture. Therefore, by selecting a synthetic image, the influence of these differences can be avoided.
[0031] JPEG2025536103000056.jpg72148
[0032] Preferably, in the first stage of the multi-stage generative training, the initial texture image is a default background image, and in the next stage, the texture image is a texture image generated in the previous stage, which is a texture image obtained by expanding a synthetic image selected from the synthetic image set according to the three-dimensional model data.
[0033] JPEG2025536103000057.jpg29148 represents taking the z value of the normal vector of the image. If the difference between the z values corresponding to the two steps is greater than 0.3, the calculated pixel dot is the area that needs to be updated. JPEG2025536103000058.jpg573, JPEG2025536103000059.jpg539 is the initial texture image JPEG2025536103000060.jpg513 and the texture image at the ith iteration stage JPEG2025536103000061.jpg513 represents the pixel dot spacing, and if the pixel dot spacing is greater than 0.7, the corresponding pixel area is the area that needs to be updated.
[0034] In this embodiment, for the texture repair mask, it is determined whether iterative updating is required based on the judgment conditions. If the difference in Z values corresponding to two stages is greater than 0.3, it is determined that the texture repair mask requires iterative updating based on the normal angle. If the pixel dot spacing is greater than 0.7, it is determined that the corresponding pixel area (the area where the pixel dot spacing is greater than 0.7) is similar to the background, and it is determined that the texture repair mask requires iterative updating based on the pixel angle.
[0035] Preferably, the first stage of multi-stage generative training involves: JPEG2025536103000062.jpg518 is blank, JPEG2025536103000063.jpg518 is the entire texture area.
[0036] In this embodiment, the updated weight score is calculated by the following formula: JPEG2025536103000064.jpg14128, where source represents the update weight score, whose value ranges from [0, 1], and M is JPEG2025536103000065.jpg28147
[0037] In this embodiment, the larger the value of source, the larger the repair texture image. JPEG2025536103000066.jpg520 shows a better match with the semantics of the written text and also ensures the effect of the final reconstructed 3D texture.
[0038] In the present invention, the repaired texture image Render and project the 3D model data based on JPEG2025536103000067.jpg520 to obtain a 2D image, and then combine the 2D image and the composite image. Calculate the similarity and update weight score with JPEG2025536103000068.jpg523, and output the current stage texture image based on the update weight score. Improve the matching accuracy of JPEG2025536103000069.jpg413.
[0039] In this embodiment, the texture image at the current stage is obtained by the following formula: JPEG2025536103000070.jpg5143where, JPEG2025536103000071.jpg412 represents the texture image at the current stage, N represents the region of the texture image where the target (the three-dimensional target to be acted upon) exists, and r is the update weight value.
[0040] Preferably, r=0.6.
[0041] In this invention, the texture image of the current stage is retained for each stage and used for updating various stages in the following iterations, and only then can a complete 3D target texture image be finally output. This 3D target texture image can be based on the previous 3D model data and employs rasterization rendering to draw the 3D target into a pattern that matches the text input description.
[0042] By using a pre-trained depth image diffusion model, the present invention can greatly improve the generalization of content generation ability, and can quickly realize the method for generating three-dimensional target texture based on text described in the present invention. The method for generating three-dimensional target texture based on text described in the present invention not only has great advancement in effectiveness, but also has high versatility and practical application significance.
[0043] As shown in FIG. 3, the present invention provides a three-dimensional model data and a corresponding descriptive text data. A data acquisition module 100 for acquiring JPEG2025536103000072.jpg413, JPEG2025536103000073.jpg413 includes visual angle information, and the data acquisition module 100 and the three-dimensional model data and descriptive text data Based on the visual angle information of JPEG2025536103000074.jpg413, the depth image at the corresponding visual angle is JPEG2025536103000075.jpg516 and current normal image Generate JPEG2025536103000076.jpg520 and descriptive text data JPEG2025536103000077.jpg413 and depth image A synthetic image set generation module 200 inputs JPEG2025536103000078.jpg516 into a depth image diffusion model to obtain a set of synthetic images, and a synthetic image set generation module 201 extracts one synthetic image from the synthetic image set. JPEG2025536103000079.jpg62148 A rendering and projection module 304 that renders and projects to obtain a two-dimensional image, and a two-dimensional image and a composite image an update weight score calculation module 305 for calculating the similarity with JPEG2025536103000080.jpg523 and an update weight score; a current texture image generation module 306 for outputting a texture image at the current stage based on the update weight score; and a 3D target texture image generation module 400 for obtaining a 3D target texture image based on JPEG2025536103000081.jpg520 and the texture image obtained by performing multiple stages of iterative updating.
[0044] In this embodiment, the data acquisition module 100 further acquires an initial texture image JPEG2025536103000082.jpg528 and the current texture image Get JPEG2025536103000083.jpg528.
[0045] An embodiment of the present invention further provides a computer-readable storage medium, which may be included in the memory in the above embodiment or may be a computer-readable storage medium that exists independently of the device. The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to realize the method for generating a three-dimensional target texture based on the text shown in Figure 1. The computer-readable storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.
[0046] In this specification, each embodiment is described step by step, and each embodiment is described by focusing on the differences from other embodiments, and identical and similar parts between embodiments may be referred to. As for the apparatus, device, and storage medium embodiments, they are almost similar to the method embodiments, and therefore their explanations are relatively simple, and for related points, the explanation of the method embodiments may be referred to.
[0047] Furthermore, as used herein, the terms "comprise," "include," or any other variation thereof are intended to encompass a non-exclusive inclusion, such that a process, method, article, or device comprising a set of elements includes not only those elements but also other elements not expressly listed or inherent in such process, method, article, or device. Unless further limited, an element defined by the phrase "comprising one of..." does not exclude the presence of other identical elements in the process, method, article, or device that comprises that element.
[0048] Although the above description shows and describes the preferred embodiment of the present invention, it should not be considered that the present invention is limited to the form disclosed herein or that other embodiments are excluded, but rather that the present invention is applicable to various other combinations, modifications, and environments, and can be changed by the above teachings or the skill or knowledge of the relevant art within the scope of the concept of the present specification. Changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention are intended to be within the scope of protection of the appended claims.
Claims
1. 1. A method for generating a three-dimensional target texture based on text, comprising: 3D model data and corresponding descriptive text data obtaining the descriptive text data the step of: 3D model data and descriptive text data Based on the visual angle information, the depth image at the corresponding visual angle is and the current normal image Generate descriptive text data and depth images inputting the vectors σ to a depth image diffusion model to obtain a set of synthetic images; conducting multi-stage generative training, projecting to obtain a two-dimensional image; 2D images and composite images and calculating a similarity and an update weight score with the texture image of the current stage based on the update weight score; Multi-stage iterative update is performed to obtain the repaired texture image. and obtaining a three-dimensional target texture image based on the texture image obtained by performing multiple stages of iterative updating.
2. 3D model data and descriptive text data is pre-processed, and the pre-processing is performing spatial position processing on the initial 3D model data to obtain a mesh rendering image thereof in a 2D texture coordinate system; The initial three-dimensional model data and the initial description text data are combined according to a plurality of preset texture synthesis viewing angles to obtain three-dimensional model data, and a plurality of description text data including viewing angle information are generated. and outputting the 2. The method for generating a three-dimensional target texture based on text according to claim 1, wherein the number of stages of the multi-stage generative training is equal to the number of preset texture synthesis viewing angles.
3. One composite image from the composite image set Specifically, the step of selecting The first stage of multi-stage generative training involves selecting one synthetic image from the synthetic image set. In the next step, a synthetic image is selected from the synthetic image set based on the synthetic image selected in the previous step.
2. The method of claim 1, wherein the step of generating a three-dimensional target texture based on text is to select:
4. Initial texture image and the texture image at the ith iteration stage and obtaining Initial texture image and the i-th iteration texture image Calculated based on the texture update mask 2. The method of claim 1, wherein the text-based three-dimensional target texture is obtained by the steps of:
5. represents taking a z value of [0.01], and if the difference between the z values corresponding to the two stages is greater than 0.3, the calculated pixel dot is an area that needs to be updated; and is the initial texture image and the texture image at the ith iteration stage 2. The method for generating a three-dimensional target texture based on text according to claim 1, wherein the pixel dot spacing is greater than 0.7, and if the pixel dot spacing is greater than 0.7, the corresponding pixel area is an area that needs to be updated.
6. When conducting the first stage of multi-stage generative training, is blank, 6. The method of generating a three-dimensional target texture based on text of claim 5, wherein: is the entire texture region.
7. The updated weight score is calculated by the following formula: where source represents the update weight score, whose value range is [0, 1], M represents the target planar region, j represents a point within the target planar region, denotes normalization, represents rendering of the 3D model data, 2. The method of claim 1, wherein: represents a projective transformation for the three-dimensional model data.
8. The texture image at the current stage is obtained by the following formula: where:
8. The method of generating a three-dimensional target texture based on text of claim 7, wherein N represents the texture image of the current stage, N represents the area of the texture image where the target exists, and r is the update weight value.
9. 1. A system for generating a three-dimensional target texture based on text, comprising: 3D model data and corresponding descriptive text data A data acquisition module for acquiring the number of descriptive texts, view a data acquisition module, the data acquisition module including the visual angle information; Repair texture image a rendering and projection module for rendering and projecting the three-dimensional model data based on the above to obtain a two-dimensional image; 2D images and composite images an update weight score calculation module for calculating a similarity and an update weight score with the a current texture image generation module that outputs a texture image of the current stage based on the updated weight scores; Repair texture image and a three-dimensional target texture image generation module that obtains a three-dimensional target texture image based on the texture image obtained by performing multiple stages of iterative updating.
10. A computer-readable storage medium, comprising: A computer-readable storage medium having stored thereon a program for generating a three-dimensional target texture based on text, the program performing the steps of the method for generating a three-dimensional target texture based on text according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Pattern generation
CN113129399A
Texture generation method of virtual object, electronic equipment and storage medium
CN116485983A