Method and apparatus for generating texture of three-dimensional scene, readable storage medium, computer program product
By performing instance segmentation and viewing angle processing in a three-dimensional scene, combining diffusion model and multi-view technology, using instance layout and orientation prompt diagrams for texture generation, the problems of texture generation quality and efficiency in complex scenes are solved, and high-quality and consistent texture generation is achieved.
Patent Information
- Application Number
- CN202411455689.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-17
AI Technical Summary
The prior art is difficult to generate high-quality, consistent style three-dimensional scene textures in complex scenarios, and requires a lot of optimization time and memory resources.
By obtaining the scene grid and prompt text of a three-dimensional scene containing multiple instances, performing instance segmentation and viewing angle processing, using diffusion model and multi-view technology to generate textures, and combining instance layout and orientation prompt diagrams for precise control.
It realizes the generation of realistic and consistent textures in large-scale scenes, improves the quality and efficiency of texture generation, and solves texture seams and artifact problems.
Smart Images

Figure CN119444961B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the technical fields of computer vision and computer graphics, and more specifically, to a method and apparatus for generating textures of a three-dimensional scene, a readable storage medium, and a computer program product. Background Art
[0002] In the fields of computer vision and graphics, automatically generating high-quality textures for three-dimensional scenes is a fundamental task and is crucial for downstream applications such as games, movies, augmented reality, and digital twins. In recent years, with the significant progress of 2D diffusion models in text-to-texture synthesis, it has become possible to generate realistic textures for individual instances. However, when applied to more complex scenes, related technologies are difficult to maintain style consistency and semantic alignment, and at the same time require a large amount of optimization time and memory resources.
[0003] Currently, related methods rely on generative models trained on specific datasets to synthesize textures, which limits the generalization ability on other category instances and the generated texture diversity is limited. In addition, these methods usually require domain-specific expertise and cumbersome manual work, and creating high-quality textures remains a daunting and time-consuming task. Despite the significant progress recently made in mesh-based texture synthesis, related technologies still face challenges when dealing with large-scale scenes, such as texture seams and accumulated artifact problems. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method and apparatus for generating textures of a three-dimensional scene, a readable storage medium, and a computer program product, which can effectively solve the problem that related technologies cannot generate high-quality textures.
[0005] In one general aspect, a method for generating textures of a three-dimensional scene is provided, including: obtaining a scene mesh of a three-dimensional scene including a plurality of instances and a prompt text, where the prompt text includes global texture information of the three-dimensional scene and texture information of each instance among the plurality of instances; performing instance segmentation on the scene mesh to obtain a three-dimensional oriented bounding box of each instance; based on the three-dimensional oriented bounding box of each instance, obtaining a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewpoints; for any one of the predetermined number of viewpoints, performing the following processing: inputting the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewpoint, and the obtained target images into a diffusion model to obtain a target image of the three-dimensional scene in the current viewpoint, where the target image includes textures of the three-dimensional scene, the obtained target images are target images of other viewpoints obtained before the target image of the current viewpoint, and the obtained target images are empty when obtaining the target image of the first viewpoint; generating a three-dimensional scene including texture information based on all the target images.
[0006] Optionally, after obtaining the two-dimensional instance layout and two-dimensional orientation hint map of each instance in each of a predetermined number of perspectives based on the three-dimensional oriented bounding box of each instance, it further includes: obtaining the depth map and line drawing of the three-dimensional scene in each of the predetermined number of perspectives; wherein, inputting the hint text, the two-dimensional instance layout and two-dimensional orientation hint map of the current perspective, and the obtained target image into the diffusion model to obtain the target image of the three-dimensional scene in the current perspective, including: inputting the hint text, the two-dimensional instance layout of the current perspective, the two-dimensional orientation hint map, the line drawing and depth map, and the obtained target image into the diffusion model to obtain the target image of the three-dimensional scene in the current perspective.
[0007] Optionally, before generating a three-dimensional scene containing texture information based on all target images, it further includes: obtaining the latent images corresponding to the target images in a predetermined number of perspectives; adding noise to the latent images to obtain noisy latent images; inputting the noisy latent images into a local multi-view diffusion model to obtain refined target images.
[0008] Optionally, the local multi-view diffusion model performs a multi-step denoising process when processing the noisy latent images. Among them, before each step of the denoising process, the following operations are performed: projecting the noisy latent images into a temporary texture space; determining the overlapping regions of the latent images of different perspectives in the temporary texture space; performing weighted processing on the parts of the latent images of different perspectives that fall within the overlapping regions; replacing the corresponding parts of the latent images of the corresponding perspectives with the weighted parts.
[0009] Optionally, generating a three-dimensional scene containing texture information based on all target images includes: for each target image, after uniformly sampling in the target image, projecting the sampling points into the UV space to obtain the UV coordinates corresponding to each sampling point; inputting the UV coordinates into a hash model to obtain the texture features corresponding to each sampling point; performing average processing on the texture features to obtain the average feature of the target image; inputting the average features of all target images into a multi-layer perceptron to obtain a three-dimensional scene containing texture information.
[0010] Optionally, the diffusion model performs a multi-step denoising process when processing the prompt text, the two-dimensional instance layout and the two-dimensional orientation hint map of the current perspective, and the acquired target image. In each step of the denoising process, the following operations are performed: When the denoising process at the current step is for the instance level, nearest neighbor sampling is performed in the latent space, where the latent space stores the texture information of the instance, and the texture information of the instance is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective; obtain the first latent image of the image and add noise to the first latent image; based on the instance-level mask of the current instance, the first latent image, and the noised first latent image, obtain the first image; When the denoising process at the current step is for the scene level, bilinear sampling is performed in the RGB space, where the RGB space stores the global texture information, and the global texture information is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective, where when the denoising process at the current step is the first processing for the scene level, the image of the current perspective is the first image; obtain the second latent image of the image and add noise to the second latent image; based on the scene-level mask, the second latent image, and the noised second latent image, obtain the second image; where, after performing the denoising process for a predetermined number of steps, the second image is used as the target image of the current perspective.
[0011] Optionally, before each step of the denoising process, it further includes: in the case that the number of pixels included in any instance projected onto the current perspective in the three-dimensional scene exceeds a preset threshold, the texture information included in the target image corresponding to any instance is only stored in the latent space.
[0012] Optionally, based on the three-dimensional oriented bounding box of each instance, obtain the two-dimensional instance layout and the two-dimensional orientation hint map of each instance in each of a predetermined number of perspectives, including: normalizing the coordinates of the three-dimensional oriented bounding box of each instance and projecting them onto a predetermined number of perspectives to obtain the two-dimensional bounding boxes of each instance in different perspectives; based on the two-dimensional bounding box of each instance and the prompt text, obtain the two-dimensional instance layout of each instance in different perspectives; normalize the coordinates of the three-dimensional oriented bounding box of each instance and convert them to the RGB format, and project them onto a predetermined number of perspectives to obtain the two-dimensional orientation hint map of each instance in different perspectives.
[0013] In another general aspect, a texture generation device for a three-dimensional scene is provided, including: a first acquisition unit configured to acquire a scene grid and a prompt text of a three-dimensional scene including a plurality of instances, wherein the prompt text includes global texture information of the three-dimensional scene and texture information of each instance among the plurality of instances; a second acquisition unit configured to perform instance segmentation on the scene grid to obtain a three-dimensional oriented bounding box of each instance; a third acquisition unit configured to obtain a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewpoints based on the three-dimensional oriented bounding box of each instance; a processing unit configured to perform the following processing for any one of the predetermined number of viewpoints: input the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewpoint, and the acquired target image into a diffusion model to obtain a target image of the three-dimensional scene in the current viewpoint, wherein the target image includes the texture of the three-dimensional scene, the acquired target image is a target image of other viewpoints acquired before the target image of the current viewpoint, and the acquired target image is empty when acquiring the target image of the first viewpoint; a generation unit configured to generate a three-dimensional scene including texture information based on all the target images.
[0014] Optionally, the third acquisition unit is further configured to, after obtaining the two-dimensional instance layout and the two-dimensional orientation prompt map of each instance in each of a predetermined number of viewpoints based on the three-dimensional oriented bounding box of each instance, obtain a depth map and a line drawing of the three-dimensional scene in each of the predetermined number of viewpoints; wherein the processing unit is further configured to input the prompt text, the two-dimensional instance layout, the two-dimensional orientation prompt map, the line drawing and the depth map of the current viewpoint, and the acquired target image into a diffusion model to obtain a target image of the three-dimensional scene in the current viewpoint.
[0015] Optionally, the generation unit is further configured to, before generating a three-dimensional scene including texture information based on all the target images, obtain latent images corresponding to the target images in a predetermined number of viewpoints; add noise to the latent images to obtain latent images with noise; and input the latent images with noise into a local multi-view diffusion model to obtain refined target images.
[0016] Optionally, the local multi-view diffusion model performs a multi-step denoising process when processing the latent images with noise, wherein, before each step of the denoising process, the following operations are performed: project the latent images with noise onto a temporary texture space; determine the overlapping regions of the latent images of different viewpoints in the temporary texture space; perform weighted processing on the parts of the latent images of different viewpoints that fall within the overlapping regions; and replace the corresponding parts of the latent images of the corresponding viewpoints with the weighted parts.
[0017] Optionally, the generating unit is further configured to, for each target image, after uniformly sampling in the target image, project the sampling points into the UV space to obtain the UV coordinates corresponding to each sampling point; input the UV coordinates into a hash model to obtain the texture features corresponding to each sampling point; perform an averaging process on the texture features to obtain the average feature of the target image; and input the average features of all target images into a multi-layer perceptron to obtain a three-dimensional scene containing texture information.
[0018] Optionally, the diffusion model performs a multi-step denoising process when processing the prompt text, the two-dimensional instance layout and the two-dimensional orientation hint map of the current perspective, and the acquired target images. In each step of the denoising process, the following operations are performed: When the current step of the denoising process is for the instance level, perform nearest neighbor sampling in the latent space, where the latent space stores the texture information of the instance, and the texture information of the instance is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective; obtain a first latent image of the image and add noise to the first latent image; based on the instance-level mask of the current instance, the first latent image, and the first latent image after adding noise, obtain a first image; When the current step of the denoising process is for the scene level, perform bilinear sampling through the RGB space, where the RGB space stores the global texture information, and the global texture information is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective, where when the current step of the denoising process is the first processing for the scene level, the image of the current perspective is the first image; obtain a second latent image of the image and add noise to the second latent image; based on the scene-level mask, the second latent image, and the second latent image after adding noise, obtain a second image; where after performing the denoising process for a predetermined number of steps, the second image is used as the target image of the current perspective.
[0019] Optionally, before each step of the denoising process, it further includes: in the case where the number of pixels included in any instance after being projected onto the current perspective in the three-dimensional scene exceeds a preset threshold, the texture information included in the target image corresponding to any instance is only stored in the latent space.
[0020] Optionally, the third acquisition unit is further configured to normalize the coordinates of the three-dimensional oriented bounding box of each instance and project them onto a predetermined number of perspectives to obtain the two-dimensional bounding boxes of each instance in different perspectives; based on the two-dimensional bounding boxes of each instance and the prompt text, obtain the two-dimensional instance layouts of each instance in different perspectives; normalize the coordinates of the three-dimensional oriented bounding box of each instance and convert them to the RGB format, and project them onto a predetermined number of perspectives to obtain the two-dimensional orientation hint maps of each instance in different perspectives.
[0021] In another general aspect, a computer-readable storage medium storing instructions is provided, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the texture generation method for any one of the above three-dimensional scenes.
[0022] In another general aspect, a system including at least one computing device and at least one storage device storing instructions is provided, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the texture generation method for any one of the above three-dimensional scenes.
[0023] In another general aspect, a computer program product including computer instructions is provided, and when the computer instructions are executed by a processor, the texture generation method for any one of the above three-dimensional scenes is implemented.
[0024] In another general aspect, an electronic device including at least one memory and at least one processor is provided, and the at least one memory is configured to store a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor, the at least one processor is caused to execute the texture generation method for any one of the above three-dimensional scenes.
[0025] According to the texture generation method and apparatus, readable storage medium, and computer program product for a three-dimensional scene according to an embodiment of the present disclosure, precise control of a single instance is achieved by means of instance layout, that is, the instance layout is used as a control condition for texture generation, so that the generated texture is more consistent with the texture in the user-specified prompt text. Moreover, the present disclosure projects the three-dimensional scene onto a two-dimensional plane and uses the two-dimensional instance layout and two-dimensional orientation prompt map as control conditions for texture generation, generating the texture of the three-dimensional scene at the two-dimensional level, avoiding the problem that it is impossible to generate a texture of good quality directly at the three-dimensional level due to the insufficient generation quality of the three-dimensional baseline model itself. Furthermore, the present disclosure also considers that the direction prompt plays a crucial role in the texture generation of a single instance, and introduces an orientation prompt map as a control condition for texture generation, further improving the quality of the generated texture. Therefore, through the present disclosure, the problem that high-quality textures cannot be generated in the related art can be effectively solved.
[0026] Additional aspects and / or advantages of the general concept of the present disclosure will be set forth in part in the following description, and in part will be obvious from the description, or may be learned through the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Through the following description with reference to the drawings showing embodiments, the above and other objects and features of the embodiments of the present disclosure will become clearer, wherein:
[0028] Figure 1It is a flowchart showing a method for generating textures of a three-dimensional scene according to an embodiment of the present disclosure;
[0029] Figure 2 It is a schematic system flowchart showing a method for generating textures of a three-dimensional scene according to an embodiment of the present disclosure;
[0030] Figure 3 It is a schematic diagram showing the denoising process of a diffusion model according to an embodiment of the present disclosure;
[0031] Figure 4 It is a first effect comparison diagram showing an embodiment of the present disclosure and related technologies;
[0032] Figure 5 It is a second effect comparison diagram showing an embodiment of the present disclosure and related technologies;
[0033] Figure 6 It is a second effect comparison diagram showing an embodiment of the present disclosure and related technologies;
[0034] Figure 7 It is a fourth effect comparison diagram showing an embodiment of the present disclosure and related technologies;
[0035] Figure 8 It is a block diagram showing a texture generation device for a three-dimensional scene according to an embodiment of the present disclosure.
[0036] Figure 9 It is a schematic structural diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0037] The following detailed embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent. For example, the order of operations described herein is merely illustrative and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0038] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples described herein are provided only to illustrate some of the many feasible ways of implementing the methods, devices, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0039] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more of them.
[0040] Although terms such as "first", "second", and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, without departing from the teachings of the examples, the first component, first element, first region, first layer, or first part referred to in the examples described herein may also be referred to as the second component, second element, second region, second layer, or second part.
[0041] In the specification, when an element (such as a layer, region, or substrate) is described as "on" another element, "connected to" or "coupled to" another element, the element may be directly "on" the other element, directly "connected to" or "coupled to" the other element, or there may be one or more other elements intervening therebetween. In contrast, when an element is described as "directly on" another element, "directly connected to" or "directly coupled to" another element, there may be no other elements intervening therebetween.
[0042] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0043] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formal manner.
[0044] Furthermore, in the description of the examples, when a detailed description of a related structure or function that is considered well-known would cause an ambiguous interpretation of the present disclosure, such a detailed description will be omitted.
[0045] The present disclosure can be applied to computer graphics and computer vision technology scenarios, especially for texture generation under multiple perspectives, significantly enhancing the realism and consistency of image and video generation.
[0046] Currently, automatically generating high-quality textures for complex scenes remains a major challenge in computer graphics. Although text-to-texture synthesis using 2D diffusion models has produced good results for individual instances, it is still difficult to maintain style consistency and semantic alignment when applied to larger scenes, requiring a large amount of optimization time and memory. To address these challenges, the texture synthesis method of the present disclosure can create realistic and style-consistent textures for large scenes with multiple instances. Its core lies in instance-level controllable texture synthesis, that is, precise semantic control of individual instances is achieved by means of instance layout while ensuring the consistency of the overall style. Furthermore, the present disclosure also introduces a locally synchronized multi-view diffusion strategy to improve local texture consistency by sharing latent denoised content in small batches of adjacent-view images. In addition, the present disclosure also introduces neural MipTexture inspired by Mipmaps, which is specifically designed for scene texture mapping and aims to minimize aliasing.
[0047] The texture generation method, device, readable storage medium, and computer program product for a three-dimensional scene of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0048] The present disclosure proposes a texture generation method for a three-dimensional scene, aiming to synthesize a highly detailed and semantically aligned texture that is suitable for large-scale scene geometry and is not affected by local or global discontinuities. Figure 1 It is a flowchart showing the texture generation method for a three-dimensional scene according to an embodiment of the present disclosure.
[0049] Referring to Figure 1 , the texture generation method for the three-dimensional scene includes the following steps:
[0050] In step S101, a scene mesh and a prompt text of a three-dimensional scene containing multiple instances are obtained, where the prompt text includes global texture information of the three-dimensional scene and texture information of each instance among the multiple instances.
[0051] As an example, the initial input of the method of the present disclosure can be understood as triple. One input is the scene mesh of a three-dimensional scene containing multiple instances, which shows a textureless 3D scene. One input is a scene-level prompt text specifying the global style (i.e., global texture information) of the three-dimensional scene by the user. Another input is an instance-level prompt text specifying the desired appearance of each instance (i.e., the texture information of each instance above) by the user. It should be noted that the latter two texts are combined together as the prompt text in step S101.
[0052] As an example, in the above steps, the prompt text contains the texture information of each instance, and this texture information can specify the texture style of each instance to guide the diffusion model to generate textures that meet specific style and content requirements, so as to generate textures that conform to the instance description and are consistent in overall style.
[0053] According to an embodiment of the present disclosure, the description information of each instance in the three-dimensional scene (i.e., the texture information of each instance above) and the description information of the three-dimensional scene (i.e., the global texture information) can be obtained through manual setting or image guidance. For example, the image guidance acquisition method is to input images of multiple angles of the three-dimensional scene into the description information extraction unit to extract the corresponding description information, and the present disclosure does not limit this.
[0054] As an example, a texture unit (texel) is assigned to each vertex of the scene mesh to store the corresponding texture information, and the texture is mapped onto the scene mesh through inverse mapping, thereby providing a more realistic appearance for the three-dimensional scene; if some vertices are not assigned texture units, the Blender software intelligent UV mapping tool is used to map the texture onto the scene mesh.
[0055] As an example, after obtaining the scene mesh of the three-dimensional scene, the scene mesh can also be preprocessed, such as normal vector calculation, point cloud simplification, etc., to ensure the efficiency and accuracy of subsequent processing.
[0056] In step S102, instance segmentation is performed on the scene mesh to obtain the three-dimensional oriented bounding box of each instance.
[0057] As an example, each instance in the scene network is segmented, and then an automatic detection method can be used to detect the segmented scene network to generate a three-dimensional oriented bounding box (abbreviated as OBB) of each instance in the scene network. When obtaining two-dimensional information subsequently, the three-dimensional oriented bounding box can be used to control a powerful two-dimensional diffusion model due to its combination of geometric and semantic information, thereby creating consistent textures aligned with the scene layout.
[0058] In step S103, based on the three-dimensional oriented bounding box of each instance, the two-dimensional instance layout and two-dimensional position map of each instance in each of a predetermined number of viewpoints are obtained.
[0059] As an example, the above-mentioned predetermined number of viewpoints can be manually preset or preset in other ways, as long as all the points on the scene grid can be observed from these preset viewpoints, so as to ensure the generation of a complete texture. The above-mentioned preset viewpoints can be represented as a sequence of a series of camera poses, that is, a sequence composed of a series of three-dimensional rotation matrices and translation vectors.
[0060] As an example, one of the key points of the present disclosure is to introduce an instance layout as a control condition for texture generation. The instance layout contains the position of each instance and the corresponding hint text, as Figure 2 shown, which makes precise and flexible instance-level control possible. Specifically, for each instance, the 3D OBB of the instance can be normalized to the range [0, 1], and then the 3D OBB is rendered into a two-dimensional space according to the preset viewpoints, so as to obtain the two-dimensional instance layout of the corresponding viewpoints.
[0061] As an example, the orientation hint plays a crucial role in the texture generation of a single instance, which prompts the introduction of an orientation hint map as a direction guide. Without the orientation hint map, the generated results lack consistency and sharp edges will be generated during the repair process. However, it is not easy to derive a feasible orientation hint for a large-scale scene because different instances have different poses in a scene. For this reason, a pose-aware position mapping can be defined to represent the relative pose between the current camera and the instance. Specifically, for each instance, the 3D OBB of the instance can be normalized to the range [0, 1], and then converted to the RGB format, and the 3D OBB is rendered into a two-dimensional space according to the preset viewpoints, so as to obtain the two-dimensional orientation hint map of the corresponding viewpoints.
[0062] It should be noted that for any point in the orientation hint map, its value represents the relative position of the point in the 3D OBB of the corresponding instance, as Figure 2 shown. Therefore, for each view of the three-dimensional scene, an orientation hint map containing multiple instances will be generated as an additional control condition for texture generation.
[0063] According to an embodiment of the present disclosure, obtaining the 2D instance layout and the 2D orientation hint map of each instance in a predetermined number of viewpoints based on the 3D oriented bounding box of each instance may include: projecting the coordinates of the 3D oriented bounding box of each instance after normalization onto a predetermined number of viewpoints to obtain the 2D bounding box of each instance in different viewpoints (2D Bounding Box); obtaining the 2D instance layout of each instance in different viewpoints based on the 2D bounding box and the hint text of each instance; converting the coordinates of the 3D oriented bounding box of each instance into the RGB format after normalization and projecting them onto a predetermined number of viewpoints to obtain the 2D orientation hint map of each instance in different viewpoints. Through this embodiment, the 2D instance layout and the 2D orientation hint map of each instance in different viewpoints can be obtained conveniently and quickly.
[0064] Return Figure 1 , in step S104, for any one of the predetermined number of viewpoints, the following processing is performed: inputting the hint text, the 2D instance layout and the 2D orientation hint map of the current viewpoint, and the obtained target image into the diffusion model to obtain the target image of the 3D scene in the current viewpoint, where the target image contains the texture of the 3D scene, and the obtained target image is the target image of other viewpoints obtained before the target image of the current viewpoint, and the obtained target image is empty when obtaining the target image of the first viewpoint.
[0065] As an example, in order to add the above instance-level control conditions, that is, the 2D instance layout and the 2D orientation hint map, during the diffusion process, Instance Diffusion can be used as the basic model of the above diffusion model.
[0066] As an example, the orientation hint map can extract corresponding features by using a control network (ControlNet), that is, using ControlNet as the encoder in the diffusion model to extract the features of the orientation hint map. However, the ordinary ControlNet is pre-trained by copying the structure of Stable Diffusion, and directly applying it to Instance Diffusion will produce obvious artifacts. In order to better achieve geometric alignment, a fine-tuning of ControlNet can be performed by training an adapter for the orientation hint map in Instance Diffusion to avoid obvious artifacts. It should be noted that both Stable Diffusion and Instance Diffusion are a type of latent diffusion model.
[0067] As an example, when the current perspective is the first perspective, the obtained target image described above can be an empty feature; when the current perspective is the second perspective, the obtained target image described above is the target image of the first perspective; when the current perspective is the third perspective, the obtained target image described above is the target image of the first perspective and the target image of the second perspective; and so on. When the current perspective is the nth perspective, the obtained target image described above is the target images of the first perspective to the (n - 1)th perspective.
[0068] According to an embodiment of the present disclosure, after obtaining the two-dimensional instance layout and two-dimensional orientation hint map of each instance in each of a predetermined number of perspectives based on the three-dimensional oriented bounding box of each instance, it is also possible to obtain the depth map and line drawing of the three-dimensional scene in each of the predetermined number of perspectives; wherein, inputting the hint text, the two-dimensional instance layout and two-dimensional orientation hint map of the current perspective, and the obtained target image into the diffusion model to obtain the target image of the three-dimensional scene in the current perspective includes: inputting the hint text, the two-dimensional instance layout of the current perspective, the two-dimensional orientation hint map, the line drawing and depth map, and the obtained target image into the diffusion model to obtain the target image of the three-dimensional scene in the current perspective. Through this embodiment, the depth map and line drawing are introduced, so that the generated texture can be geometrically aligned with the three-dimensional scene.
[0069] As an example, the above two-dimensional instance layout and hint text can be directly input into Instance Diffusion. However, the texture of the synthesized size is not only restricted by the instance layout but also needs to be geometrically aligned with the scene. Therefore, for each perspective, a depth map and a line drawing can be rendered as geometric cues, as Figure 2 shown. In this embodiment, a set of perspectives (i.e., the above-mentioned predetermined number of perspectives) can be defined. For each perspective, the depth map, line drawing of the scene, and two-dimensional orientation hint map of the instance are rendered. These three conditions, together with the two-dimensional instance layout, are input into the diffusion model to generate the target image of each perspective.
[0070] As an example, the depth map and line drawing can be obtained in the following way: By rendering the geometric information of the scene grid according to the pose and internal parameters of the camera in the standard rasterization manner for a predetermined number of perspectives, what is obtained is multi-channel geometric cues, including multi-perspective depth maps and multi-perspective line drawings, that is, the depth map of each perspective and the line drawing of each perspective. The feature extraction of the depth map and line drawing can adopt the same way as the orientation hint Figure 1 and will not be elaborated here.
[0071] As an example, to ensure that the images generated by the diffusion model are consistent with the control conditions of the image domain, in this disclosure, for geometric conditions and camera pose conditions, a number of control networks (ControlNets) are trained using a large amount of 2D / 3D datasets with text to extract the features of depth maps, linear maps, and azimuth hint maps respectively. For instance layout, an additional model based on a cross-modal attention mechanism is trained on image data with instance-level semantics to extract the instance layout model.
[0072] In summary, through the rendering method, the present embodiment generates the image domain control conditions of the scene grid texture from multiple preset perspectives, mainly including the following conditions: geometric conditions (depth map and line drawing), text conditions (instance layout), and camera pose conditions (azimuth hint map).
[0073] According to an embodiment of the present disclosure, the diffusion model performs a multi-step denoising process when processing the prompt text, the two-dimensional instance layout and two-dimensional azimuth hint map of the current perspective, and the obtained target image. Among them, in each step of the denoising process, the following operations are performed: When the current step of the denoising process is for the instance level, nearest neighbor sampling is performed in the latent space, where the latent space stores the texture information of the instance, and the texture information of the instance is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective; obtain the first latent image of the image and add noise to the first latent image; based on the instance-level mask of the current instance, the first latent image, and the noise-added first latent image, obtain the first image; When the current step of the denoising process is for the scene level, bilinear sampling is performed in the RGB space, where the RGB space stores the global texture information, and the global texture information is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective, where when the current step of the denoising process is the first processing for the scene level, the image of the current perspective is the first image; obtain the second latent image of the image and add noise to the second latent image; based on the scene-level mask, the second latent image, and the noise-added second latent image, obtain the second image; Among them, after performing the denoising process for a predetermined number of steps, the second image is used as the target image of the current perspective. Through this embodiment, the consistency of the images generated from multiple perspectives can be preliminarily ensured.
[0074] As an example, after inputting the prompt text, the 2D instance layout of the current perspective, the 2D orientation hint map, the linear map, the depth map, and the acquired target image into the diffusion model, a specified UniFusion block can be used to pre-encode the 2D instance layout, determine the texture information and global texture information of the instance based on the encoded 2D instance layout, and map the texture information and global texture information of the instance to the latent space and the RGB space respectively, where the latent space is used for instance-level denoising and the RGB space is used for scene-level denoising. It should be noted that after using the instance layout as the input, the denoising process of the image in this embodiment is divided into two stages: instance-level denoising and scene-level denoising, as Figure 2 shown. Specifically, after receiving the 2D instance layout, first, a certain number of steps are used to denoise the instance alone to achieve instance-level denoising, then the scene-level latent representation is generated by merging the latent representations generated by instance-level denoising, and then, scene-level denoising can be performed by denoising the remaining steps based on the generated scene-level latent representation.
[0075] As an example, to initially ensure the consistency of the images generated from multiple perspectives, and to fuse the diffusion prior of instance-level image generation into the image inpainting architecture, this embodiment adopts a generation strategy of iterative image inpainting that includes two stages. During the instance-level denoising process, a latent space is maintained to store the texture information of the instance; during the scene-level denoising process, an RGB space is maintained to store the global texture information of the 3D scene. The above two spaces continuously update the stored texture information during the multi-perspective iterative image inpainting process. It should be noted that during each denoising step, it can be first determined whether the current step belongs to the instance-level denoising process or the scene-level denoising process.
[0076] As an example, after generating the texture of one perspective, that is, after generating the target image of one perspective, based on the generated target image, it is possible to continue generating the target images (which can also be called texture images) of other perspectives with global and local style consistency. If the directly applied generated target image in Instance Diffusion of a large-scale scene for subsequent image generation will produce serious artifacts, because introducing the scene-level context of previous frames during the instance-level denoising process will disrupt the generation of the images of different instances, resulting in information leakage between instances, that is: the prompt text of one instance will affect the appearance of another instance. In addition, due to the semantic complexity of the scene, only using scene-level inpainting increases the difficulty for the diffusion model to produce coordinated results, that is, it is challenging to draw the entire scene, which may cause significant inconsistencies in the generated image texture, making it difficult to infer reasonable textures.
[0077] To solve this problem, in this embodiment, the denoising process of generating images of multiple instances is decomposed into two stages: the instance-level denoising process and the scene-level denoising process, which largely solves the problems of information leakage and texture incoherence. To achieve two-stage image generation, separate masks can be generated for each instance and the entire scene. In addition, this embodiment also maintains two independent spaces, namely the latent space and the RGB space, which are used to record the texture information of instances and the global texture information of the three-dimensional scene respectively.
[0078] Specifically, when it is determined that the current step belongs to the instance-level denoising process, nearest neighbor sampling is performed through the maintained latent space. After projecting the sampled existing texture information onto the image in the current perspective, local noise is added to the sampled existing part, and the latent variables of the part to be generated are fused, thereby performing the denoising process of this stage, which is expressed in the following form:
[0079]
[0080] Among them, represents the inpainting mask corresponding to instance i, represents the latent variable generated in the denoising process of instance i at the t-th step, while represents the part of the latent variable generated after noise addition for instance i. It should be noted that the iterative process of the denoising process can be regarded as a reverse process, that is, the latent variable with T times of noise added is denoised T times.
[0081] When it is determined that the current step belongs to the scene-level denoising process, first, the latent variables of the above instance-level denoising process are obtained Furthermore, bilinear sampling is performed through the maintained RGB space. After projecting the sampled existing texture information onto the image in the current perspective, local noise is added to the sampled existing part, and the latent variables of the part to be generated are fused, thereby performing the denoising process of this stage, which is expressed in the following form:
[0082]
[0083] Among them, refers to the inpainting mask at the scene level, represents the latent variable generated in the denoising process of the three-dimensional scene at the t-th step, while represents the part of the latent variable generated after noise addition for the three-dimensional scene.
[0084] In summary, as Figure 3As shown, for each instance in a perspective, the instance-level denoising process first projects the known context of the instance (i.e., the sampled texture information mentioned above) onto the image in the current perspective. Then, in order to constrain the instance-level denoising process with the known context of the instance, noise can be added to the projected known context, and the denoised latent is mixed with the estimated denoised latent to obtain the updated latent variable of instance i. Then, the image of the 3D scene is generated through the remaining denoising steps, which are conditioned on the scene-level inpainting mask and the known context of the entire scene (i.e., the sampled texture information).
[0085] According to an embodiment of the present disclosure, before each denoising step of the above diffusion model, when the number of pixels included in the projection of any instance in the 3D scene onto the current perspective exceeds a preset threshold, the texture information included in the target image corresponding to the any instance is only stored in the latent space. Through this embodiment, selectively storing texture information can avoid unnecessary processing to optimize computing resources, thereby improving the generation efficiency.
[0086] As an example, multi-layer latent variables help reduce information leakage and achieve more consistent scene images with higher fidelity. Specifically, in a far perspective, although more complete textures are generated for the instance, due to its small size in the target image, the texture quality is low. In this embodiment, this low-quality texture can be retained as part of the initialization, that is, when viewing this example from a closer perspective, this initialization can serve as a consistency hint, and through inpainting, the texture can be transformed into a high-quality and style-consistent texture. This method helps reduce the complexity of the entire scene while ensuring the continuity of the generated texture through instance-level image generation operations.
[0087] As an example, this embodiment proposes the following strategy to select instances far away in the 3D scene and perform different instance-level and scene-level image generation operations on instances under different occlusion conditions. Specifically, first, a strict threshold is established to identify instances far away from the current camera viewpoint. When the threshold is greater than the number of pixels included in the projection of the instance onto the current perspective, it can be considered that the instance is far away from the current camera, and the texture included in the image generated by the instance is only projected onto the latent space. This selective projection can avoid unnecessary processing to optimize computing resources, thereby improving the generation efficiency. It should be noted that the number of pixels included in the projection of the instance onto the current perspective can be determined by the pixels included in the projection of the triangle of the instance onto the current perspective.
[0088] As an example, the above examples depend on their proximity to the viewpoint and occlusion state of the current camera.
[0089] Return Figure 1 , in step S105, a three-dimensional scene including texture information is generated based on all target images.
[0090] According to an embodiment of the present disclosure, before generating a three-dimensional scene including texture information based on all target images, potential images corresponding to target images at a predetermined number of viewpoints may also be obtained; noise is added to the potential images to obtain noisy potential images; and the noisy potential images are input into a local multi-view diffusion model to obtain refined target images. Through this embodiment, the target images are refined using the local multi-view diffusion model, which can improve the texture consistency of multi-view images.
[0091] As an example, for large-scale scene texture generation, it is usually difficult to maintain geometric and semantic consistency over a long trajectory. To improve the consistency of multi-view images, this embodiment introduces a local multi-view diffusion (LMVD) model to refine the target images generated previously.
[0092] Specifically, the LMVD model can be inserted into the image inpainting process of the above embodiment. For example, after the first batch of target images including images = {1, 2,...} are generated, each target image is encoded into a latent space by an encoder, and noise is gradually added to obtain a noisy latent image. The noisy latent image is input into the LMVD model to obtain refined target images, so as to refine the local details of each target image while maintaining its overall structure. Next, a denoising process can be performed through the latent space, which makes it possible to exchange information between different viewpoints.
[0093] As an example, the above LMVD model includes an encoder module, a diffusion module, and a decoder module. Among them, the encoder module extracts features from the input multi-view target images to generate a feature representation of the target image corresponding to the input; the diffusion model processes the feature representation output by the encoder module to generate high-quality textures that are multi-view consistent, that is, high-dimensional features; the decoder module maps the high-dimensional features generated by the diffusion module back to the original target image space to generate the final high-quality textures.
[0094] It should be noted that the encoder module can be composed of multiple parallel convolutional neural networks (CNNs) to process image data from different viewpoints respectively; the diffusion module can gradually optimize the details and consistency of the texture through iterative calculations; the decoder module can adopt a transposed convolution network to restore the high-dimensional features layer by layer to image textures.
[0095] According to an embodiment of the present disclosure, the local multi-view diffusion model performs a multi-step denoising process when processing a noisy latent image. Before each step of the denoising process, the following operations are performed: project the noisy latent image onto a temporary texture space; determine the overlapping regions of the latent images from different perspectives in the temporary texture space; weight the parts of the latent images from different perspectives that fall within the overlapping regions; and replace the corresponding parts of the latent images from the corresponding perspectives with the weighted parts. Through this embodiment, the latent denoised content is shared during the processing of a small batch of images, thereby further improving the local texture consistency and detail fidelity.
[0096] As an example, before each step of the denoising process, latent images can be shared between the target images from different perspectives to synchronize the diffusion process. For example, in each denoising step, all the noisy latent images can first be projected onto a temporary texture space (such as a UV texture space), and then information sharing can be achieved through the overlapping regions between the target images from different perspectives in the temporary texture space. It should be noted that the temporary texture space may contain multiple overlapping regions, and the target images involved in each overlapping region can be regarded as a batch, that is, the small batch of images in the above embodiment. Specifically, the latent values from different perspectives within the overlapping regions (i.e., the parts of the latent images that fall within the overlapping regions) are weighted and mixed. Finally, the weighted and mixed parts are mapped back to the corresponding parts of the latent images from the corresponding perspectives to generate updated latent images for denoising. It should be noted that the weights for the above weighted mixing can be calculated as the cosine value between the normal vector of the visible surface in the 3D grid and the viewing direction, and the present disclosure does not limit this.
[0097] In summary, better local consistency can be obtained by sharing the denoised content between local views. In addition, in order to ensure global consistency as much as possible, overlapping views can also be set between two consecutive small batch processes.
[0098] It should be noted that the UV space is a two-dimensional coordinate space related to the surface texture mapping of a 3D model. In the UV space, the U-axis and the V-axis represent two independent directions. Usually, U corresponds to the horizontal direction and V corresponds to the vertical direction. Their coordinate ranges are usually between 0 and 1. (0, 0) represents the upper left corner of the texture, and (1, 1) represents the lower right corner of the texture. This two-dimensional space is used to determine how each pixel on the texture image corresponds to the surface of the 3D model.
[0099] According to an embodiment of the present disclosure, generating a three-dimensional scene including texture information based on all target images may include: for each target image, after uniformly sampling in the target image, projecting the sampling points into the UV space to obtain the UV coordinates corresponding to each sampling point; inputting the UV coordinates into a hash model to obtain the texture features corresponding to each sampling point; averaging the texture features to obtain the average feature of the target image; and inputting the average features of all target images into a multi-layer perceptron to obtain a three-dimensional scene including texture information. Through this embodiment, the use of a hash model and multi-sampling technology can reduce the aliasing effect and ensure high fidelity and visual quality of texture synthesis.
[0100] As an example, with the primary generation of consistent multi-view images, a texture mapping needs to be unfolded for the entire scene, and directly projecting the generated two-dimensional images onto the texture space of a large-scale three-dimensional scene often faces aliasing problems, and noise artifacts will be generated due to the different distances between different instances. Traditionally, mipmapping is a very effective anti-aliasing technique that constructs a pre-filtered image pyramid by building texture images with gradually decreasing resolutions. During the rendering process, the pixel color is determined by selecting the most appropriate texture map according to the pixel and UV sizes.
[0101] And this embodiment proposes a Neural MipTexture network, that is, a multi-resolution hash model and mipmapping technology are used to unfold the texture in a large scene. In fact, as Figure 2 shown, for a given target image, uniformly sample within the pixels, and then project these sampling points onto the UV space to obtain the corresponding UV coordinates. Based on these UV coordinates, the texture features corresponding to each sampling point can be extracted through a pre-trained multi-resolution hash model for efficient lookup, and these texture features can be arranged, and each texture feature contains a high-dimensional feature vector. Then, these texture features are averaged to obtain an average feature (such as, miptexture feature). Then the average feature is fed into a compact multi-layer perceptron (MLP) to derive the final pixel color, that is, a three-dimensional scene including texture information is obtained. It should be noted that the entire optimization process can be driven by the L2 loss between the true color and the predicted color, and the present disclosure does not limit this.
[0102] The Neural MipTexture network proposed in this embodiment combines deep learning technology and traditional MipMapping technology. Through multi-level hash models and multi-sampling technology processing, the aliasing effect is reduced, and high fidelity and visual quality of texture synthesis are ensured. Its implementation steps may include:
[0103] Multi-level detail processing: Use a multi-resolution hash model to generate texture features at different levels. These feature levels can better retain and synthesize texture details by learning the multi-scale features of the image.
[0104] Multi-sampling technique: Uniformly sample each pixel of the target image multiple times, and input these multi-sampled sub-pixels into the multi-resolution hash model to obtain corresponding texture features. Replace the original pixel point features with the average features of the sampled sub-pixels to achieve the goal of pre-filtering. Then, input this average feature into a trainable multi-layer perceptron to predict the final texture color.
[0105] Jaggies reduction: Through the multi-resolution hash model and the multi-sampling technique, it is possible to identify and reduce jaggies in the texture, making the generated texture edges smoother and more natural, and improving the fidelity and detail performance of the texture.
[0106] In summary, this disclosure is oriented to 3D scenes. Through instance-level controllable texture synthesis, it uses instance layout to achieve precise semantic control of individual instances while maintaining overall style consistency, significantly improving the quality and efficiency of 3D scene texture synthesis, solving the challenges encountered in the prior art. Moreover, this disclosure also significantly improves the local consistency of the texture and reduces the aliasing effect through the local synchronous multi-view diffusion strategy and the neural MipTexture technique, bringing an innovative solution to the field of 3D scene texture synthesis.
[0107] Specifically, this disclosure performs instance segmentation on the input scene mesh, constructs an instance layout based on the segmentation result, and this instance layout includes the position (i.e., two-dimensional bounding box) and appearance style of each instance. Given that only processing each instance separately may affect the consistency of the scene style, an instance-conditioned image generation method based on a flexible image rendering framework can be adopted. To solve the common discontinuity problem in the restoration stage, the image denoising process can be decomposed into two stages: an instance-level denoising process and a scene-level denoising process. The former realizes precise control of each instance, and the latter ensures the overall style consistency. In addition, this disclosure develops a texture refinement strategy by innovatively integrating the local multi-view diffusion (LMVD) model into the image restoration framework. This strategy unifies the diffusion process between adjacent perspective images and collaboratively denoises to enhance local consistency. Finally, to efficiently project the generated multi-perspective images back to the UV atlas, neural MipTexture is introduced. This is a neural multi-scale texture mapping algorithm specifically designed to create high-fidelity texture maps. The neural MipTexture of this disclosure is specifically designed to solve the aliasing problem often encountered in large-scale scene textures.
[0108] Through all the above embodiments, the advantages of this disclosure are summarized as follows:
[0109] 1. Improved the quality of texture generation: Through the synchronous diffusion of multi-view images and integrating the information of multi-view images, high-quality texture generation is achieved. Compared with traditional methods, the present disclosure has significant advantages in terms of texture details, color consistency, geometric consistency, etc.
[0110] 2. Extended the application scope: The generated high-quality texture images can be applied to the texture covering of various 3D models and are widely used in fields such as game development, virtual reality, and film and television production. At the same time, multi-view synchronous diffusion can also be extended to other fields such as medical image processing and remote sensing image analysis, having broad application value.
[0111] 3. Improved the processing efficiency: By optimizing the algorithm structure and processing flow, the present disclosure improves the processing efficiency of texture generation and can generate high-quality texture images in a relatively short time to meet the requirements of practical applications.
[0112] 4. Enhanced the robustness of the system: Through the comprehensive application of steps such as feature matching, feature fusion, and diffusion processing, the present disclosure enhances the robustness of the system and can achieve stable texture generation effects between different perspective images.
[0113] The method for generating textures of a three-dimensional scene in an embodiment of the present disclosure has important innovation and practicality in the field of texture generation, can significantly improve the generation efficiency and quality of textures of a three-dimensional scene, solve problems such as local texture inconsistency, detail loss, and jagged effects in the prior art, can significantly enhance the visual effect of a 3D model, and meet the requirements for high-quality textures in fields such as game development, virtual reality, film and television production, and digital display. At the same time, the method structure of the present disclosure is scientifically reasonable, the processing flow is rigorous and efficient, and it has broad application prospects and commercial value. Through the above detailed implementation scheme, the present disclosure not only achieves high-quality texture generation but also provides a solid foundation for further research and development of the generation and application of high-quality three-dimensional scenes in the future.
[0114] The present disclosure has conducted experiments on a large number of indoor and outdoor scenes, all indicating that compared with related texture generation technologies, the method of the present disclosure can generate high-quality and consistent textures, superior to existing texture generation methods. The experimental effect comparisons are as Figures 4 to 7 shown.
[0115] Figure 8 is a block diagram showing a texture generation device for a three-dimensional scene according to an embodiment of the present disclosure, as Figure 8 shown, the device includes a first acquisition unit 80, a second acquisition unit 82, a third acquisition unit 84, a processing unit 86, and a generation unit 88.
[0116] The first acquisition unit 80 is configured to acquire a scene grid and prompt text of a three-dimensional scene including multiple instances, where the prompt text includes global texture information of the three-dimensional scene and texture information of each instance among the multiple instances; the second acquisition unit 82 is configured to perform instance segmentation on the scene grid to acquire a three-dimensional oriented bounding box of each instance; the third acquisition unit 84 is configured to, based on the three-dimensional oriented bounding box of each instance, acquire a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewpoints; the processing unit 86 is configured to perform the following processing for any one of the predetermined number of viewpoints: input the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewpoint, and the acquired target image into a diffusion model to obtain a target image of the three-dimensional scene in the current viewpoint, where the target image includes the texture of the three-dimensional scene, the acquired target image is the target image of other viewpoints acquired before the target image of the current viewpoint, and the acquired target image is empty when acquiring the target image of the first viewpoint; the generation unit 88 is configured to generate a three-dimensional scene including texture information based on all the target images.
[0117] According to an embodiment of the present disclosure, the third acquisition unit 84 is further configured to, after acquiring the two-dimensional instance layout and the two-dimensional orientation prompt map of each instance in each of a predetermined number of viewpoints based on the three-dimensional oriented bounding box of each instance, acquire a depth map and a line drawing of the three-dimensional scene in each of the predetermined number of viewpoints; where the processing unit is further configured to input the prompt text, the two-dimensional instance layout, the two-dimensional orientation prompt map, the line drawing and the depth map of the current viewpoint, and the acquired target image into a diffusion model to obtain a target image of the three-dimensional scene in the current viewpoint.
[0118] According to an embodiment of the present disclosure, the generation unit 88 is further configured to, before generating a three-dimensional scene including texture information based on all the target images, acquire latent images corresponding to the target images in a predetermined number of viewpoints; add noise to the latent images to obtain noisy latent images; and input the noisy latent images into a local multi-view diffusion model to obtain refined target images.
[0119] According to an embodiment of the present disclosure, the local multi-view diffusion model performs a multi-step denoising process when processing the noisy latent images, where, before each step of the denoising process, the following operations are performed: project the noisy latent images onto a temporary texture space; determine the overlapping regions of the latent images of different viewpoints in the temporary texture space; perform weighted processing on the parts of the latent images of different viewpoints that fall within the overlapping regions; and replace the corresponding parts of the latent images of the corresponding viewpoints with the weighted parts.
[0120] According to an embodiment of the present disclosure, the generation unit 88 is further configured to, for each target image, after uniformly sampling in the target image, project the sampling points into the UV space to obtain the UV coordinates corresponding to each sampling point; input the UV coordinates into the hash model to obtain the texture features corresponding to each sampling point; perform an averaging process on the texture features to obtain the average feature of the target image; input the average features of all target images into a multi-layer perceptron to obtain a three-dimensional scene containing texture information.
[0121] According to an embodiment of the present disclosure, the diffusion model performs a multi-step denoising process when processing the prompt text, the two-dimensional instance layout and the two-dimensional orientation hint map of the current perspective, and the acquired target images. In each step of the denoising process, the following operations are performed: When the current step of the denoising process is for the instance level, perform nearest neighbor sampling in the latent space, where the latent space stores the texture information of the instance, and the texture information of the instance is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective; obtain the first latent image of the image and add noise to the first latent image; based on the instance-level mask of the current instance, the first latent image, and the first latent image after adding noise, obtain the first image; When the current step of the denoising process is for the scene level, perform bilinear sampling through the RGB space, where the RGB space stores the global texture information, and the global texture information is determined based on the two-dimensional instance layout; project the sampled texture information onto the image of the current perspective, where when the current step of the denoising process is the first processing for the scene level, the image of the current perspective is the first image; obtain the second latent image of the image and add noise to the second latent image; based on the scene-level mask, the second latent image, and the second latent image after adding noise, obtain the second image; wherein, after performing the denoising process for a predetermined number of steps, the second image is used as the target image of the current perspective.
[0122] According to an embodiment of the present disclosure, before each step of the denoising process, it further includes: in the case where the number of pixels included in any instance projected onto the current perspective in the three-dimensional scene exceeds a preset threshold, the texture information included in the target image corresponding to any instance is only stored in the latent space.
[0123] According to an embodiment of the present disclosure, the third acquisition unit 84 is further configured to normalize the coordinates of the three-dimensional oriented bounding box of each instance and project them onto a predetermined number of perspectives to obtain the two-dimensional bounding boxes of each instance in different perspectives; based on the two-dimensional bounding boxes of each instance and the prompt text, obtain the two-dimensional instance layout of each instance in different perspectives; normalize the coordinates of the three-dimensional oriented bounding box of each instance and convert them into the RGB format, and project them onto a predetermined number of perspectives to obtain the two-dimensional orientation hint maps of each instance in different perspectives.
[0124] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for generating textures of a three-dimensional scene as described above. Examples of such computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0125] According to an embodiment of the present disclosure, there is provided a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the method for generating textures of a three-dimensional scene as described above.
[0126] According to an embodiment of the present disclosure, there is provided a computer program product including computer instructions that, when executed by a processor, implement the method for generating textures of a three-dimensional scene as described above.
[0127] Figure 9 is a schematic structural diagram of an electronic device showing an embodiment of the present disclosure, as Figure 9As shown, the electronic device includes at least one memory 930 and at least one processor 910. The at least one memory 930 is used to store a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor 910, it causes the at least one processor 910 to execute any of the above texture generation methods for three-dimensional scenes. The electronic device transmits data internally through a communication bus 940 and interacts with other devices through a communication interface 920.
[0128] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A texture generation method for a three-dimensional scene, characterized in that: include: Acquire a scene mesh and a prompt text of a three-dimensional scene including a plurality of instances, wherein the prompt text includes global texture information of the three-dimensional scene and texture information of each of the plurality of instances; Performing instance segmentation on the scene grid to obtain a three-dimensional oriented bounding box for each instance; Based on the three-dimensional directional bounding box of each instance, obtaining a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewing angles; For any of the predetermined number of viewing angles, the following processing is performed: the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewing angle, and the acquired target image are input into a diffusion model to obtain a target image of the three-dimensional scene at the current viewing angle, wherein the target image includes a texture of the three-dimensional scene, and the acquired target image is a target image of other viewing angles acquired before the target image of the current viewing angle, and the acquired target image is empty when the target image of the first viewing angle is acquired; Based on all target images, generating the three-dimensional scene containing texture information; The step of obtaining a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewing angles based on the three-dimensional directional bounding box of each instance includes: Normalizing the coordinates of the three-dimensional oriented bounding box of each instance and projecting them to the predetermined number of viewing angles to obtain a two-dimensional bounding box of each instance at different viewing angles; Based on the two-dimensional bounding box and prompt text of each instance, a two-dimensional instance layout of each instance at different viewing angles is obtained; The coordinates of the three-dimensional directional bounding box of each instance are normalized and converted into RGB format, and projected to the predetermined number of viewing angles to obtain a two-dimensional orientation prompt map of each instance at different viewing angles.
2. The texture generation method according to claim 1, characterized in that: After obtaining a two-dimensional instance layout and a two-dimensional orientation hint map of each instance in each of a predetermined number of viewing angles based on the three-dimensional directional bounding box of each instance, the method further includes: Acquire a depth map and a line map of the three-dimensional scene at each of the predetermined number of viewing angles; The step of inputting the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewing angle, and the acquired target image into the diffusion model to obtain the target image of the three-dimensional scene at the current viewing angle includes: The prompt text, the two-dimensional instance layout of the current viewing angle, the two-dimensional orientation prompt map, the linear map and the depth map, and the acquired target image are input into the diffusion model to obtain the target image of the three-dimensional scene at the current viewing angle.
3. The texture generation method according to claim 1, characterized in that: Before generating the three-dimensional scene containing texture information based on all target images, the method further includes: Acquire potential images corresponding to the target images under the predetermined number of viewing angles; adding noise to the latent image to obtain a latent image containing noise; The noisy latent image is input into a local multi-view diffusion model to obtain a refined target image.
4. The texture generation method according to claim 3, characterized in that: The local multi-view diffusion model performs a multi-step denoising process when processing a noisy latent image, wherein before each step of the denoising process, the following operations are performed: Projecting the noisy latent image into a temporary texture space; Determining overlapping areas of potential images at different perspectives in the temporary texture space; weighting the parts of the latent images of different perspectives falling in the overlapping area; The weighted portion is used to replace the corresponding portion of the latent image of the corresponding perspective.
5. The texture generation method according to claim 1, characterized in that: The step of generating the three-dimensional scene containing texture information based on all target images comprises: For each target image, after uniform sampling in the target image, the sampling points are projected into the UV space to obtain the UV coordinates corresponding to each sampling point; the UV coordinates are input into the hash model to obtain the texture features corresponding to each sampling point; the texture features are averaged to obtain the average features of the target image; The average features of all target images are input into a multilayer perceptron to obtain the three-dimensional scene containing texture information.
6. The texture generation method according to claim 1, characterized in that: The diffusion model performs a multi-step denoising process when processing the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewing angle, and the acquired target image, wherein in each step of the denoising process, the following operations are performed: When the current denoising process is for the instance level, nearest neighbor sampling is performed in the latent space, wherein the latent space stores texture information of the instance, and the texture information of the instance is determined based on the two-dimensional instance layout; the sampled texture information is projected onto the image of the current perspective; a first latent image of the image is obtained, and the first latent image is denoised; a first image is obtained based on the instance level mask of the current instance, the first latent image, and the first latent image after denoising; When the current denoising process is for the scene level, bilinear sampling is performed through the RGB space, wherein the RGB space stores global texture information, and the global texture information is determined based on the two-dimensional instance layout; the sampled texture information is projected to the image of the current perspective, wherein when the current denoising process is the first processing for the scene level, the image of the current perspective is the first image; a second latent image of the image is obtained, and the second latent image is denoised; a second image is obtained based on the scene level mask, the second latent image and the denoised second latent image; After executing a denoising process for a predetermined number of steps, the second image is used as the target image of the current viewing angle.
7. The texture generation method according to claim 6, characterized in that: Before each denoising process, it also includes: When the number of pixels contained in any instance of the three-dimensional scene after being projected to the current viewing angle exceeds a preset threshold, texture information contained in the target image corresponding to the any instance is only stored in the latent space.
8. A texture generation device for a three-dimensional scene, characterized in that: include: A first acquisition unit is configured to acquire a scene mesh and a prompt text of a three-dimensional scene including a plurality of instances, wherein the prompt text includes global texture information of the three-dimensional scene and texture information of each of the plurality of instances; A second acquisition unit is configured to perform instance segmentation on the scene grid and acquire a three-dimensional oriented bounding box of each instance; A third acquisition unit is configured to acquire a two-dimensional instance layout and a two-dimensional orientation prompt map of each instance in each of a predetermined number of viewing angles based on the three-dimensional directional bounding box of each instance; The processing unit is configured to perform the following processing for any one of the predetermined number of viewing angles: input the prompt text, the two-dimensional instance layout and the two-dimensional orientation prompt map of the current viewing angle, and the acquired target image into a diffusion model to obtain a target image of the three-dimensional scene at the current viewing angle, wherein the target image includes a texture of the three-dimensional scene, and the acquired target image is a target image of other viewing angles acquired before the target image of the current viewing angle, and the acquired target image is empty when the target image of the first viewing angle is acquired; A generating unit, configured to generate the three-dimensional scene containing texture information based on all target images; Among them, the third acquisition unit is also configured to normalize the coordinates of the three-dimensional directional bounding box of each instance and project them to the predetermined number of viewing angles to obtain the two-dimensional bounding box of each instance at different viewing angles; based on the two-dimensional bounding box and prompt text of each instance, obtain the two-dimensional instance layout of each instance at different viewing angles; normalize the coordinates of the three-dimensional directional bounding box of each instance and convert them into RGB format, and project them to the predetermined number of viewing angles to obtain the two-dimensional orientation prompt map of each instance at different viewing angles.
9. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the texture generation method for a three-dimensional scene as claimed in any one of claims 1 to 7.
10. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, the at least one computing device is prompted to execute the texture generation method for a three-dimensional scene according to any one of claims 1 to 7.
11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the texture generation method for a three-dimensional scene according to any one of claims 1 to 7 is implemented.
12. An electronic device, characterized in that: It comprises at least one memory and at least one processor, wherein the at least one memory is used to store a set of computer executable instructions, and when the set of computer executable instructions is executed by the at least one processor, it prompts the at least one processor to execute the texture generation method for a three-dimensional scene as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic texture construction method for white mold based on live-action three dimensions
CN114998503A
Three-dimensional texture image generation method and device, computer equipment and storage medium
CN116977531A