Texture generation method and apparatus, and computer device, storage medium and program product
By acquiring the guidance information and mesh file of the geometric model, using a depth rendering camera to render depth maps from multiple perspectives, and combining texture information for texture rendering, the problem of low texture template adaptability is solved, and the generation quality and adaptability of texture maps are improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-11-21
- Publication Date
- 2026-06-04
AI Technical Summary
In existing texture generation schemes, the adaptability of fixed texture templates is low, resulting in poor texture map generation quality.
By acquiring the guidance information and mesh file of the geometric model, a depth map is rendered from multiple perspectives using a depth rendering camera, and texture rendering is performed by combining texture information and depth map to generate a texture map consistent with the preset texture style.
It improves the quality of texture mapping generation, making the generated texture style consistent with the preset texture style and better adapting to the geometry of the geometric model.
Smart Images

Figure CN2025136758_04062026_PF_FP_ABST
Abstract
Description
Texture generation methods, apparatus and computer equipment, storage media and program products
[0001] This application claims priority to Chinese Patent Application No. 202411751322.X, filed on November 28, 2024, entitled “Texture Generation Method, Apparatus, Computer Equipment, Storage Medium, and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and more particularly to the field of texture generation technology, specifically to a texture generation method, a texture generation apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0003] Texture generation techniques can generate texture maps for geometric models. Mapping these texture maps onto the surface of the geometric model can simulate its material and appearance. Current texture generation schemes use fixed texture templates for simple transformations to obtain texture maps of the geometric model. In such schemes, the texture styles of the fixed templates are limited, and the compatibility between the fixed templates and the geometric model is low, resulting in low-quality texture maps. Therefore, improving the quality of generated texture maps has become a current research hotspot. Summary of the Invention
[0004] This application provides a texture generation method, apparatus, computer device, storage medium, and program product, which can improve the generation quality of texture maps.
[0005] On one hand, embodiments of this application provide a texture generation method, which includes:
[0006] Obtain the guidance information and mesh file of the geometric model. The guidance information contains texture information, which is used to indicate the preset texture style.
[0007] Based on the mesh file of the geometric model, at least one depth rendering camera is used to perform depth rendering on the geometric model from at least one viewpoint to obtain the depth map of the geometric model in each viewpoint; at least one depth rendering camera corresponds to one viewpoint respectively;
[0008] Based on the texture information in the guidance information and the depth map of the geometric model in each view, the geometric model is processed by texture rendering to obtain the texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
[0009] Accordingly, embodiments of this application provide a texture generation apparatus, which includes:
[0010] The acquisition unit is used to acquire the guidance information and mesh file of the geometric model. The guidance information includes texture information, which is used to indicate the preset texture style.
[0011] The processing unit is used to perform depth rendering processing on the geometric model from at least one viewpoint using at least one depth rendering camera based on the mesh file of the geometric model, and to obtain the depth map of the geometric model in each viewpoint; the at least one depth rendering camera corresponds to one viewpoint respectively;
[0012] The processing unit is also used to perform texture rendering on the geometric model based on the geometric information in the guidance information and the depth map of the geometric model in each viewpoint, to obtain the texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
[0013] Accordingly, embodiments of this application provide a computer device, which includes:
[0014] A processor is a tool for implementing computer programs.
[0015] A computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the texture generation method described above.
[0016] Accordingly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when read and executed by a processor of a computer device, causes the computer device to perform the texture generation method described above.
[0017] Accordingly, this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the texture generation method described above.
[0018] In this embodiment, the texture information in the guidance information of the geometric model can be used to indicate a preset texture style, and the depth map of the geometric model at each viewpoint can be used to indicate the geometric information of the geometric model at each viewpoint. Based on the texture information in the guidance information and the depth map of the geometric model at each viewpoint, texture rendering processing can be performed on the geometric model to obtain a texture map of the geometric model. It can be seen that the guidance information in this embodiment provides texture generation guidance, specifically texture style generation guidance. Using the texture information in the guidance information to guide the generation of the geometric model's texture ensures that the generated actual texture style is consistent with the preset texture style indicated by the texture information. That is, given any preset texture style, based on the guidance of this texture style, a texture map with a texture style consistent with the given preset texture style can be generated. Therefore, the generation quality of the texture map can be improved from the consistency dimension between the generated actual texture style and the given preset texture style. Furthermore, this embodiment utilizes the depth map of the geometric model at each viewpoint to guide texture generation. The depth map of the geometric model at each viewpoint acts as a geometric guide during texture generation, enabling the generated texture map to better adapt to the geometric model. This improves the quality of texture map generation by addressing the compatibility between the generated texture style and the geometric model. In summary, this embodiment can improve the quality of texture map generation. Attached Figure Description
[0019] Figure 1 is a schematic diagram illustrating the application of a texture map according to an embodiment of this application;
[0020] Figure 2a is a schematic diagram of an interface for uploading geometric models provided in an embodiment of this application;
[0021] Figure 2b is a schematic diagram of an interface for inputting guidance information provided in an embodiment of this application;
[0022] Figure 2c is a schematic diagram of an interface for accepting texture generation effects provided in an embodiment of this application;
[0023] Figure 3 is a schematic diagram of the architecture of a texture generation system provided in an embodiment of this application;
[0024] Figure 4 is a flowchart illustrating a texture generation method provided in an embodiment of this application;
[0025] Figure 5 is a schematic diagram of a text-to-image conversion process provided in an embodiment of this application;
[0026] Figure 6 is a schematic diagram of a depth rendering process provided in an embodiment of this application;
[0027] Figure 7 is a flowchart illustrating another texture generation method provided in an embodiment of this application;
[0028] Figure 8 is a schematic diagram of the large visual model provided in an embodiment of this application;
[0029] Figure 9 is a comparison diagram of a texture generation effect provided in an embodiment of this application;
[0030] Figure 10 is a comparison diagram of another texture generation effect provided by an embodiment of this application;
[0031] Figure 11 is a schematic diagram of a texture generation device provided in an embodiment of this application;
[0032] Figure 12 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0033] To better understand the technical solutions provided in the embodiments of this application, the technical terms involved in the embodiments of this application will be introduced first:
[0034] I. Geometric Model:
[0035] A geometric model refers to a model possessing a geometric shape, which refers to a geometric shape or outline. In the embodiments of this application, the geometric model can be a three-dimensional model, which refers to a geometric object in three-dimensional space, obtained through three-dimensional modeling software. A three-dimensional model visually possesses not only depth and width but also height, providing a more realistic and three-dimensional visual effect. In the embodiments of this application, a geometric model refers to the geometric model of an object, representing the geometric shape of the object; for example, in a game scene, the object can be a game character, and the geometric model can be a model of the game character; similarly, within a game scene, the object can be a game prop, and the geometric model can be a model of the game prop; furthermore, in the field of architectural visualization, the object can be a building, and the geometric model can be a model of the building.
[0036] A mesh is the basic structure of a geometric model. A geometric model is usually composed of multiple interconnected polygons (e.g., triangles, quadrilaterals, and other polygons). These polygons are formed by points forming lines, and lines forming faces, ultimately forming the mesh structure of the geometric model. In other words, the mesh of a geometric model is composed of vertices, edges, and faces, and these elements together constitute the geometric outline of the geometric model.
[0037] A mesh file is a data file used to describe the mesh structure of a geometric model. A mesh file can include information describing the mesh structure of the geometric model, such as vertex information and texture coordinates. Vertex information includes the coordinates of each vertex in the mesh structure that makes up the geometric model; texture coordinates are used to define the mapping relationship between vertices in the mesh structure of the geometric model and pixels in texture space. Based on texture coordinates, a two-dimensional texture map can be mapped onto the surface of the geometric model.
[0038] II. Texture generation techniques:
[0039] Texture generation technology refers to the technique of creating or synthesizing texture maps of geometric models using computational methods. Texture generation technology can be applied in fields such as computer graphics, game development, virtual reality, and architectural visualization. A texture map is a two-dimensional image containing the visual features of the surface of a geometric model. These visual features may include color, pattern, material, detail, and structure. Texture maps can be mapped onto the surface of a geometric model to simulate its material and appearance. The uses of texture maps are shown in Figure 1. Taking a three-dimensional model as an example, a texture map 102 generated for geometric model 101 can be mapped onto the surface of geometric model 101, giving the surface of geometric model 101 visual features.
[0040] Based on the above introduction of geometric models and texture generation techniques, this application provides a texture generation method. This texture generation method can utilize a large visual model, combined with guiding information (the texture information in the guiding information can be used to indicate a preset texture style), the depth map of the geometric model in at least one viewpoint, and the direction of each viewpoint, to guide the large visual model to generate a high-quality texture map for the geometric model. The high-quality texture map generated for the geometric model can refer to a texture map with rich texture details that, after being mapped onto the surface of the geometric model, makes the texture of the geometric model consistent across multiple viewpoints.
[0041] Among them, visual big models are a class of deep learning-based models designed to generate visual content, which may include images, videos or other visual data. These models typically learn from large amounts of labeled or unlabeled visual data to capture the latent features and structures in the data, thereby generating new and creative visual content. For example, visual big models may be SD (Stable Diffusion) models (a type of diffusion model) or SD XL (Stable Diffusion XL) models (another type of diffusion model), etc.
[0042] The following describes the application scenarios of the texture generation method provided in the embodiments of this application:
[0043] Step 1: Upload the model.
[0044] As shown in Figure 2a, a texture request object (referring to an object requesting the generation of a texture map) can submit a geometric model in the UI (User Interface) provided by the texture generation client. The geometric model can be uploaded locally by the texture request object, or it can be a ready-made model provided by the model library in the texture generation client. This embodiment uses a 3D dinosaur model submitted by the texture request object as an example. The texture generation client refers to a client that provides texture generation capabilities.
[0045] After submitting the geometric model, this step can also automatically correct the frontal view of the geometric model for subsequent texture generation model initialization. Automatic frontal view correction refers to the process of automatically correcting the frontal view of the geometric model, aligning the front of the geometric model with the front of the model in its 3D space, so that the front of the geometric model can be observed from a frontal perspective. For example, if the geometric model is an animal model, before automatic frontal view correction, the frontal view shows the back of the animal model; after automatic correction, the frontal view shows the front of the animal model, including its face.
[0046] Step 2: Enter the guidance information.
[0047] This application's embodiments support text-based texture generation, image-based texture generation, and voice-based texture generation tasks. A text-based texture generation task refers to the process of using text (which can be called texture generation guidance text) to guide a large visual model in generating texture maps; in this task, the texture style is described through text. An image-based texture generation task refers to the process of using an image (which can be called texture generation guidance image) to guide a large visual model in generating texture maps; in this task, the texture style is described through image. A voice-based texture generation task refers to the process of using voice (which can be called texture generation guidance voice) to guide a large visual model in generating texture maps; in this task, the texture style is described through voice. Based on this, the texture requesting object can select any one of these three tasks—text-based texture generation, image-based texture generation, or voice-based texture generation—from the interface provided by the texture generation client.
[0048] For the case of selecting a text-based texture task, as shown in Figure 2b, the texture requester can input texture generation guidance text (for example, the guidance text is the prompt "a bright green dinosaur with scaly skin" in Figure 2b). The guidance text indicates the texture style to be generated, and the texture generation guidance image obtained by performing text-to-image conversion on the guidance text can be used as guidance information. Furthermore, this embodiment does not limit the language type of the texture generation guidance text. For example, Chinese texture generation guidance text, or English texture generation guidance text, or texture generation guidance text in other languages can be used. That is to say, this embodiment can support text-based texture tasks in different language types. In addition, the texture generation client can also provide the function of expanding the texture generation guidance text (for example, the "help me polish" function in Figure 2b). The so-called text expansion refers to expanding the texture generation guidance text to a texture generation guidance text with richer indicated texture styles. For example, the texture generation guidance text before expansion only indicates the texture color, while the expanded texture generation guidance text also indicates the texture material. The texture styles indicated by the expanded texture generation guidance text are richer. Through the text expansion function, the expanded texture generation guidance text can provide more accurate and richer texture generation guidance during the texture generation process, which can improve the generation quality of texture maps.
[0049] When selecting the image-generated texture task, the texture request object can input a texture generation guide image, which indicates the texture style to be generated and can be used as guide information.
[0050] When selecting the voice-to-texture task, the texture requester can input guiding voice for texture generation. This guiding voice indicates the requested texture style. Speech recognition processing can be performed on the guiding voice to obtain guiding text. The resulting guiding image is then obtained by performing text-to-image conversion on the guiding text and used as the guiding information. Furthermore, this embodiment does not limit the language type of the guiding voice. For example, it can use Chinese, English, or other languages.
[0051] Step 3: Results verification.
[0052] After submitting the geometric model and inputting guidance information, the texture generation client's backend can guide the visual large model to generate texture maps for the geometric model based on the guidance information. The texture generation client can map the texture maps onto the surface of the geometric model, as shown in Figure 2c. The texture generation client can present the geometric model after mapping the texture maps, and it can also support viewing the texture generation effect of the geometric model from a free perspective.
[0053] Based on the above description of the application scenarios of the texture generation method, it can be seen that the embodiments of this application can support text-generated texture tasks, image-generated texture tasks, and speech-generated texture tasks. It can generate high-quality texture maps for geometric models based on input texture generation guide text, generate high-quality texture maps for geometric models based on input texture generation guide images, and generate high-quality texture maps for geometric models based on input texture generation guide speech.
[0054] The texture generation method proposed in this application can be executed by a computer device, such as a terminal device, or a server, or a texture generation system composed of a terminal device and a server.
[0055] When the computer device is the terminal device, both the texture generation client and its backend are deployed on the terminal device. The backend of the texture generation client contains a large visual model. On the terminal device, the texture generation client can receive the geometric model and input guidance information submitted by the texture request object. The texture generation client can send the mesh file of the geometric model and the guidance information to its backend. The backend of the texture generation client can use the guidance information to instruct the large visual model to generate texture maps for the geometric model and then send the generated texture maps to the texture generation client. The texture generation client can then map the texture maps onto the surface of the geometric model, presenting the texture generation effect of the geometric model.
[0056] When the computer device is a server, a large visual model is deployed on the server. The server can obtain the mesh file and guidance information of the geometric model (for example, it can obtain the mesh file and guidance information of the geometric model from memory), and guide the large visual model to generate texture maps for the geometric model using the guidance information.
[0057] When the computer device is a texture generation system composed of a terminal device and a server, as shown in Figure 3, the texture generation system may include a terminal device 301 and a server 302. The terminal device 301 and the server 302 can establish a direct communication connection through wired communication, or the terminal device 301 and the server 302 can establish an indirect communication connection through wireless communication. This application embodiment does not limit this.
[0058] In the texture generation system, the texture generation client is installed and runs on terminal device 301, while the background of the texture generation client is deployed on server 302. A large visual model is deployed in the background of the texture generation client. In the texture generation system, the texture generation client can receive the geometric model and input guidance information submitted by the texture request object. The texture generation client can send the mesh file of the geometric model and the guidance information to server 302 via terminal device 301. The background of the texture generation client on server 302 can guide the large visual model to generate texture maps for the geometric model using the guidance information, and then send the generated texture maps to terminal device 301. The texture generation client on terminal device 301 can map the texture maps onto the surface of the geometric model, presenting the texture generation effect of the geometric model. Alternatively, the background of the texture generation client on server 302 can guide the large visual model to generate texture maps for the geometric model using the guidance information, and then map the texture maps onto the surface of the geometric model. The server can then send the mesh file of the geometric model after mapping the texture maps to terminal device 301, and terminal device 301 can present the texture generation effect of the geometric model.
[0059] The texture generation system shown in Figure 3 is intended to more clearly illustrate the technical solutions of the embodiments of this application and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0060] The client mentioned in this application refers to a program that provides services locally on a terminal device, and may include, but is not limited to, any of the following: application client, mini-program client, and web (World Wide Web) program client. The terminal device mentioned in this application may include, but is not limited to, any of the following: smartphone, tablet computer, laptop computer, desktop computer, smartwatch, smart home appliance, smart vehicle terminal, and aircraft. The server mentioned in this application may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server, etc.
[0061] This application provides a texture generation method, which includes: a method for obtaining guidance information and a method for generating texture maps guided by the guidance information. As shown in Figure 4, the texture generation method may include, but is not limited to, the following steps S401-S403:
[0062] S401, Obtain the guidance information and mesh file of the geometric model. The guidance information contains texture information, which is used to indicate the preset texture style.
[0063] In step S401, the guiding information can refer to visual information that guides the texture generation of the geometric model; visual information refers to information that can be perceived visually, for example, visual information can be an image, that is, the guiding information can refer to an image that guides the texture generation of the geometric model. The guiding information may include texture information, which can refer to information representing texture in the guiding information, specifically the texture pixel values in the image that serves as the guiding information; for example, for an image containing "mystical mage", the texture information is the pixel values in the image that represent the texture of "mystical mage".
[0064] The texture information in the guidance information can be used to indicate a preset texture style, and it can guide the texture generation of the geometric model. Specifically, the texture style guidance is that the texture information in the guidance information is equivalent to providing a texture style template, guiding the generation of texture maps for the geometric model according to this template, so that the actual texture style in the generated texture map is consistent with the preset texture style indicated by the texture information.
[0065] The guidance information can be a directly obtained image. It can also be an image obtained by converting text-based prompts (e.g., texture-generated guidance text), speech-based prompts (e.g., texture-generated guidance speech), or other help-type prompts. Specifically, the methods for obtaining guidance information can include any of the following:
[0066] The first method involves obtaining guidance information through text-to-image conversion. Specifically, this involves acquiring the texture generation guidance text from the geometric model, performing text-to-image conversion on the text to obtain a texture generation guidance image, and then using this image as the guidance information. The texture generation guidance text can be input by the texture request object through the texture generation client. Texture generation guidance text refers to the text that guides texture generation; that is, text guides texture generation. Texture generation guidance image refers to the image that guides texture generation; that is, an image guides texture generation.
[0067] In this method of obtaining guidance information, text-to-image conversion refers to text processing that converts text into an image. The image content in the texture-generated guidance image obtained after text conversion is the content expressed by the texture-generated guidance text. Text-to-image conversion can be performed using an existing Text2Image Model, a large-scale model based on deep learning. This model parses and understands the input text to generate a corresponding image. For example, as shown in Figure 5, if the texture-generated guidance text is "Mysterious Mage," after performing text-to-image conversion, a texture-generated guidance image containing "Mysterious Mage" can be obtained. It can be seen that the first method of obtaining guidance information is applicable to text-to-texture tasks.
[0068] Optionally, the texture generation guide image can be selected from multiple candidate images obtained by performing text-to-image conversion on the texture generation guide text. Specifically, the texture generation guide text can be obtained, and text-to-image conversion can be performed on the texture generation guide text to obtain multiple candidate images. In response to an image selection operation, the target image selected from the multiple candidate images can be used as the texture generation guide image, and the texture generation guide image can be used as the guide information. The image content of the multiple candidate images is different. For example, if the texture generation guide text is "mysterious mage," and text-to-image conversion yields three candidate images, the clothes worn by the mage in the three candidate images are different.
[0069] In this method of obtaining guidance information, multiple candidate images can be output to the texture request object. The texture request object can perform an image selection operation to select a texture generation guidance image from the multiple candidate images as guidance information. This allows the selected texture generation guidance image to more accurately express the content of the texture generation guidance text, improves the guidance accuracy of the guidance information in the texture generation process, and the texture style in the generated texture map can more accurately match the texture style required by the texture request object.
[0070] The second method involves obtaining guidance information through images. Specifically, a texture generation guidance image of the geometric model can be obtained and identified as the guidance information; this guidance image can be input by the texture request object through the texture generation client. It can be seen that this second method of obtaining guidance information is applicable to image-to-texture tasks.
[0071] The third method involves obtaining guidance information through speech-to-image conversion. Specifically, this involves acquiring the texture generation guidance speech for the geometric model, performing speech recognition processing on the guidance speech to obtain texture generation guidance text, performing text-to-image conversion on the guidance text to obtain texture generation guidance image, and then using the texture generation guidance image as the guidance information. The texture generation guidance speech can be input by the texture request object through the texture generation client. Texture generation guidance speech refers to the speech that guides texture generation; that is, speech guides the generation of the texture.
[0072] Alternatively, similar to the method of obtaining guidance information through text-image conversion processing, in this method of obtaining guidance information, the texture-generated guidance image can be selected from multiple candidate images obtained by text-image conversion processing of the texture-generated guidance text.
[0073] In this method of obtaining guidance information, speech recognition processing refers to converting speech into text, and text-to-image conversion processing refers to converting text into an image. After speech recognition and text-to-image conversion processing, the texture generation guidance speech is converted into a texture generation guidance image, and the image content in the texture generation guidance image is the content expressed by the texture generation guidance speech. It can be seen that the third method of obtaining guidance information can be applied to speech-to-texture tasks.
[0074] As can be seen from the above description of the three methods of obtaining guidance information, the embodiments of this application use images as guidance information. Images can express more precise content than text or voice. The texture style indicated by images is more precise than the texture style indicated by text or voice. The texture style has richer details, thereby guiding the generation of texture maps with rich texture style details and improving the generation quality of texture maps.
[0075] In step S401, the mesh file of the geometric model may include vertex information and texture coordinates. The vertex information may include the vertex coordinates of each vertex in the mesh structure that makes up the geometric model; the texture coordinates can be used to define the mapping relationship between vertices in the mesh structure of the geometric model and pixels in texture space.
[0076] S402, based on the mesh file of the geometric model, use at least one depth rendering camera to perform depth rendering processing on the geometric model from at least one viewpoint to obtain the depth map of the geometric model in each viewpoint; at least one depth rendering camera corresponds to one viewpoint respectively.
[0077] In step S402, the depth rendering camera refers to the camera used for depth rendering processing. It is a virtual camera, essentially an observation point. In other words, the depth rendering camera refers to the observation point used for depth rendering processing. One observation point corresponds to one viewpoint, and different observation points correspond to different viewpoints. One viewpoint corresponds to one direction, and different viewpoints correspond to different directions. The viewpoint corresponding to an observation point defines the observation range of that observation point in the corresponding direction. At least one depth rendering camera (i.e., at least one observation point) corresponds to one viewpoint.
[0078] Depth rendering for any given viewpoint refers to the process of rendering a depth map of a geometric model from that viewpoint. The process can be simply understood as follows, as shown in Figure 6: At least one depth rendering camera is pre-set around the geometric model in the 3D space (e.g., 2, 4, 6, or 8 depth rendering cameras). Each depth rendering camera corresponds to a viewpoint, and the viewpoints corresponding to each depth rendering camera are different. Using the pre-set depth rendering camera surrounding the geometric model, depth maps of the geometric model are rendered from their respective viewpoints.
[0079] The depth of a geometric model at any given viewpoint refers to the distance between the vertices observable from that viewpoint and the corresponding observation point. The observable vertices are those within the mesh structure of the geometric model that fall within the observation range of that viewpoint. Rendering the depth map of a geometric model at any given viewpoint can involve using different grayscale values to represent the different distances between the observable vertices and the corresponding observation point. The closer the distance, the higher the grayscale value, resulting in a brighter, whiter image in the depth map; conversely, the farther the distance, the lower the grayscale value, resulting in a darker, blacker image. In other words, a depth map is a grayscale image that uses different grayscale values to represent the different distances between the observed vertices and the observation point in the geometric model. The depth map of a geometric model at any viewpoint can represent the distance between the vertex that the geometric model can observe at that viewpoint and the observation point corresponding to that viewpoint. The distance between the vertex and the observation point reflects the geometric contour (or geometric shape) of the geometric model. Therefore, the depth map of a geometric model at any viewpoint can be used to characterize the geometric contour of the geometric model at that viewpoint.
[0080] The depth map of the geometric model at each viewpoint can be used to indicate the geometric information of the geometric model at each viewpoint. Geometric information refers to the geometric contour of the geometric model at each viewpoint. Therefore, the geometric information of the geometric model at each viewpoint can be used to characterize the geometric contour of the geometric model at each viewpoint. The geometric information of the geometric model at each viewpoint can include the vertices that the geometric model can observe at each viewpoint and the distances to the observation points at each viewpoint.
[0081] In step S402, the geometric model undergoing depth rendering is the geometric model after frontal view correction processing. As described above, frontal view correction processing refers to the process of aligning the front of the geometric model with the front of the three-dimensional space in which the geometric model resides, so that the front of the geometric model can be observed from a frontal viewpoint. Frontal view correction processing can be obtained by performing coordinate transformation on the vertex coordinates of each vertex in the mesh file of the geometric model.
[0082] Frontal view correction is particularly important for geometric models containing prominent faces (e.g., people or animals) or those with significant differences between their faces (e.g., the doors, windows, and walls of a house are significantly different, or the screen and non-screen surfaces of a phone model are significantly different). For example, the front side of a geometric model containing a face is its front view. The goal of frontal view correction is to align this front side with the frontal view. Without correction, the face might be observed from a perspective where it shouldn't be (e.g., a rear view), resulting in facial textures being rendered on the back of the model. Texture generation, on the other hand, aims to render facial textures on the front of the model. Similarly, for a phone model, the screen is its front. Frontal view correction aims to align the screen with the frontal view. Without correction, the screen might be observed from a perspective where it shouldn't be (e.g., a rear view), resulting in screen textures being rendered on the back of the model. Texture generation, on the other hand, aims to render screen textures on the front of the model.
[0083] In other words, failing to perform frontal view correction on the geometric model may lead to the generation of incorrect textures; for example, the multi-face problem, where frontal face textures are rendered for both the front and back of a human or animal model; or the multi-screen problem, where screen textures are rendered for both the front and back of a mobile phone model; therefore, performing frontal view correction on the geometric model can improve the accuracy of texture generation.
[0084] S403, based on the texture information in the guidance information and the depth map of the geometric model in each view, perform texture rendering processing on the geometric model to obtain the texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
[0085] In step S403, texture rendering refers to the process of generating texture maps for the geometric model. Based on the texture information in the guidance information and the depth maps of the geometric model at each viewpoint, texture rendering can be performed on the geometric model to obtain its texture map. The texture information in the guidance information and the depth maps of the geometric model at each viewpoint can play a guiding role in the texture map generation process. This guiding role can include: the texture information in the guidance information can guide the actual texture style in the generated texture map to be consistent with the preset texture style indicated by the texture information in the guidance information; and the depth maps of the geometric model at each viewpoint can guide the generated texture map to better fit the geometric contours of the geometric model. Therefore, based on the guiding role of the texture information in the guidance information and the guiding role of the depth maps, the quality of texture map generation can be improved.
[0086] Optionally, the directional information of each viewpoint can be used to guide the texture mapping generation process. Specifically, the directional information of each viewpoint in at least one viewpoint can be obtained, and the directional information of each viewpoint can be used to indicate the direction of each viewpoint. Based on the guiding information, the depth map of the geometric model in each viewpoint, and the directional information of each viewpoint, the geometric model can be texture-rendered to obtain the texture map of the geometric model. The guiding role of directional information in the texture mapping generation process can include: the direction of each viewpoint can take into account the natural flow of the texture during the texture mapping generation process, thereby solving the visual inconsistency problem caused by insufficient directionality. The visual inconsistency problem refers to the mismatch between the texture map of the geometric model and the geometric contour of the geometric model, resulting in visual misalignment or distortion.
[0087] The texture rendering process can be performed N times, where N is a positive integer. A single texture rendering process can include texture color rendering and texture map conversion. Texture color rendering refers to rendering the texture color of the geometric model from each viewpoint, resulting in a color view of the geometric model at each viewpoint. Rendering the texture color of the geometric model means expressing the texture style through color rendering; the color view of the geometric model at each viewpoint can include the texture generated for the geometric model at each viewpoint. Texture map conversion is the process of fusing the color views of the geometric model from various viewpoints into a texture map of the geometric model. This fusing eliminates differences in texture styles between different viewpoints, thereby improving texture generation quality. For example, different viewpoints may observe the same part of the geometric model (e.g., a hat), and the texture style of the same part of the geometric model may differ from different viewpoints (e.g., the color of the hat is inconsistent from different viewpoints). This difference can be eliminated through fusing, making the texture style of the same part of the geometric model consistent across different viewpoints (e.g., the color of the hat is consistent across different viewpoints).
[0088] The process of the i-th texture rendering process in N texture rendering processes can include: when i=1, texture color rendering processing of the geometric model can be performed based on the texture information in the guidance information, the depth map of the geometric model in each view, and the orientation information of each view to obtain the color view of the geometric model in each view; texture map conversion processing is performed on the color view of the geometric model in each view to obtain the shared texture map corresponding to the i-th texture rendering process. When i∈[2,N], texture color rendering processing of the geometric model can be performed based on the texture information in the guidance information, the depth map of the geometric model in each view, the orientation information of each view, and the shared texture map corresponding to the (i-1)-th texture rendering process to obtain the color view of the geometric model in each view; texture map conversion processing is performed on the color view of the geometric model in each view to obtain the shared texture map corresponding to the i-th texture rendering process. The texture map of the geometric model can be generated based on the shared texture map corresponding to the N-th texture rendering process.
[0089] As can be seen, the N texture rendering processes are iterative. The first texture rendering process yields a shared texture map corresponding to the first texture rendering process. For the second and subsequent texture rendering processes, the current texture rendering process can be executed based on the shared texture map corresponding to the previous texture rendering process. The shared texture map has undergone texture map transformation processing, eliminating the differences in texture styles between geometric models from different perspectives. Therefore, using the shared texture map corresponding to the previous texture rendering process to participate in the current texture rendering process helps to ensure the consistency of texture rendering across multiple perspectives, thereby improving the texture rendering effect and the quality of texture generation.
[0090] The following section describes how to set the number of texture rendering processes, N:
[0091] Setting Method 1: The number of texture rendering processes, N, can be preset. For example, the number of texture rendering processes, N, can be set based on empirical values. These empirical values are values summarized from multiple texture generation experiences. Setting the number of texture rendering processes, N, based on these empirical values can ensure that a high-quality texture map is generated after an appropriate number of iterations of texture rendering.
[0092] Setting Method 2: The number of texture rendering processes, N, can be controlled based on the similarity between the texture style in the shared texture map and the preset texture style indicated by the texture information in the guidance information. Specifically, after each texture rendering process, a similarity calculation can be performed between the texture style in the shared texture map obtained from the texture rendering process and the preset texture style indicated by the texture information in the guidance information. If the calculated similarity score is greater than or equal to a similarity threshold, the iteration can stop. If the calculated similarity score is less than the similarity threshold, the iteration can continue, and subsequent texture rendering processes can be performed until the calculated similarity score is greater than or equal to the similarity threshold. The method for calculating the similarity between the texture style in the shared texture map obtained from the texture rendering process and the preset texture style indicated by the guidance information can include: extracting a first feature vector from the shared texture map obtained from the texture rendering process (the first feature vector refers to the feature vector used to represent the texture style in the shared texture map), extracting a second feature vector from the texture information in the guidance information (the second feature vector refers to the feature vector used to represent the preset texture style indicated by the texture information in the guidance information), and performing vector similarity calculation on the first feature vector and the second feature vector to obtain a similarity score. Therefore, the consistency between the actual texture style in the texture map and the preset texture style indicated by the texture information in the guidance information can mean that the similarity score between the actual texture style in the texture map and the preset texture style indicated by the texture information in the guidance information is greater than the similarity threshold.
[0093] In this embodiment, the texture information in the guidance information can guide the actual texture style generated during texture generation to be consistent with the preset texture style indicated by the texture information in the guidance information; the depth map of the geometric model at each viewpoint can guide the texture style generated during texture generation to adapt to the geometric contour of the geometric model; the direction of each viewpoint can increase the texture directionality during texture generation, avoiding the generation of visually inconsistent textures; thus, based on the guidance of the texture information in the guidance information, the depth map of the geometric model at each viewpoint, and the guidance of the direction of each viewpoint, the texture generation quality can be improved. Furthermore, since the guidance information is an image, the image can indicate a more accurate texture style, which helps to optimize texture details during texture generation, prevents the texture generation result from being too smooth, and preserves the richness and realism of the texture style in the image, thereby improving the texture generation quality.
[0094] This application provides a texture generation method, which includes: a texture color rendering process, a texture map conversion process, and a texture map multi-view splitting process. As shown in Figure 7, the texture generation method may include, but is not limited to, the following steps S701-S707:
[0095] S701, obtain the guidance information and mesh file of the geometric model. The guidance information contains texture information, which is used to indicate the preset texture style.
[0096] The execution process of step S701 in this embodiment is the same as the execution process of step S401 in the embodiment shown in Figure 4 above. For details of the execution process of step S701, please refer to the execution process described in step S401 in the embodiment shown in Figure 4 above, which will not be repeated here.
[0097] S702, based on the mesh file of the geometric model, use at least one depth rendering camera to perform depth rendering processing on the geometric model from at least one viewpoint to obtain the depth map of the geometric model in each viewpoint; at least one depth rendering camera corresponds to one viewpoint respectively.
[0098] The execution process of step S702 in this embodiment is the same as the execution process of step S402 in the embodiment shown in Figure 4 above. For details of the execution process of step S702, please refer to the execution process described in step S402 in the embodiment shown in Figure 4 above, which will not be repeated here.
[0099] S703, Obtain the direction information of each view in at least one view, the direction information of each view is used to indicate the direction of each view.
[0100] The execution process of step S703 in this embodiment is the same as the process of obtaining the direction information of each view in at least one view in step S403 of the embodiment shown in FIG4 above. The specific execution process of step S703 can be found in the process of obtaining the direction information of each view in at least one view described in step S403 of the embodiment shown in FIG4 above, which will not be repeated here.
[0101] S704, when i=1, based on the texture information in the guidance information, the depth map of the geometric model in each view, and the direction information of each view, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model in each view.
[0102] In step S704, the texture color rendering process is performed separately on a viewpoint basis, and any one of the at least one viewpoint can be represented as the target viewpoint. For the texture color rendering process in the first (i.e., i=1) texture rendering process, the process of performing texture color rendering on the geometric model under the target viewpoint may include: performing feature extraction processing on the guiding information to obtain the texture features corresponding to the texture information in the guiding information; performing feature extraction processing on the depth map under the target viewpoint to obtain the geometric features corresponding to the target viewpoint; performing feature extraction processing on the direction information of the target viewpoint to obtain the direction features corresponding to the target viewpoint; and performing feature integration processing on the texture features, the geometric features corresponding to the target viewpoint, and the direction features corresponding to the target viewpoint to obtain the color view of the geometric model under the target viewpoint.
[0103] In the process of texture and color rendering under the first target perspective mentioned above:
[0104] ① The feature extraction process for the guiding information can be called the first feature extraction process. The essence of the guiding information is an image, and the first feature extraction process refers to image encoding. This first feature extraction process can be performed by an image encoder. The image encoder can be used to extract useful information (such as texture information) from the input image and convert it into a feature representation that a large visual model can understand. This feature representation can be called texture features. Therefore, the first feature extraction process can be used to encode the texture information in the guiding information into texture features. Texture features refer to feature vectors used to characterize a preset texture pattern.
[0105] ② The feature extraction process on the depth map can be called the second feature extraction process, which refers to geometric information extraction. Geometric information extraction is used to extract geometric information from the depth map and convert it into a feature representation that the large visual model can understand. This feature representation can be called geometric features. Geometric information extraction can be performed through a ControlNet, which is a type of neural network. Based on this, the geometric features corresponding to the target viewpoint refer to feature vectors that characterize the geometric information of the geometric model under the target viewpoint. The geometric information of the geometric model under the target viewpoint can reflect several contours of the geometric model under the target viewpoint. Therefore, the geometric features corresponding to the target viewpoint refer to feature vectors used to characterize the geometric contours of the geometric model under the target viewpoint.
[0106] ③ The feature extraction process for directional information can be called the third feature extraction process. The essence of directional information is text, and the third feature extraction process refers to text encoding. This third feature extraction process can be performed through a text encoder. The text encoder can be used to extract useful information (such as direction) from the input text and convert it into a feature representation that the large visual model can understand. This feature representation can be called directional features. Therefore, the third feature extraction process can be used to encode the directional information of the target viewpoint into directional features corresponding to the target viewpoint. The directional features corresponding to the target viewpoint are feature vectors used to characterize the direction of the target viewpoint.
[0107] ④ Feature integration processing can be used to integrate texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint into the denoising process of the large visual model. The large visual model can predict the noise that should be removed from the noisy image based on the texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint. After removing the predicted noise from the noisy image, the color view (Latent Views) of the geometric model under the target viewpoint can be obtained. Feature integration processing can be performed by a feature integration network (such as Unet), which is a convolutional neural network model.
[0108] The process of integrating texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint to obtain a color view of the geometric model from the target viewpoint can include: performing noise prediction processing on the noisy image based on texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint to obtain predicted noise; and performing denoising processing on the noisy image based on the predicted noise to obtain a noisy image of the geometric model from the target viewpoint. Texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint can be integrated into the denoising process of the large visual model through a decoupled cross-attention mechanism. This decoupled cross-attention mechanism can effectively handle the interaction between different features (i.e., texture features, geometric features, and directional features).
[0109] S705, when i∈[2,N], based on the texture information in the guidance information, the depth map of the geometric model in each view, the direction information of each view, and the shared texture map corresponding to the i-1th texture rendering process, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model in each view.
[0110] In step S705, the texture color rendering process is performed on a per-viewpoint basis, and any one of the at least one viewpoint can be represented as the target viewpoint. For subsequent texture color rendering processes (i.e., i∈[2,N]) other than the first one, the process of performing texture color rendering on the geometric model under the target viewpoint may include the following sub-steps s11-s12:
[0111] s11 performs view splitting on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model from each perspective.
[0112] View splitting refers to the rasterization process, which splits a shared texture map into views of the geometric model from each viewpoint. The resulting views are called feedback views. The view splitting process can include: based on the mapping relationship between pixels in the shared texture map corresponding to the (i-1)th texture rendering process and vertices of the geometric model, using the pixel values of the pixels in the shared texture map corresponding to the (i-1)th texture rendering process as the pixel values of the vertices mapped onto the geometric model; determining the vertices observed on the geometric model from each viewpoint; and performing view rendering on the geometric model from each viewpoint based on the pixel values of the vertices observed from each viewpoint to obtain the feedback view of the geometric model from each viewpoint. The view rendering is performed separately for each viewpoint. For any given viewpoint, the view rendering for that viewpoint can be used to render the pixel values of the vertices observed on the geometric model from that viewpoint onto an image, obtaining the feedback view of the geometric model from that viewpoint. As can be seen, based on the mapping relationship between pixels in the shared texture map and vertices of the geometric model, the shared texture map can be quickly mapped onto the geometric model. By performing view rendering processing on the geometric model mapped with the shared texture map from different perspectives, feedback views from different perspectives can be obtained. In other words, the view splitting process realizes the rapid conversion from texture map to multiple views of the geometric model mapped with texture in the texture map.
[0113] s12, based on the texture information in the guidance information, the depth map of the geometric model in each view, the direction information of each view, and the feedback view in each view, performs texture color rendering processing on the geometric model to obtain the color view of the geometric model in each view.
[0114] For the target viewpoint, after introducing the feedback view of the geometric model in the target viewpoint, the process of performing texture and color rendering on the geometric model in the target viewpoint can include: obtaining the texture features corresponding to the texture information in the guidance information, where the texture features are obtained by performing feature extraction processing on the guidance information (the feature extraction processing here corresponds to the first feature extraction processing mentioned above); obtaining the geometric features corresponding to the target viewpoint, where the geometric features corresponding to the target viewpoint are obtained by performing feature extraction processing on the depth map in the target viewpoint (the feature extraction processing here corresponds to the second feature extraction processing mentioned above); obtaining the directional features corresponding to the target viewpoint, where the directional features corresponding to the target viewpoint are obtained by performing feature extraction processing on the directional information in the target viewpoint (the feature extraction processing here corresponds to the third feature extraction processing mentioned above); performing feature extraction processing on the feedback view in the target viewpoint to obtain the feedback features corresponding to the target viewpoint; and performing feature integration processing on the texture features, the geometric features corresponding to the target viewpoint, the directional features corresponding to the target viewpoint, and the feedback features corresponding to the target viewpoint to obtain the color view of the geometric model in the target viewpoint.
[0115] In the process of texture and color rendering from the subsequent secondary target perspective described above:
[0116] ①The first feature extraction process, the second feature extraction process, and the third feature extraction process in the subsequent texture and color rendering process are the same as those in the first texture and color rendering process, and will not be described again here.
[0117] ② In subsequent texture and color rendering processes, the feature extraction process for the feedback view can be called the fourth feature extraction process. The feedback view is essentially an image, and the fourth feature extraction process refers to image encoding. This fourth feature extraction process can be performed through an image encoder. The image encoder can extract useful information from the input image (such as texture patterns) and convert it into a feature representation that the large visual model can understand. This feature representation can be called the feedback feature. Therefore, the fourth feature extraction process can be used to encode the feedback view into feedback features. For any given viewpoint, the feedback feature corresponding to that viewpoint refers to the feature vector that characterizes the texture pattern of the geometric model in the feedback view at that viewpoint.
[0118] ③ Feature integration processing can be used to integrate texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint into the denoising process of the large visual model. The large visual model can predict the noise that should be removed from the noisy image based on the texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint. After removing the predicted noise from the noisy image, a color view of the geometric model under the target viewpoint can be obtained. Feature integration processing can be performed through a feature integration network (e.g., Unet), which is a convolutional neural network model.
[0119] The process of integrating texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint to obtain a color view of the geometric model from the target viewpoint can include: performing noise prediction processing on the noisy image based on texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint to obtain predicted noise; and performing denoising processing on the noisy image based on the predicted noise to obtain a noisy image of the geometric model from the target viewpoint. Texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint can be integrated into the denoising process of the large visual model through a decoupled cross-attention mechanism. This decoupled cross-attention mechanism can effectively handle the interaction between different features (i.e., texture features, geometric features, directional features, and feedback features).
[0120] Based on the above sub-steps s11-s12, it can be seen that when i∈[2,N], the shared texture map corresponding to the (i-1)th texture rendering process is shared by the texture color rendering processes corresponding to each viewpoint in the i-th texture rendering process. The sharing method is: the shared texture map corresponding to the (i-1)th texture rendering process is subjected to view splitting processing to obtain the feedback view of the geometric model in each viewpoint in at least one viewpoint. The feedback view of the geometric model in each viewpoint is used to participate in the texture color rendering process corresponding to each viewpoint in the i-th texture rendering process. The benefit of sharing is that the shared texture map has undergone texture map conversion processing, eliminating the difference between the texture styles of the geometric model in different viewpoints. Therefore, using the shared texture map corresponding to the (i-1)th texture rendering process to participate in the i-th texture rendering process helps to improve the consistency of multi-view texture rendering, which can improve the texture rendering effect and the texture generation quality.
[0121] S706 performs texture mapping conversion on the color view of the geometric model from various perspectives to obtain the shared texture map corresponding to the i-th texture rendering process.
[0122] In step S706, texture mapping conversion processing refers to the process of texture baking. Texture mapping conversion processing can be used to convert the texture information of the geometric model in the color views from various perspectives into a shared texture map. The texture mapping conversion process may include: mapping the color view of the geometric model in each perspective onto a texture template of the geometric model to obtain a mapped texture template; and performing texture fusion processing on the different textures mapped to the color views from different perspectives in the mapped texture template to obtain the shared texture map corresponding to the i-th texture rendering process.
[0123] In the above texture mapping conversion process:
[0124] ① The texture template of a geometric model refers to a blank texture map of the geometric model, which does not contain any texture. Mapping specifically refers to mapping the texture of the geometric model in the color view of each perspective onto the texture template of the geometric model. Mapping can be based on the mapping relationship between the vertices observed on the geometric model from each perspective and the pixels in the texture template. The mapping process for any perspective may include: based on the mapping relationship between the vertices observed on the geometric model from that perspective and the pixels in the texture template, determining the pixel value of the vertex observed on the geometric model from that perspective in the corresponding color view as the pixel value of the mapped pixel in the texture template.
[0125] ② Texture fusion processing refers to the process of fusing different textures mapped to a texture template from different viewpoints in a color view. Specifically, for the same vertex observed on the geometric model from different viewpoints, the pixel values in the color view may differ. Based on the mapping relationship between vertices and pixels in the texture template, mapping the pixel values of that vertex in the color view from different viewpoints to the texture template results in pixels with different pixel values. Different textures refer to the same pixel having different pixel values in the mapped texture template. The texture fusion process can include fusing different textures; that is, performing pixel value fusion processing on the same pixel with different pixel values. Pixel value fusion processing can be, for example, averaging different pixel values of the same pixel to eliminate the differences between the different pixel values. It can be seen that texture fusion processing can be used to eliminate the differences between different textures mapped to a texture template from different viewpoints in a color view, thus ensuring the consistency of the texture of the geometric model from different viewpoints and improving the quality of texture generation.
[0126] S707, Generate a texture map of the geometric model based on the shared texture map corresponding to the Nth texture rendering process. The actual texture style of the texture map is consistent with the preset texture style indicated by the texture information in the guidance information.
[0127] In step S707, the Nth texture rendering process is performed in the encoding space. Therefore, the shared texture map corresponding to the Nth texture rendering process needs to be converted from the encoding space to the non-encoding space before generating the texture map of the geometric model. The encoding space, also known as the latent space, refers to the dataset containing low-dimensional data, while the non-encoding space refers to the dataset containing high-dimensional data. Converting data from the non-encoding space to the encoding space is achieved by encoding the data, such as text encoding for text or image encoding for images. Converting data from the encoding space to the non-encoding space is achieved by decoding the data, such as decoding text features or image features. High-dimensional data refers to data of the original data dimension, such as text or images, while low-dimensional data refers to the representation data of high-dimensional data, such as text features of text or image features of images. It is easy to see that low-dimensional features are processed in the encoding space. Low-dimensional features are easier for the model to understand and utilize. Processing low-dimensional data can speed up the processing speed. Therefore, performing N texture rendering processes in the encoding space can improve the speed of texture rendering and thus improve texture generation efficiency.
[0128] The process of generating a texture map of the geometric model based on the shared texture map corresponding to the Nth texture rendering process can include: performing view splitting processing on the shared texture map corresponding to the Nth texture rendering process to obtain feedback views of the geometric model from each viewpoint; decoding processing on the feedback views of the geometric model from each viewpoint to obtain decoded views of the geometric model from each viewpoint; and performing texture map conversion processing on the decoded views from each viewpoint to obtain the final texture map of the geometric model. It can be seen that this embodiment does not directly decode the shared texture map corresponding to the Nth texture rendering process to obtain the texture map of the geometric model. Instead, it splits the shared texture map corresponding to the Nth texture rendering process into feedback views from each viewpoint, decodes them separately, and then performs texture map conversion processing to obtain the texture map of the geometric model. This can further eliminate texture differences from different viewpoints, thereby improving the texture generation quality.
[0129] In this embodiment, color views can be generated simultaneously from different perspectives, and texture mapping conversion is performed on the color views from different perspectives to generate the final texture map, instead of performing texture rendering across different perspectives. This achieves a fast and efficient texture generation process. Furthermore, by eliminating the differences between textures generated from different perspectives, this embodiment ensures that the generated textures maintain visual consistency across different perspectives, thereby improving the quality of texture generation.
[0130] The structure of the large-scale visual model is shown in Figure 8. The large-scale visual model can include a direction-aware adaptation module, a visual guidance enhancement module, and a texture fusion module. Specifically:
[0131] (1) Direction perception module:
[0132] The input to the orientation sensing module is the mesh file of the geometric model after frontal view correction.
[0133] The output of the orientation-aware module is the depth map of the geometric model in each viewpoint at least once, as well as the orientation information for each viewpoint.
[0134] The orientation awareness module can pre-set at least one depth rendering camera (e.g., 6 or 8 depth rendering cameras) surrounding the geometric model. Based on the viewpoint corresponding to each of the at least one depth rendering camera, it renders a corresponding depth map for each viewpoint. Furthermore, it can consider the different poses (here, pose refers to orientation) of different depth rendering cameras and extract orientation information related to each viewpoint. The depth map can be processed by a ControlNet to obtain geometric features, and the orientation information can be processed by a text encoder to obtain orientation features.
[0135] The orientation awareness module provides depth maps and orientation information of the geometric model from different perspectives. Based on the guiding role of this information in the texture generation process, it ensures the natural flow and consistency of the texture from different perspectives and avoids common visual misalignment problems.
[0136] (2) Visual guidance enhancement module:
[0137] The input to the visual guidance enhancement module is texture-generated guidance text, texture-generated guidance speech, or texture-generated guidance image.
[0138] The output of the visual guidance enhancement module is texture features.
[0139] The visual guidance enhancement module can be used to extract texture features. For input texture-generated guidance text, this module can utilize an existing Text2Image Model for pre-processing text-to-image conversion, then input the converted image into an image encoder to extract texture features. For input texture-generated guidance images, this module can directly input the generated guidance image into the image encoder to extract texture features. For input texture-generated guidance speech, this module can perform speech recognition processing on the generated guidance speech, utilize an existing Text2Image Model to perform text-to-image conversion on the text obtained from the speech recognition, and then input the converted image into an image encoder to extract texture features. The texture features output by this module can be incorporated into the inference process of the large visual model through cross-attention, ensuring the completeness of texture detail generation and the consistency of the global texture style.
[0140] (3) Texture blending module:
[0141] The input to the texture fusion module is texture features, geometric features corresponding to different viewpoints, and directional features corresponding to different viewpoints.
[0142] The output of the texture blending module is the geometric model after texture mapping.
[0143] The texture fusion module integrates texture features, directional features corresponding to different viewpoints, and geometric features corresponding to different viewpoints into UNet. Through a decoupled cross-attention mechanism, it effectively handles the interactions between different features. Simultaneously, it repeatedly performs texture fusion (texture baking, as mentioned earlier) and re-rasterization on the shared global texture (the shared texture map mentioned earlier), significantly improving the detail and realism of the generated texture. The texture fusion module can also map the generated final texture map onto the surface of the geometric model, obtaining a texture-mapped geometric model. The texture fusion module maintains texture consistency across multiple viewpoints, ensuring that the final generated texture is not only highly detailed but also visually high-fidelity.
[0144] The texture mapping generation process described above demonstrates that by comprehensively utilizing texture styles indicated in text or images, the generated textures not only possess high-quality visual effects but also maintain a high degree of consistency with the geometric shape and features of the geometric model. The large visual model, as a powerful generation tool, effectively captures complex visual features, providing a solid foundation for texture generation. The visual guidance enhancement module optimizes texture details, preventing the generated results from being overly smooth, thus preserving the richness and realism of the textures. The orientation-aware module, by introducing directional information, ensures that the natural flow of the texture is considered during generation, thereby resolving the visual inconsistency problem caused by insufficient directionality. The combination of these modules results in excellent performance in texture generation tasks, better meeting the needs of practical applications.
[0145] The texture generation effects of the texture generation methods are described below.
[0146] The texture generation scheme provided by this texture generation method (hereinafter referred to as "this scheme") is compared with the texture generation effects of several existing texture generation schemes (hereinafter referred to as "existing schemes") under text-based texture tasks, image-based texture tasks, and numerical quantization. The comparison results are as follows:
[0147] I. Comparison results of the textural texture task:
[0148] Existing texture generation schemes involved in texturing tasks may include Latent-Paint (existing scheme 1), Paint3D (existing scheme 2), TexPainter (existing scheme 3), Text2Tex (existing scheme 4), and SyncMVD (existing scheme 5).
[0149] Figure 9 shows a comparison of the texture generation results of this solution and existing solutions on the textural texture task. The texture generated by Latent-Paint is damaged due to noise and impurities, which is a result of the inherent limitations of SDS (Score Distillation Sampling). Paint3D and Text2Tex show obvious blurring at texture seams and exhibit polyhedral issues, thus failing to maintain consistency across multiple views. TexPainter and SyncMVD perform poorly in maintaining texture clarity, ultimately resulting in overly smooth and monotonous textures. Our solution addresses the oversmoothing problem through a visual guidance enhancement module and retains more high-frequency details during generation. Simultaneously, the orientation-aware module in this solution enhances the geometric alignment of side or back views, alleviating the multi-face problem. These innovations significantly improve the detail and consistency of the textures generated by this solution, exhibiting higher visual quality.
[0150] II. Comparison results of the image-generated texture task:
[0151] Existing texture generation schemes involved in the image-based texture task may include TEXTure (existing scheme 6), PGC-3D (existing scheme 7), and Paint3D (existing scheme 2).
[0152] Figure 10 shows a comparison of the texture generation performance of our proposed solution with existing solutions in the image-based texture task. In the image-based texture task, both Texture and PGC-3D struggle to retain the semantic information from image cues, resulting in low-quality and noisy generated textures. While Paint3D partially retains semantic information, it generates a large number of texture fragments. In contrast, our proposed solution implements a decoupled cross-attention strategy, starting from texture and orientation features, to achieve semantic alignment with the texture style cued by the image on the mesh of the geometric model, exhibiting a distinct visual effect.
[0153] III. Comparison results of numerical quantization:
[0154] Four numerical metrics—Fréchet Inception Distance (FID), Kernel Inception Distance (KID), CLIP Score (Contrastive Language-Image Pre-training Score), and Time—are used to quantify texture generation quality. FID is calculated by comparing the texture-generated guidance text (or image) with the color view (specifically, using the mean and covariance), and is used to characterize the difference between the texture-generated guidance text and the color view. A smaller FID value indicates better texture generation, and vice versa. KID is similar to FID, calculated by comparing the texture-generated guidance text (or image) with the color view (specifically, using a kernel method), and is used to characterize the difference between the texture-generated guidance text and the color view. A smaller KID value indicates better texture generation, and vice versa. The CLIP Score is calculated by comparing the texture generation guide text (or texture generation guide image) with the color view (specifically, using cosine similarity). It represents the texture similarity between the texture generation guide text and the color view. A higher CLIP Score indicates better texture generation, and vice versa. Time represents the texture generation time; a lower Time value indicates better texture generation, and vice versa.
[0155] The comparison of texture generation results between this scheme and existing schemes under numerical quantization is shown in Tables 1 and 2 below:
[0156] Table 1. Numerical quantization comparison results of textured surfaces.
[0157] Table 2 shows the numerical quantization comparison results of the image-generated textures.
[0158] Based on Tables 1 and 2 above, it can be seen that in the text-to-texture task, due to the use of the text-to-image conversion module, the texture-guided image obtained by text-to-image conversion of the texture generation guidance text may have semantic consistency deviations with the texture generation guidance text, resulting in a CLIP Score slightly lower than TexPainter (existing solution 3); however, our method improves texture quality and realism. For the image-to-texture task, Paint3D (existing solution 2) improves alignment by injecting image features into the refinement of the UV map (a graph used to define the coordinates of texture mapping on the surface of a geometric model), but performs poorly when processing complex UV maps due to semantic differences. In contrast, our solution achieves the highest CLIP Score by employing visual guidance in multi-view inference, ensuring semantic consistency with the input texture generation guidance image. Compared with optimization-based and repair-based methods that require sequential sampling, our solution can generate multiple views simultaneously, and rapid texture deformation further accelerates the generation process, while TexPainter (existing solution 3) requires forced differentiable rendering in each denoising step, thus taking up to 40 minutes.
[0159] Please refer to Figure 11, which is a schematic diagram of a texture generation device provided in an embodiment of this application. The texture generation device may include an acquisition unit 1101 and a processing unit 1102. The texture generation device may be disposed in the computer device provided in the embodiment of this application. The computer device may be, for example, a terminal device, or, for example, a server, or, the computer device may be a texture generation system composed of a terminal device and a server.
[0160] When a computer device is a texture generation system composed of terminal devices and servers, the texture generation devices are distributed throughout the texture generation system. The so-called distributed distribution means that some units of the texture generation device are located in the terminal devices of the texture generation system, and other units are located in the servers of the texture generation system. For example, the acquisition unit 1101 of the texture generation device is located in the terminal devices of the texture generation system, and the processing unit 1102 of the texture generation device is located in the servers of the texture generation system.
[0161] The texture generation apparatus shown in Figure 11 can be a computer program running on a computer device, which can be used to perform some or all of the steps in the method embodiments shown in Figure 4 or Figure 7. In the texture generation apparatus:
[0162] The acquisition unit 1101 is used to acquire the guidance information of the geometric model and the mesh file of the geometric model. The guidance information includes texture information, which is used to indicate the preset texture style.
[0163] The processing unit 1102 is used to perform depth rendering processing on the geometric model from at least one viewpoint using at least one depth rendering camera based on the mesh file of the geometric model, so as to obtain a depth map of the geometric model in each viewpoint; the at least one depth rendering camera corresponds to one viewpoint respectively.
[0164] The processing unit 1102 is further configured to perform texture rendering processing on the geometric model based on the texture information in the guidance information and the depth map of the geometric model under each viewpoint, to obtain a texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
[0165] In one implementation, when the acquisition unit 1101 acquires the guidance information of the geometric model, it is specifically used to perform any of the following:
[0166] Obtain the texture generation guidance text of the geometric model, perform text-to-image conversion on the texture generation guidance text to obtain the texture generation guidance image, and determine the texture generation guidance image as the guidance information;
[0167] Obtain the texture generation guide image of the geometric model and determine the texture generation guide image as the guide information;
[0168] The text generation guidance speech of the geometric model is obtained, and speech recognition processing is performed on the text generation guidance speech to obtain the text generation guidance text. The text generation guidance text is then processed by text-to-image conversion to obtain the text generation guidance image, and the text generation guidance image is identified as the guidance information.
[0169] In one implementation, the acquisition unit 1101 is also used to perform the following steps:
[0170] Acquire the orientation information of each viewpoint in at least one viewpoint, and the orientation information of each viewpoint is used to indicate the orientation of each viewpoint;
[0171] Processing unit 1102 is used to perform texture rendering processing on the geometric model based on the texture information in the guidance information and the depth map of the geometric model in each viewpoint. When obtaining the texture map of the geometric model, it specifically performs the following steps:
[0172] Based on the texture information in the guidance information, the depth map of the geometric model in each view, and the orientation information of each view, the geometric model is processed by texture rendering to obtain the texture map of the geometric model.
[0173] In one implementation, the texture rendering process is performed N times, where N is a positive integer. The processing unit 1102, based on the texture information in the guidance information, the depth map of the geometric model at each viewpoint, and the orientation information at each viewpoint, specifically performs the following steps when performing the i-th texture rendering process on the geometric model:
[0174] When i=1, based on the texture information in the guidance information, the depth map of the geometric model in each view, and the orientation information of each view, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model in each view; the color view of the geometric model in each view is subjected to texture map conversion processing to obtain the shared texture map corresponding to the i-th texture rendering processing.
[0175] When i∈[2,N], based on the texture information in the guidance information, the depth map of the geometric model in each view, the direction information of each view, and the shared texture map corresponding to the i-1th texture rendering process, the geometric model is subjected to texture color rendering process to obtain the color view of the geometric model in each view; the color view of the geometric model in each view is subjected to texture map conversion process to obtain the shared texture map corresponding to the i-th texture rendering process.
[0176] The texture map of the geometric model is generated based on the shared texture map corresponding to the Nth texture rendering process.
[0177] In one implementation, any one of the at least three perspectives is represented as the target perspective;
[0178] Processing unit 1102 is used to perform texture and color rendering processing on the geometric model based on the texture information in the guidance information, the depth map of the geometric model in the target view, and the orientation information of the target view, to obtain a color view of the geometric model in the target view. Specifically, it performs the following steps:
[0179] The guidance information is processed by feature extraction to obtain the texture features corresponding to the texture information;
[0180] Feature extraction is performed on the depth map from the target's perspective to obtain the geometric features corresponding to the target's perspective.
[0181] The directional information of the target viewpoint is processed by feature extraction to obtain the directional features corresponding to the target viewpoint.
[0182] The texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model from the target viewpoint.
[0183] In one implementation, the processing unit 1102 is used to perform texture color rendering processing on the geometric model based on the texture information in the guidance information, the depth map of the geometric model in each view, the orientation information of each view, and the shared texture map corresponding to the (i-1)th texture rendering process, to obtain a color view of the geometric model in each view. Specifically, it is used to perform the following steps:
[0184] Perform view splitting on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model from each perspective.
[0185] Based on the texture information in the guidance information, the depth map of the geometric model in each view, the orientation information in each view, and the feedback view in each view, the geometric model is processed for texture color rendering to obtain the color view of the geometric model in each view.
[0186] In one implementation, any one of the at least three perspectives is represented as the target perspective;
[0187] Processing unit 1102 is used to perform texture and color rendering processing on the geometric model based on guidance information, the depth map of the geometric model in the target view, the orientation information of the target view, and the feedback view in the target view, to obtain a color view of the geometric model in the target view. Specifically, it performs the following steps:
[0188] Obtain the texture features corresponding to the texture information in the guidance information. The texture features are obtained by performing feature extraction processing on the guidance information.
[0189] Obtain the geometric features corresponding to the target viewpoint. The geometric features corresponding to the target viewpoint are obtained by feature extraction processing of the depth map under the target viewpoint.
[0190] The directional features corresponding to the target viewpoint are obtained by performing feature extraction processing on the directional information of the target viewpoint.
[0191] Feature extraction processing is performed on the feedback view from the target's perspective to obtain the feedback features corresponding to the target's perspective.
[0192] The texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model from the target viewpoint.
[0193] In one implementation, when processing unit 1102 performs view splitting processing on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model from each perspective, it specifically performs the following steps:
[0194] Based on the mapping relationship between the pixels in the shared texture map corresponding to the i-1th texture rendering process and the vertices of the geometric model, the pixel value of the pixel in the shared texture map corresponding to the i-1th texture rendering process is used as the pixel value of the vertex mapped by the pixel on the geometric model.
[0195] Determine the vertices observed on the geometric model from each viewpoint. Based on the pixel values of the vertices observed from each viewpoint, perform view rendering processing on the geometric model from each viewpoint to obtain the feedback view of the geometric model from each viewpoint.
[0196] In one implementation, the processing unit 1102, when performing texture mapping conversion on the color view of the geometric model from various perspectives to obtain the shared texture map corresponding to the i-th texture rendering process, specifically performs the following steps:
[0197] Map the color view of the geometric model from each perspective onto the texture template of the geometric model to obtain the mapped texture template;
[0198] In the mapped texture template, different textures mapped to the color view from different perspectives are subjected to texture fusion processing to obtain the shared texture map corresponding to the i-th texture rendering process.
[0199] In one implementation, the Nth texture rendering process is performed in the encoding space; the processing unit 1102, when generating the texture map of the geometric model based on the shared texture map corresponding to the Nth texture rendering process, specifically performs the following steps:
[0200] Perform view splitting on the shared texture map corresponding to the Nth texture rendering process to obtain the feedback view of the geometric model from each perspective.
[0201] The feedback view of the geometric model under each viewpoint is decoded to obtain the decoded view of the geometric model under each viewpoint.
[0202] Texture mapping is performed on the decoded views from various perspectives to obtain the texture maps of the geometric models.
[0203] According to one embodiment of this application, the various units in the texture generation apparatus shown in FIG11 can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This can achieve the same operation without affecting the technical effect of the embodiment of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the texture generation apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0204] According to another embodiment of this application, the texture generation apparatus shown in FIG11, and the texture generation method of this application embodiment, can be constructed and implemented by running a computer program capable of performing some or all of the steps involved in the methods shown in FIG4 or FIG7 on a general-purpose computing device including processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM). The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the aforementioned computing device through the computer-readable storage medium, and run therein.
[0205] In this embodiment, the texture information in the guidance information of the geometric model can be used to indicate a preset texture style, and the depth map of the geometric model at each viewpoint can be used to indicate the geometric information of the geometric model at each viewpoint. Based on the texture information in the guidance information and the depth map of the geometric model at each viewpoint, texture rendering processing can be performed on the geometric model to obtain a texture map of the geometric model. It can be seen that the guidance information in this embodiment provides texture generation guidance, specifically texture style generation guidance. Using the texture information in the guidance information to guide the generation of the geometric model's texture ensures that the generated actual texture style is consistent with the preset texture style indicated by the texture information. That is, given any preset texture style, based on the guidance of this texture style, a texture map with a texture style consistent with the given preset texture style can be generated. Therefore, the generation quality of the texture map can be improved from the consistency dimension between the generated actual texture style and the given preset texture style. Furthermore, this embodiment utilizes the depth map of the geometric model at each viewpoint to guide texture generation. The depth map of the geometric model at each viewpoint acts as a geometric guide during texture generation, enabling the generated texture map to better adapt to the geometric model. This improves the quality of texture map generation by addressing the compatibility between the generated texture style and the geometric model. In summary, this embodiment can improve the quality of texture map generation.
[0206] Please refer to Figure 12, which is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device shown in Figure 12 includes at least a processor 1201, an input interface 1202, an output interface 1203, and a computer-readable storage medium 1204. The processor 1201, input interface 1202, output interface 1203, and computer-readable storage medium 1204 can be connected via a bus or other means.
[0207] The computer-readable storage medium 1204 can be stored in the memory of a computer device. The computer-readable storage medium 1204 is used to store computer programs, which include computer instructions. The processor 1201 is used to execute the computer program stored in the computer-readable storage medium 1204. The processor 1201 (or CPU (Central Processing Unit)) is the computing and control core of the computer device. It is suitable for implementing computer programs, specifically for loading and executing computer programs to achieve corresponding methods or functions.
[0208] This application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space for storing the operating system of the computer device. Furthermore, the storage space also stores computer programs suitable for loading and execution by a processor. It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.
[0209] The computer device may be, for example, a terminal device, or, for example, a server, or, alternatively, a texture generation system composed of a terminal device and a server. When the computer device is a texture generation system composed of a terminal device and a server, the processor 1201, input interface 1202, output interface 1203, and computer-readable storage medium 1204 are distributed throughout the texture generation system. This distributed distribution means that a portion of the processor 1201, input interface 1202, output interface 1203, and computer-readable storage medium 1204 are located in the terminal device of the texture generation system, while another portion is located in the server of the texture generation system. For example, the input interface 1202 and output interface 1203 of the computer device are located in the terminal device of the texture generation system, while the processor 1201 and computer-readable storage medium 1204 of the computer device are located in the server of the texture generation system.
[0210] In a specific implementation, the processor 1201 can load and execute the computer program stored in the computer-readable storage medium 1204 to implement the corresponding steps in the methods shown in Figure 4 or Figure 7. Specifically, the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 in the following steps:
[0211] Obtain the guidance information and mesh file of the geometric model. The guidance information contains texture information, which is used to indicate the preset texture style.
[0212] Based on the mesh file of the geometric model, at least one depth rendering camera is used to perform depth rendering on the geometric model from at least one viewpoint to obtain the depth map of the geometric model in each viewpoint; at least one depth rendering camera corresponds to one viewpoint respectively;
[0213] Based on the texture information in the guidance information and the depth map of the geometric model in each view, the geometric model is processed by texture rendering to obtain the texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
[0214] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to obtain the guidance information of the geometric model, it is specifically used to perform any of the following:
[0215] Obtain the texture generation guidance text of the geometric model, perform text-to-image conversion on the texture generation guidance text to obtain the texture generation guidance image, and determine the texture generation guidance image as the guidance information;
[0216] Obtain the texture generation guide image of the geometric model and determine the texture generation guide image as the guide information;
[0217] The text generation guidance speech of the geometric model is obtained, and speech recognition processing is performed on the text generation guidance speech to obtain the text generation guidance text. The text generation guidance text is then processed by text-to-image conversion to obtain the text generation guidance image, and the text generation guidance image is identified as the guidance information.
[0218] In one implementation, the computer program in the computer-readable storage medium 1204 is loaded by the processor 1201 and is also used to perform the following steps:
[0219] Acquire the orientation information of each viewpoint in at least one viewpoint, and the orientation information of each viewpoint is used to indicate the orientation of each viewpoint;
[0220] The computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201. When performing texture rendering processing on the geometric model based on the texture information in the boot information and the depth map of the geometric model at each viewpoint to obtain the texture map of the geometric model, it is specifically used to perform the following steps:
[0221] Based on the texture information in the guidance information, the depth map of the geometric model in each view, and the orientation information of each view, the geometric model is processed by texture rendering to obtain the texture map of the geometric model.
[0222] In one implementation, the texture rendering process is performed N times, where N is a positive integer. The computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201. Based on the guidance information, the depth map of the geometric model at each viewpoint, and the orientation information of each viewpoint, when performing the i-th texture rendering process on the geometric model, it is specifically used to perform the following steps:
[0223] When i=1, based on the texture information in the guidance information, the depth map of the geometric model in each view, and the orientation information of each view, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model in each view; the color view of the geometric model in each view is subjected to texture map conversion processing to obtain the shared texture map corresponding to the i-th texture rendering processing.
[0224] When i∈[2,N], based on the texture information in the guidance information, the depth map of the geometric model in each view, the direction information of each view, and the shared texture map corresponding to the i-1th texture rendering process, the geometric model is subjected to texture color rendering process to obtain the color view of the geometric model in each view; the color view of the geometric model in each view is subjected to texture map conversion process to obtain the shared texture map corresponding to the i-th texture rendering process.
[0225] The texture map of the geometric model is generated based on the shared texture map corresponding to the Nth texture rendering process.
[0226] In one implementation, any one of the at least three perspectives is represented as the target perspective;
[0227] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201, based on the texture information in the boot information, the depth map of the geometric model from the target viewpoint, and the orientation information from the target viewpoint, performs texture and color rendering processing on the geometric model to obtain a color view of the geometric model from the target viewpoint, it specifically performs the following steps:
[0228] The guidance information is processed by feature extraction to obtain the texture features corresponding to the texture information;
[0229] Feature extraction is performed on the depth map from the target's perspective to obtain the geometric features corresponding to the target's perspective.
[0230] The directional information of the target viewpoint is processed by feature extraction to obtain the directional features corresponding to the target viewpoint.
[0231] The texture features, geometric features corresponding to the target viewpoint, and directional features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model from the target viewpoint.
[0232] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to perform texture color rendering processing on the geometric model based on the texture information in the boot information, the depth map of the geometric model in each view, the orientation information of each view, and the shared texture map corresponding to the (i-1)th texture rendering process, and to obtain a color view of the geometric model in each view, the program specifically performs the following steps:
[0233] Perform view splitting on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model from each perspective.
[0234] Based on the texture information in the guidance information, the depth map of the geometric model in each view, the orientation information in each view, and the feedback view in each view, the geometric model is processed for texture color rendering to obtain the color view of the geometric model in each view.
[0235] In one implementation, any one of the at least three perspectives is represented as the target perspective;
[0236] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201, based on the texture information in the boot information, the depth map of the geometric model in the target view, the orientation information of the target view, and the feedback view in the target view, performs texture color rendering processing on the geometric model to obtain a color view of the geometric model in the target view, it is specifically used to perform the following steps:
[0237] Obtain the texture features corresponding to the texture information in the guidance information. The texture features are obtained by performing feature extraction processing on the guidance information.
[0238] Obtain the geometric features corresponding to the target viewpoint. The geometric features corresponding to the target viewpoint are obtained by feature extraction processing of the depth map under the target viewpoint.
[0239] The directional features corresponding to the target viewpoint are obtained by performing feature extraction processing on the directional information of the target viewpoint.
[0240] Feature extraction processing is performed on the feedback view from the target's perspective to obtain the feedback features corresponding to the target's perspective.
[0241] The texture features, geometric features corresponding to the target viewpoint, directional features corresponding to the target viewpoint, and feedback features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model from the target viewpoint.
[0242] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to perform view splitting processing on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model in each viewpoint, it specifically performs the following steps:
[0243] Based on the mapping relationship between the pixels in the shared texture map corresponding to the i-1th texture rendering process and the vertices of the geometric model, the pixel value of the pixel in the shared texture map corresponding to the i-1th texture rendering process is used as the pixel value of the vertex mapped by the pixel on the geometric model.
[0244] Determine the vertices observed on the geometric model from each viewpoint. Based on the pixel values of the vertices observed from each viewpoint, perform view rendering processing on the geometric model from each viewpoint to obtain the feedback view of the geometric model from each viewpoint.
[0245] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to perform texture mapping conversion processing on the color views of the geometric model from various perspectives to obtain the shared texture map corresponding to the i-th texture rendering process, it specifically performs the following steps:
[0246] Map the color view of the geometric model from each perspective onto the texture template of the geometric model to obtain the mapped texture template;
[0247] In the mapped texture template, different textures mapped to the color view from different perspectives are subjected to texture fusion processing to obtain the shared texture map corresponding to the i-th texture rendering process.
[0248] In one implementation, the Nth texture rendering process is performed in the encoding space; when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to generate a texture map of the geometric model based on the shared texture map corresponding to the Nth texture rendering process, it specifically performs the following steps:
[0249] Perform view splitting on the shared texture map corresponding to the Nth texture rendering process to obtain the feedback view of the geometric model from each perspective.
[0250] The feedback view of the geometric model under each viewpoint is decoded to obtain the decoded view of the geometric model under each viewpoint.
[0251] Texture mapping is performed on the decoded views from various perspectives to obtain the texture maps of the geometric models.
[0252] In this embodiment, the texture information in the guidance information of the geometric model can be used to indicate a preset texture style, and the depth map of the geometric model at each viewpoint can be used to indicate the geometric information of the geometric model at each viewpoint. Based on the texture information in the guidance information and the depth map of the geometric model at each viewpoint, texture rendering processing can be performed on the geometric model to obtain a texture map of the geometric model. It can be seen that the guidance information in this embodiment provides texture generation guidance, specifically texture style generation guidance. Using the texture information in the guidance information to guide the generation of the geometric model's texture ensures that the generated actual texture style is consistent with the preset texture style indicated by the texture information. That is, given any preset texture style, based on the guidance of this texture style, a texture map with a texture style consistent with the given preset texture style can be generated. Therefore, the generation quality of the texture map can be improved from the consistency dimension between the generated actual texture style and the given preset texture style. Furthermore, this embodiment utilizes the depth map of the geometric model at each viewpoint to guide texture generation. The depth map of the geometric model at each viewpoint acts as a geometric guide during texture generation, enabling the generated texture map to better adapt to the geometric model. This improves the quality of texture map generation by addressing the compatibility between the generated texture style and the geometric model. In summary, this embodiment can improve the quality of texture map generation.
[0253] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the texture generation method described above.
Claims
1. A method of texture generation, characterized by, The method is performed by a computer device, and the method includes: Obtain the guidance information of the geometric model and the mesh file of the geometric model. The guidance information includes texture information, which is used to indicate a preset texture style. Based on the mesh file of the geometric model, at least one depth rendering camera is used to perform depth rendering processing on the geometric model from at least one viewpoint to obtain a depth map of the geometric model under each viewpoint, wherein the at least one depth rendering camera corresponds to one viewpoint respectively; Based on the texture information in the guidance information and the depth map of the geometric model under each viewpoint, the geometric model is subjected to texture rendering processing to obtain a texture map of the geometric model. The actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
2. The method of claim 1, wherein, The guidance information for obtaining the geometric model includes any of the following: The texture generation guidance text of the geometric model is obtained, the text-to-image conversion processing of the texture generation guidance text is performed to obtain the texture generation guidance image, and the texture generation guidance image is determined as the guidance information; Obtain the texture generation guide image of the geometric model, and determine the texture generation guide image as the guide information; The texturing guidance speech of the geometric model is obtained, and speech recognition processing is performed on the texturing guidance speech to obtain texturing guidance text; the texturing guidance text is subjected to text-to-image conversion processing to obtain a texturing guidance image, and the texturing guidance image is determined as the guidance information.
3. The method of claim 1 or 2, wherein, The method further includes: Obtain the direction information of each of the at least one viewpoints, wherein the direction information of each viewpoint is used to indicate the direction of each viewpoint; The process of performing texture rendering on the geometric model based on the texture information in the guidance information and the depth map of the geometric model under each viewpoint includes: Based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, and the direction information of each viewpoint, the geometric model is subjected to texture rendering processing to obtain the texture map of the geometric model.
4. The method according to any one of claims 1 to 3, characterized in that, The texture rendering process is performed N times, where N is a positive integer; the process of performing the i-th texture rendering process on the geometric model based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, and the direction information of each viewpoint includes: When i=1, based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, and the direction information of each viewpoint, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model under each viewpoint; the color view of the geometric model under each viewpoint is subjected to texture map conversion processing to obtain the shared texture map corresponding to the i-th texture rendering processing. When i∈[2,N], based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, the direction information of each viewpoint, and the shared texture map corresponding to the (i-1)th texture rendering process, the geometric model is subjected to texture color rendering to obtain the color view of the geometric model under each viewpoint; the color view of the geometric model under each viewpoint is subjected to texture map conversion to obtain the shared texture map corresponding to the i-th texture rendering process. The texture map of the geometric model is generated based on the shared texture map corresponding to the Nth texture rendering process.
5. The method according to any one of claims 1 to 4, wherein Any one of the at least one perspectives is referred to as the target perspective; Based on the texture information in the guidance information, the depth map of the geometric model from the target viewpoint, and the orientation information of the target viewpoint, the process of performing texture color rendering processing on the geometric model to obtain a color view of the geometric model from the target viewpoint includes: The guidance information is subjected to feature extraction processing to obtain the texture features corresponding to the texture information; The depth map from the target viewpoint is subjected to feature extraction processing to obtain the geometric features corresponding to the target viewpoint. The directional information of the target viewpoint is processed by feature extraction to obtain the directional features corresponding to the target viewpoint; The texture features, the geometric features corresponding to the target viewpoint, and the directional features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model under the target viewpoint.
6. The method according to any one of claims 1-5, wherein performing texture color rendering processing on the geometric model based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, the orientation information of each viewpoint, and the shared texture map corresponding to the (i-1)th texture rendering process to obtain a color view of the geometric model under each viewpoint includes: The shared texture map corresponding to the (i-1)th texture rendering process is split into views to obtain the feedback view of the geometric model under each viewpoint. Based on the texture information in the guidance information, the depth map of the geometric model under each viewpoint, the direction information of each viewpoint, and the feedback view under each viewpoint, the geometric model is subjected to texture color rendering processing to obtain the color view of the geometric model under each viewpoint.
7. The method according to any one of claims 1 to 6, wherein Any one of the at least one perspectives is referred to as the target perspective; Based on the texture information in the guidance information, the depth map of the geometric model from the target viewpoint, the orientation information of the target viewpoint, and the feedback view from the target viewpoint, the process of performing texture color rendering processing on the geometric model to obtain a color view of the geometric model from the target viewpoint includes: Obtain the texture features corresponding to the texture information in the guidance information, wherein the texture features are obtained by performing feature extraction processing on the guidance information; Obtain the geometric features corresponding to the target viewpoint, which are obtained by performing feature extraction processing on the depth map under the target viewpoint; The directional features corresponding to the target viewpoint are obtained by performing feature extraction processing on the directional information of the target viewpoint. The feedback view from the target perspective is subjected to feature extraction processing to obtain the feedback features corresponding to the target perspective. The texture features, the geometric features corresponding to the target viewpoint, the directional features corresponding to the target viewpoint, and the feedback features corresponding to the target viewpoint are integrated to obtain a color view of the geometric model under the target viewpoint.
8. The method according to any one of claims 1 to 7, wherein, The step of performing view splitting on the shared texture map corresponding to the (i-1)th texture rendering process to obtain the feedback view of the geometric model under each viewpoint includes: Based on the mapping relationship between the pixels in the shared texture map corresponding to the (i-1)th texture rendering process and the vertices of the geometric model, the pixel values of the pixels in the shared texture map corresponding to the (i-1)th texture rendering process are used as the pixel values of the vertices mapped by the pixels on the geometric model. Determine the vertices observed on the geometric model from each viewpoint, and perform view rendering processing on the geometric model from each viewpoint based on the pixel values of the vertices observed from each viewpoint to obtain the feedback view of the geometric model from each viewpoint.
9. The method according to any one of claims 1 to 8, wherein, The process of performing texture mapping conversion on the color views of the geometric model under each of the aforementioned viewpoints to obtain the shared texture map corresponding to the i-th texture rendering process includes: The color view of the geometric model under each viewpoint is mapped onto the texture template of the geometric model to obtain the mapped texture template; In the mapped texture template, different textures mapped to the color view from different perspectives are subjected to texture fusion processing to obtain the shared texture map corresponding to the i-th texture rendering process.
10. The method of any one of claims 1-9, wherein, N texture rendering processes are performed in the encoding space; The process of generating the texture map of the geometric model based on the shared texture map corresponding to the Nth texture rendering process includes: The shared texture map corresponding to the Nth texture rendering process is split into views to obtain the feedback view of the geometric model under each viewpoint. The feedback view of the geometric model under each viewpoint is decoded to obtain the decoded view of the geometric model under each viewpoint. The decoded views from each of the aforementioned perspectives are subjected to texture mapping conversion to obtain the texture maps of the geometric model.
11. A texture generation apparatus characterized by comprising: include: The acquisition unit is used to acquire the guidance information of the geometric model and the mesh file of the geometric model. The guidance information includes texture information, which is used to indicate a preset texture style. The processing unit is configured to perform depth rendering processing on the geometric model from at least one viewpoint using at least one depth rendering camera based on the mesh file of the geometric model, and obtain a depth map of the geometric model under each viewpoint. Each of the at least one depth rendering camera corresponds to one of the viewpoints; The processing unit is further configured to perform texture rendering processing on the geometric model based on the texture information in the guidance information and the depth map of the geometric model under each viewpoint, to obtain a texture map of the geometric model; the actual texture style of the texture map is consistent with the preset texture style indicated by the texture information.
12. A computer device, comprising: The computer device includes: A processor is a tool for implementing computer programs. A computer-readable storage medium storing a computer program adapted to be loaded by the processor and executed as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-10.
14. A computer program product, characterised in that, The computer program product includes a computer program that, when executed by a processor, implements the texture generation method as described in any one of claims 1-10.