Texture map generation method, device and equipment
By fusing and processing noise maps from multiple viewing angles in the UV space, the problem of poor UV map quality is solved, and the view consistency of texture and color is achieved, and the map defects are avoided.
Patent Information
- Application Number
- CN202411710655.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-02
AI Technical Summary
In the prior art, when generating UV maps, it is easy to cause map flaws, such as color misalignment at the joints, resulting in poor quality of UV maps.
By obtaining the basic information of the three-dimensional model and multi-view feature maps, the trained diffusion model predicts the noise maps of multiple viewing angles, and project them into a unified UV space for fusion processing. After denoising, the target UV feature map is generated, and then texture maps are generated.
Ensure consistency of texture and/or colors in different views, avoid color misalignment and texture distortion at seams, and improve the quality and consistency of UV maps.
Smart Images

Figure CN119919560A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a texture map generation method, device and equipment. Background Art
[0002] In the field of three-dimensional content creation, especially in the refined texture editing of map building models and character models, UV mapping generation technology plays a vital role. Among them, UV mapping is a technology that maps two-dimensional images (textures) to the surface of three-dimensional models. Through UV mapping, detailed textures and colors can be added to three-dimensional models, thereby enhancing the realism and visual effects of the models.
[0003] In the related art, when generating UV maps for map building models, character models and other models, detailed text information about the three-dimensional model is usually obtained, including information such as shape, structure, texture features, etc., and then multiple model views from different perspectives are generated based on the text information, and then the multiple model views are projected into the UV space to generate a UV map.
[0004] However, the display effects corresponding to the same position in different views may be different, such as color misalignment at the seams, which can easily cause mapping defects and result in poor quality of the generated UV map. Summary of the invention
[0005] The present application provides a texture map generation method, device and equipment, which are used to solve the problem of poor quality of existing generated UV maps.
[0006] In a first aspect, the present application provides a texture map generation method, the method comprising:
[0007] Acquire basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles;
[0008] According to the basic information and the initial multi-view feature map, noise maps of multiple viewpoints are predicted using a trained diffusion model;
[0009] The noise maps of the multiple viewing angles are projected into the UV space and then fused, and the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-view feature map in the UV space;
[0010] A texture map of the three-dimensional model is generated based on the target UV feature map.
[0011] Optionally, the noise maps of the multiple perspectives are projected into a UV space and then fused, including:
[0012] Projecting the noise maps of the multiple viewing angles into a UV space to obtain a UV noise map;
[0013] A weighted average operation is performed on the pixel values of the pixel points in the UV noise map to obtain a fused noise map.
[0014] Optionally, the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, including:
[0015] The UV feature map is subtracted from the fused noise map to obtain a target UV feature map.
[0016] Optionally, the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, including:
[0017] De-noising the UV feature map according to the fused noise map to obtain a first UV feature map;
[0018] Rendering the first UV feature map to obtain a first multi-view feature map;
[0019] According to the basic information and the first multi-view feature map, at least one round of denoising processing is performed based on the diffusion model to obtain a target UV feature map, wherein the target UV feature map is the first UV feature map obtained in the last round;
[0020] Among them, the denoising processing of each round includes: according to the characteristics of the basic information and the first multi-view feature map of the previous round, based on the diffusion model, obtaining the first noise maps of multiple viewpoints in the current round, projecting the first noise maps of the multiple viewpoints into the UV space and performing fusion processing, denoising the first UV feature map of the previous round according to the first noise map after the fusion processing to obtain the first UV feature map of the current round, and obtaining the first multi-view feature map of the current round according to the first UV feature map.
[0021] Optionally, the method further includes:
[0022] Extracting a normal map corresponding to the initial multi-view feature map;
[0023] The method of predicting noise maps of multiple perspectives using a trained diffusion model according to the basic information and the initial multi-perspective feature map includes:
[0024] The normal map, the features of the basic information and the initial multi-view feature map are input into the diffusion model to predict noise maps of multiple viewpoints.
[0025] Optionally, the basic information is an initial image of the three-dimensional model at any viewing angle, and the method further includes:
[0026] The feature extraction is performed on the initial image of any perspective through a feature extraction model to obtain the features of the basic information, wherein the feature extraction model includes a plurality of feature extraction layers connected in sequence, which are used to extract and output features of different levels respectively.
[0027] Optionally, the training process of the diffusion model includes:
[0028] Acquire multiple training sets, where each training set includes basic information corresponding to the three-dimensional model, an initial multi-view feature map, and manually processed noise maps of multiple viewpoints;
[0029] For one of the training sets, noise prediction is performed based on the initial diffusion model according to the basic information in the training set and the initial multi-view feature map, and the parameters of the initial diffusion model are adjusted according to the predicted noise maps of multiple views and the noise maps of multiple views in the training set to obtain a first diffusion model;
[0030] The remaining training sets are used to adjust the parameters of the first diffusion model one by one to obtain a trained diffusion model.
[0031] Optionally, the method further includes:
[0032] Displaying the texture map through a terminal interface, and updating the texture map in response to a user's adjustment operation on the texture map; and / or,
[0033] According to the multi-view feature map rendered by the target UV feature map, images corresponding to the three-dimensional model at the multiple view angles are generated, the images corresponding to the multiple view angles are displayed through a terminal interface, and in response to a user's adjustment operation on the image corresponding to at least one view angle, the target UV feature map and the texture map are updated according to the adjusted image.
[0034] In a second aspect, the present application provides a texture map generation device, the device comprising:
[0035] An acquisition module, used to acquire basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles;
[0036] A prediction module, configured to predict noise maps of multiple perspectives using a trained diffusion model based on the basic information and the initial multi-perspective feature map;
[0037] A denoising module is used to project the noise maps of the multiple viewing angles into the UV space and then perform fusion processing, and denoise the UV feature map according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-view feature map in the UV space;
[0038] A generation module is used to generate a texture map of the three-dimensional model based on the target UV feature map.
[0039] In a third aspect, the present application provides an electronic device, including:
[0040] at least one processor; and
[0041] a memory communicatively coupled to the at least one processor;
[0042] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method as described in any one of the first aspects.
[0043] In summary, the present application provides a texture map generation method, device and equipment. By using the diffusion model technology, the basic information that can characterize the texture and / or shape of the three-dimensional model and the preset noise images under multiple perspectives are used as the input of the model, and the noise maps of multiple perspectives are predicted by the diffusion model trained in advance, and then the noise maps of multiple perspectives are projected into a unified UV space for fusion processing. The unified UV space provides a consistent processing framework, so that the noise processing under different perspectives can be performed in the same space to ensure the consistency of the processed noise map under different perspectives, and avoid processing differences caused by different perspectives. Further, the UV feature map is denoised according to the fused noise map to obtain the target UV feature map, and then the texture map of the three-dimensional model is generated based on the target UV feature map. Because the present application can ensure the consistency of texture and / or color under different views by fusing the noise map in the same UV space, it can avoid the color dislocation and texture distortion at the seams. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application;
[0046] Figure 2 A schematic diagram of a process for generating a texture map provided in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of a flow chart of an optional texture map generation method provided in an embodiment of the present application;
[0048] Figure 4 A schematic diagram of the structure of a texture map generating device provided in an embodiment of the present application;
[0049] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0050] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0051] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0052] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0053] First, some of the terms involved in this application are explained:
[0054] Diffusion model: refers to a generative model technology that generates images, texts and other content by iteratively diffusing and contracting Gaussian noise. Usually, the processing process of the diffusion model is: gradually adding noise to the data through a forward process, and then predicting the noise added at each step through a reverse process, and gradually restoring the noise-free data by removing the noise. However, in this application, the trained diffusion model is used to predict the noise added at each step, and there is no need to restore the noise-free data.
[0055] Taking an image as an example, through the forward process, Gaussian noise is gradually added to the image in a diffusion manner to generate a Gaussian noise map. Further, based on the Gaussian noise map, the inverse process of the forward process is used, that is, through the reverse process, the Gaussian noise is gradually added to the image in a diffusion manner to generate a Gaussian noise map. Figure 1 The added Gaussian noise is subtracted step by step and then converted into the original image. The above is the processing process of the diffusion model, while the reverse process is used in the process of texture map generation.
[0056] SD (Stable Diffusion): An open source image generation model. This application can be implemented using the open source code of the SD model.
[0057] UV Mapping is a technique widely used in computer graphics and 3D modeling to apply a 2D image (texture) to the surface of a 3D model.
[0058] Among them, UV map can also be called texture map.
[0059] In one possible implementation, detailed text information about the three-dimensional model, including shape, structure, texture features and other information, can be obtained, and multiple model views from different perspectives can be generated based on the text information, and then the multiple model views can be projected into the UV space to generate a UV map.
[0060] In the process of generating a plurality of model views under different perspectives, the model views under different perspectives may also be manually created or adjusted.
[0061] However, when using text information to generate UV maps, the description of the text information may not be detailed enough, such as it is difficult to clearly indicate texture details, which leads to poor quality of the generated UV maps.
[0062] In another possible implementation, an image of the three-dimensional model can be obtained and then converted into text information. The text information can be used as the input of the cross-attention network model to generate multiple model views from different perspectives. Furthermore, the multiple model views can be denoised and the denoised multiple model views can be projected into the UV space to generate a UV map, wherein the image is used as a reference feature of the network model.
[0063] It should be noted that when the above two methods generate model views under different viewing angles, each viewing angle needs to be processed by a denoising algorithm. After processing, the display results corresponding to the same position under different views may be different, such as color misalignment at the seams, that is, the color of the seams of the model at two adjacent viewing angles is inconsistent. This can easily cause mapping defects, which in turn causes the generated UV mapping to have poor quality and consistency.
[0064] In response to the above problems, the present application provides a texture map generation method. By using the diffusion model technology, the basic information that can characterize the texture and / or shape of the three-dimensional model and the preset noise images under multiple viewing angles are used as the input of the model. The noise maps of multiple viewing angles are predicted by the diffusion model trained in advance, and then the noise maps of multiple viewing angles are projected into a unified UV space for fusion processing. The unified UV space provides a consistent processing framework, so that the noise processing under different viewing angles can be performed in the same space to ensure the consistency of the processed noise map under different viewing angles, and avoid processing differences caused by different viewing angles. Furthermore, the UV feature map is denoised according to the fused noise map to obtain the target UV feature map, and then the texture map of the three-dimensional model is generated based on the target UV feature map. Because the present application can ensure the consistency of texture and / or color under different views by fusing the noise map in the same UV space, it can avoid the color dislocation and texture distortion at the seams.
[0065] For example, Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application, such as Figure 1 As shown, the application scenario can be applied to a texture map generation scenario of any map building, and the application scenario includes a user's terminal device 101 and a data processing system 102 .
[0066] Taking the generation of a texture map of a map building as an example, when a user reconstructs a three-dimensional model of a map building based on the terminal device 101, it is necessary to generate the texture map of the map building so as to accurately place the texture map on the surface of the three-dimensional model.
[0067] For example, the user can obtain the image of the map building based on the terminal device 101, such as Figure 1 The image shown in A in FIG. 1 , further, the terminal device 101 can send the image to the data processing system 102 for processing. Taking the multi-view feature map as a six-view feature map as an example, the processing process of the data processing system 102 includes: taking the initial six-view feature map and the image as the input of the model, using the trained diffusion model to predict the noise map corresponding to the initial six-view feature map, that is, the noise image under multiple perspectives, and then projecting the noise map corresponding to the six-view feature map into a unified UV space for fusion processing, and further, using the fused noise map to denoise the initial UV feature map corresponding to the image to obtain the target UV feature map, and then generating a texture map of the three-dimensional model based on the target UV feature map.
[0068] Optionally, after the fusion-processed noise map is used to denoise the initial UV feature map corresponding to the image to update the UV feature map, the quality of the updated UV feature map may be poor or may not meet the user's requirements. In this case, the updated UV feature map may be used to render a new six-view feature map, such as Figure 1 The six-view feature maps corresponding to B, C, D, E, F, and G are obtained, and then the next round of denoising is performed based on the new six-view feature maps, that is, the new six-view feature maps and the image are used as the input of the model, and the above process is performed again. The above process can be performed at least once, so that a texture map with quality that meets user needs or with better quality can be generated.
[0069] Optionally, the six viewing angles can be: the first four viewing angles are horizontal viewing angles, and the rotation angles are 0, 90, 180, and 270 degrees; the last two viewing angles are 30-degree overhead viewing angles, and the rotation angles are 0 and 180 degrees. Under these six viewing angles, projection is performed to obtain a six-view feature map.
[0070] Furthermore, the data processing system 102 may send the generated texture map to the user's terminal device 101 for visual display so that the user can view it.
[0071] Optionally, the initial six-view feature map can be obtained from the data processing system 102 through the terminal device 101, and then the terminal device 101 generates a texture map based on the image and the initial six-view feature map using the texture map generation method provided in the embodiment of the present application. The embodiment of the present application does not specifically limit the execution entity of the texture map generation method provided in the present application.
[0072] The terminal device may also be referred to as user equipment (UE), mobile station (MS), mobile terminal (Mobile Terminal), terminal, etc. In practical applications, terminal devices include desktop computers, notebooks, personal digital assistants (PDA), smart phones, tablet computers, vehicle-mounted devices, wearable devices (such as smart watches and smart bracelets), smart home devices (such as smart display devices), etc.
[0073] It should be noted that the embodiments of the present application do not specifically limit the application scenarios of the present application. The above is only an example. For example, it can be applied to application scenarios with extremely high requirements on texture details, such as game development, film and television special effects, and virtual reality environment construction.
[0074] For example, Figure 2 A flow chart of a texture map generation method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the texture map generation method includes the following steps:
[0075] S201. Obtain basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles.
[0076] In an embodiment of the present application, the basic information may include image information or text information. The image information may be a two-dimensional image of a map building model, a character model, etc. including a three-dimensional model. The embodiment of the present application does not limit the specific three-dimensional model corresponding to the image information. The text information is detailed text information used to describe the three-dimensional model. The specific content can refer to the existing definition and will not be repeated here.
[0077] The initial multi-view feature map refers to the Gaussian noise map used in the reverse process. The Gaussian noise map can predict the noise map corresponding to the three-dimensional model through the reverse process of the diffusion model. The Gaussian noise map is obtained by gradually adding Gaussian noise to the basic information corresponding to the three-dimensional model in a diffusion manner based on the forward process of the diffusion model. It should be noted that the multi-view feature map is also used to indicate the three-dimensional model to be projected into several views under different perspectives. Taking the multi-view feature map as a six-view feature map as an example, the six perspectives are: the first four perspectives are horizontal perspectives with rotation angles of 0, 90, 180, and 270 degrees, and the last two perspectives are 30-degree overhead perspectives with rotation angles of 0 and 180 degrees, that is, the three-dimensional model is projected under these six perspectives.
[0078] Optionally, the six viewing angles can also be respectively that the first four viewing angles are horizontal viewing angles with rotation angles of 0, 90, 180, and 270 degrees, and the last two viewing angles are vertical overhead viewing angles with rotation angles of 0 and 180 degrees, that is, the six-view feature maps correspond to the feature maps of the front view, rear view, left view, right view, top view and bottom view.
[0079] It can be understood that the above-mentioned multi-view feature map can also be a four-view feature map, that is, the feature maps corresponding to the front view, the rear view, the left view and the right view, that is, four horizontal viewing angles, and the views obtained by projecting the three-dimensional model at rotation angles of 0, 90, 180, and 270 degrees. The embodiment of the present application does not specifically limit the number of views and viewing angles of the multi-view feature map, which can be determined based on the actual application scenario or the type of model.
[0080] It should be noted that the embodiment of the present application does not specifically limit the size of the multi-view feature map, that is, the arrangement of each view. Figure 1 As shown in , the size corresponding to the image of the six-view feature map can be 640*960, two rows and three columns, or 960*640, three rows and two columns.
[0081] S202: According to the basic information and the initial multi-view feature map, noise maps of multiple viewpoints are predicted using a trained diffusion model.
[0082] Optionally, the noise maps of multiple perspectives can be predicted using a trained diffusion model based on the features of the basic information and the initial multi-perspective feature map, wherein the features of the basic information correspond to the texture or color features of the three-dimensional model at different resolutions. The embodiments of the present application do not limit the specific content corresponding to the features. For example, the features of the basic information may be hierarchical features, which may be extracted through an attention cross network. The embodiments of the present application do not specifically limit the method for extracting hierarchical features, which may refer to existing methods or redefine new methods.
[0083] Exemplarily, noise maps of multiple perspectives can be predicted based on the trained diffusion model according to the hierarchical features of the basic information and the initial multi-perspective feature map.
[0084] S203, projecting the noise maps of the multiple perspectives into the UV space and performing fusion processing, denoising the UV feature map according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-perspective feature map in the UV space.
[0085] Exemplarily, after noise maps of multiple perspectives are predicted based on the trained diffusion model, the noise maps of multiple perspectives can be projected into the UV space for fusion processing, and the initial UV feature map is denoised according to the fused noise map to obtain the target UV feature map.
[0086] The initial UV feature map is a feature map corresponding to the initial multi-view feature map in the UV space.
[0087] It should be noted that the UV space is a two-dimensional plane space used to represent the texture mapping of the surface of a three-dimensional model, and the noise maps of multiple perspectives correspond to the noise to be subtracted from the UV feature map.
[0088] Optionally, the UV feature map may contain some noise or inaccurate details, and the noise map represents the noise components that need to be removed from the UV feature map. Further, the noise map is subtracted from the UV feature map pixel by pixel to obtain a denoised target UV feature map. The embodiment of the present application does not specifically limit the method of denoising the UV feature map using the noise map, such as it can be a simple subtraction or a weighted subtraction.
[0089] Optionally, since overlapping areas may appear in noise maps of different viewing angles when the noise map is projected into the UV space, the noise map in the UV space may be fused after the noise map is projected into the UV space. The noise fusion process may be a process of processing pixel data corresponding to pixel points in overlapping areas or edge areas using a weighted average algorithm or an average algorithm. The embodiment of the present application does not specifically limit the fusion process. The fusion process may ensure that the noise is consistent at different viewing angles.
[0090] S204: Generate a texture map of the three-dimensional model based on the target UV feature map.
[0091] Exemplarily, after obtaining the target UV feature map, texture drawing can be performed in 2D (Two-Dimensional) image editing software or three-dimensional drawing software to generate a texture map. The embodiment of the present application does not specifically limit the method of generating a texture map using the target UV feature map. It can use tools or existing map generation methods.
[0092] Therefore, the embodiment of the present application uses a combination of a diffusion model and UV space projection to project the noise maps under multiple perspectives predicted by the diffusion model into a unified UV space for fusion processing, thereby ensuring the coherence and harmony between the various noise maps, and then using the fused noise map to denoise the UV feature map. This can ensure the consistency of the denoising effect under different perspectives, and can also effectively reduce the noise points and defects in the obtained UV feature map, thereby improving the overall quality of the generated texture map.
[0093] Optionally, the noise maps of the multiple perspectives are projected into a UV space and then fused, including:
[0094] Projecting the noise maps of the multiple viewing angles into a UV space to obtain a UV noise map;
[0095] A weighted average operation is performed on the pixel values of the pixel points in the UV noise map to obtain a fused noise map.
[0096] Exemplarily, after the noise maps of multiple perspectives are projected into the UV space to obtain the UV noise map, there may be overlapping areas in the image features at different perspectives in the UV noise map. For the overlapping areas, the weight values corresponding to the image features at the perspective corresponding to the overlapping areas can be determined, and then the weight values and the pixel values of the pixels corresponding to the overlapping areas can be used to perform weighted averaging. The pixel values of the pixels outside the overlapping areas can be weighted or not. The embodiment of the present application does not specifically limit this, and the noise map after fusion processing can be obtained.
[0097] Optionally, after projecting the noise maps of multiple perspectives into the UV space to obtain the UV noise map, the edges of the image features at different perspectives can be determined, and then different weight values can be determined for the edges and transition areas connected to the edges, and then the pixel values of the pixel points in the UV noise map can be processed using a weighted average operation to obtain a fused noise map.
[0098] It should be noted that the embodiment of the present application does not limit the specific area corresponding to the weighted averaging operation of the pixel values of the pixel points in the UV noise image. For example, in order to determine the weight values corresponding to the pixel points in different areas, weighted averaging processing can be performed based on the weight values corresponding to the pixel points in each area and the pixel values corresponding to the pixel points.
[0099] Therefore, the embodiment of the present application uses projection and weighted averaging operations to fuse noise maps from multiple perspectives, which can integrate the details in the image and reduce the randomness and inconsistency of noise. Even if the image quality from certain perspectives is poor or there are overlapping areas in the images from certain perspectives, images from other perspectives can be used for compensation or smoothing, thereby effectively suppressing random noise under the perspective, making the denoising effect more stable and reliable.
[0100] Optionally, the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, including:
[0101] The UV feature map is subtracted from the fused noise map to obtain a target UV feature map.
[0102] In some embodiments, the pixel value corresponding to the pixel point in the UV feature map may be subtracted from the pixel value corresponding to the pixel point in the noise map to obtain the target UV feature map.
[0103] In other embodiments, the target UV feature map can be obtained by determining a first weight corresponding to the noise map after fusion processing and a second weight corresponding to the UV feature map, and based on the first weight and the second weight, using a weighted subtraction method, that is, the pixel value corresponding to the pixel point in the UV feature map * the first weight - the pixel value corresponding to the pixel point in the noise map * the second weight, to obtain the target UV feature map.
[0104] It should be noted that the embodiment of the present application does not limit the specific method of subtraction. The above is only an example. The subtraction method can also be a method of subtracting using a subtraction function.
[0105] In this way, by performing denoising by subtracting the UV feature map from the noise map, noise and defects can be removed quickly and effectively to improve the quality of the UV feature map.
[0106] Optionally, the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, including:
[0107] De-noising the UV feature map according to the fused noise map to obtain a first UV feature map;
[0108] Rendering the first UV feature map to obtain a first multi-view feature map;
[0109] According to the basic information and the first multi-view feature map, at least one round of denoising processing is performed based on the diffusion model to obtain a target UV feature map, wherein the target UV feature map is the first UV feature map obtained in the last round;
[0110] Among them, the denoising processing of each round includes: according to the characteristics of the basic information and the first multi-view feature map of the previous round, based on the diffusion model, obtaining the first noise maps of multiple viewpoints in the current round, projecting the first noise maps of the multiple viewpoints into the UV space and performing fusion processing, denoising the first UV feature map of the previous round according to the first noise map after the fusion processing to obtain the first UV feature map of the current round, and obtaining the first multi-view feature map of the current round according to the first UV feature map.
[0111] In the present application, after obtaining the first UV feature map of the current round, the first UV feature map of the current round can be re-rendered to obtain the first multi-view feature map of the current round, so as to provide more accurate input for the next stage of denoising.
[0112] Therefore, in this step, the first multi-view feature map can be continuously refined and enhanced until the iteration ends.
[0113] Exemplarily, the denoising process of the first round includes: obtaining the first noise map of multiple perspectives of the current round based on the diffusion model according to the hierarchical features of the basic information and the first multi-view feature map, and then projecting the first noise maps of the multiple perspectives into the UV space for fusion processing, denoising the first UV feature map according to the first noise map after fusion processing to obtain a new first UV feature map, and rendering according to the new first UV feature map to obtain a new first multi-view feature map.
[0114] Correspondingly, the second round of denoising process includes: obtaining the first noise map of multiple perspectives of the current round based on the diffusion model according to the hierarchical features of the basic information and the first multi-perspective feature map obtained by the first denoising process, and then projecting the first noise maps of multiple perspectives into the UV space for fusion processing, denoising the first UV feature map obtained by the first denoising process according to the first noise map after the fusion processing to obtain a new first UV feature map, and rendering according to the new first UV feature map to obtain a new first multi-perspective feature map.
[0115] It should be noted that the denoising process from the third round to the Nth round is similar to the denoising process of the first and second rounds mentioned above, and will not be repeated here. For details, please refer to the description of the above embodiment, and the size of N corresponds to the number of iterations.
[0116] It can be understood that the number of iterations corresponding to the above-mentioned denoising process is determined by the sampler, and different samplers correspond to different numbers of iterations. It can also be understood that the number of iterations is set in advance. The embodiment of the present application does not specifically limit the number of iterations. The sampler is used to select sample points from a continuous or discrete space, that is, to obtain pixel points from an image. The more sample points obtained, the more iterations there are, and vice versa, the fewer iterations there are.
[0117] Exemplarily, based on the trained diffusion model, the first multi-view feature map and the first UV feature map are subjected to at least one round of denoising processing to obtain a target UV feature map, i.e., the first UV feature map of the last round, which can characterize the features of the images corresponding to the three-dimensional model at multiple viewpoints.
[0118] Optionally, the first UV feature map of the previous round is denoised according to the first noise map after fusion processing to obtain the first UV feature map of the current round, including: subtracting the first UV feature map of the previous round from the first noise map after fusion processing to obtain the first UV feature map of the current round.
[0119] In this way, through multiple rounds of denoising processing, the noise in the UV feature map can be effectively reduced, the clarity and detail of the texture can be improved, and each round of denoising processing is performed on the basis of the previous round, so that through multiple rounds of iterative processing, high-quality UV feature maps and multi-view feature maps can be obtained.
[0120] Optionally, the method further includes:
[0121] Extracting a normal map corresponding to the initial multi-view feature map;
[0122] The method of predicting noise maps of multiple perspectives using a trained diffusion model according to the basic information and the initial multi-perspective feature map includes:
[0123] The normal map, the features of the basic information and the initial multi-view feature map are input into the diffusion model to predict noise maps of multiple viewpoints.
[0124] In an embodiment of the present application, a normal map is a texture image that describes texture details, wherein each pixel stores a normal vector, and the normal vector is usually encoded with an RGB color value, wherein R, G, and B correspond to the X, Y, and Z components of the normal vector, respectively. The normal map is used to simulate surface details, such as bumps, wrinkles, and scratches, without increasing the geometric complexity of the model. Therefore, in an embodiment of the present application, the normal map can be used to characterize the shape characteristics of a three-dimensional model at different viewing angles.
[0125] Thus, by using the normal map, the diffusion model can more accurately capture the subtle details of the 3D model surface, which helps to take into account the complexity of surface features in noise prediction.
[0126] Optionally, in each round of denoising processing, using the features of the basic information and the first multi-view feature map of the previous round, obtaining the first noise map of multiple views in the current round based on the diffusion model may include:
[0127] The normal map of the previous round, the features of the basic information and the multi-view feature map of the previous round are input into the diffusion model to obtain the noise maps of multiple viewpoints in the current round.
[0128] Exemplarily, each round of denoising processing includes: obtaining the first noise maps of multiple perspectives in the current round based on the diffusion model according to the normal map of the previous round, the features of the basic information and the first multi-view feature map of the previous round, projecting the first noise maps of the multiple perspectives into the UV space and performing fusion processing, denoising the first UV feature map of the previous round according to the first noise map after the fusion processing to obtain the first UV feature map of the current round, generating the first multi-view feature map of the current round according to the first UV feature map, and then extracting the normal map of the current round after obtaining the first multi-view feature map of the current round.
[0129] In this way, the embodiment of the present application simplifies the complexity of feature extraction by extracting the normal map, and the normal map can constrain the texture shape of the three-dimensional model and determine that the three-dimensional model maintains consistent surface details at different viewing angles. The use of features and normal maps can ensure the shape and texture consistency of the three-dimensional model at different viewing angles, which helps to more accurately predict the noise map.
[0130] Optionally, the basic information is an initial image of the three-dimensional model at any viewing angle, and the method further includes:
[0131] The feature extraction is performed on the initial image of any perspective through a feature extraction model to obtain the features of the basic information, wherein the feature extraction model includes a plurality of feature extraction layers connected in sequence, which are used to extract and output features of different levels respectively.
[0132] In the present application, any viewing angle can be selected from the three-dimensional model to obtain an initial image at the viewing angle, and the initial image is a two-dimensional representation of the three-dimensional model at a specific viewing angle.
[0133] In this step, the initial image is processed using a feature extraction model. The feature extraction model can be a deep learning model composed of multiple feature extraction layers. After being processed by the feature extraction model, the features of the initial image can be output. Since the image can better restore the texture details in the texture map, the features of the image at any viewing angle are used as the diffusion model input to accurately predict the noise maps of multiple viewing angles.
[0134] It should be noted that the embodiment of the present application does not specifically limit the feature extraction model. For example, the feature extraction model can also be an attention cross network model, a VGG (Visual Geometry Group) network, etc. Different feature extraction layers are used to extract image features of different resolutions. The embodiment of the present application does not specifically limit the number of feature extraction layers.
[0135] Thus, in the embodiment of the present application, features at different levels can be extracted through multiple feature extraction layers of the feature extraction model. The hierarchical features include image features of different resolutions, so that the diffusion model can understand the image content more comprehensively and thus perform noise prediction more accurately.
[0136] Optionally, the training process of the diffusion model includes:
[0137] Acquire multiple training sets, where each training set includes basic information corresponding to the three-dimensional model, an initial multi-view feature map, and manually processed noise maps of multiple viewpoints;
[0138] For one of the training sets, noise prediction is performed based on the initial diffusion model according to the basic information in the training set and the initial multi-view feature map, and the parameters of the initial diffusion model are adjusted according to the predicted noise maps of multiple views and the noise maps of multiple views in the training set to obtain a first diffusion model;
[0139] The remaining training sets are used to adjust the parameters of the first diffusion model one by one to obtain a trained diffusion model.
[0140] Optionally, adjusting the parameters of the initial diffusion model to obtain a first diffusion model includes:
[0141] Compare and analyze the second noise images of multiple viewing angles predicted by the initial diffusion model with the noise images of multiple viewing angles after manual processing to obtain analysis results;
[0142] The parameters of the initial diffusion model are adjusted based on the analysis results to obtain a first diffusion model.
[0143] Optionally, in order to determine whether the second noise maps of multiple perspectives predicted by the initial diffusion model can achieve denoising of the UV feature map, the UV feature map after manual processing and the initial UV feature map corresponding to the three-dimensional model can be obtained first. The manually processed UV feature map can generate a texture map with quality that meets user requirements or has good quality. Further, the predicted second noise maps of multiple perspectives are projected into the UV space and fused to obtain a second UV noise map, and then the initial UV feature map is denoised using the second UV noise map to obtain a second UV feature map. Further, the second UV feature map and the manually processed UV feature map can be compared and calculated for similarity, and then it is determined based on the similarity whether the denoised second UV feature map meets the requirements. Correspondingly, when the second UV feature map meets the requirements, it is determined that the predicted second noise maps of multiple perspectives also meet the requirements, which means that the diffusion model has been trained. Otherwise, the parameter tuning process of the diffusion model is continued until noise maps of multiple perspectives that meet the requirements are predicted.
[0144] Among them, whether the second UV characteristic map obtained by denoising meets the requirements is determined according to the similarity. For example, a similarity threshold can be set. When the similarity is greater than the threshold, it means that the second UV characteristic map meets the requirements. It should be noted that the embodiment of the present application does not specifically limit the method used for comparing and analyzing the second UV characteristic map with the UV characteristic map to be compared. The above method for calculating the similarity is only an example.
[0145] It should be noted that, in the process of adjusting the parameters of the first diffusion model one by one using the remaining training sets, the processing of the obtained analysis results is similar to the processing of the initial diffusion model, which will not be repeated here.
[0146] In this application, the effect and quality of the denoising process can be evaluated through the above comparative analysis, and the analysis results can guide further optimization and adjustment.
[0147] It should be noted that the trained diffusion model may also include a process of extracting a normal map and extracting features from basic information. Optionally, the training process of the diffusion model includes:
[0148] Acquire multiple training sets, wherein each training image set includes basic information corresponding to the three-dimensional model, an initial multi-view feature map, and noise maps of multiple viewpoints after manual processing;
[0149] For one of the training sets, extract the features of the basic information and the normal map of the initial multi-view feature map;
[0150] The initial multi-view feature map, the features of the basic information and the normal map are input into the initial diffusion model to predict the third noise map of multiple viewpoints;
[0151] Analyzing the third noise map to obtain an analysis result, and adjusting the parameters of the initial diffusion model based on the analysis result to obtain a second diffusion model;
[0152] The remaining training sets are used to adjust the parameters of the second diffusion model one by one to obtain a trained diffusion model.
[0153] Optionally, the training process of the diffusion model may also include a process of training to obtain a preset noise image, that is, training of the forward process of the diffusion model. For example, a multi-view feature map corresponding to the three-dimensional model may be obtained, and the diffusion model may be trained through a forward noisy process so that the trained diffusion model can process the preset noise image to obtain a noise map.
[0154] Optionally, the input of the training of the forward process of the diffusion model can also be a clear image of the three-dimensional model. Using this image, the diffusion model is trained through a forward noisy process. The embodiment of the present application does not specifically limit the training process of forward noisy. It can be understood that the training of the forward process is the inverse training of the reverse process.
[0155] For example, any one of the training sets is obtained from multiple training sets and input into the initial diffusion model for prediction to obtain noise from multiple perspectives. Figure 1 , further, for noise from multiple perspectives Figure 1 Analyze and generate analysis results, and adjust the parameters of the initial diffusion model based on the analysis results to obtain a first diffusion model; further, input another training set into the first diffusion model for prediction to obtain noise from multiple perspectives Figure 2 , noise from multiple perspectives Figure 2 Perform analysis, generate analysis results, adjust the parameters of the first diffusion model based on the analysis results, obtain a third diffusion model, obtain a training set, and repeat the above process until a diffusion model that meets the requirements is obtained.
[0156] Among them, whether the diffusion model meets the requirements can be determined by the user, or a noise comparison process can be set, such as comparing the similarity of the predicted noise maps of multiple perspectives with the preset noise, and judging whether it meets the requirements based on the comparison results. The embodiment of the present application does not specifically limit the method for judging whether the diffusion model meets the requirements.
[0157] Therefore, the embodiment of the present application adjusts the model parameters one by one so that the performance of the model is gradually optimized, thereby ensuring that the diffusion model finally obtained has a good prediction effect.
[0158] Optionally, the method further includes:
[0159] Displaying the texture map through a terminal interface, and updating the texture map in response to a user's adjustment operation on the texture map; and / or,
[0160] According to the multi-view feature map rendered by the target UV feature map, images corresponding to the three-dimensional model at the multiple view angles are generated, the images corresponding to the multiple view angles are displayed through a terminal interface, and in response to a user's adjustment operation on the image corresponding to at least one view angle, the target UV feature map and the texture map are updated according to the adjusted image.
[0161] Exemplarily, the texture map is displayed through the terminal interface, and in response to the user's adjustment operation on the image size at each viewing angle in the texture map, the size of the image at each viewing angle is updated, and the positions of the updated images at each viewing angle are arranged to form a target texture map.
[0162] Among them, by rearranging all the images under the updated perspective, it is ensured that they are arranged reasonably on the terminal interface, do not overlap and make full use of the space, and then through reasonable arrangement, a new target texture map can be formed, so that the texture map remains neat and consistent after adjustment.
[0163] In an embodiment of the present application, the generated texture map can be displayed on the terminal interface, and the texture map can be adjusted through comprehensive evaluation and verification by the user to ensure that the generated target texture map is not only visually continuous and seamless, but also highly restores the texture details of the original input image, thereby meeting the needs of high-end modeling and rendering.
[0164] Therefore, in this application, users can freely adjust the size of texture maps as needed to meet different application requirements, and the system can automatically arrange the updated texture maps, reducing the workload of manual adjustments by users and improving work efficiency.
[0165] Optionally, the multi-view feature map obtained by rendering the target UV feature map is restored to the corresponding images under the multiple viewpoints through a decoder, so that the user can adjust the corresponding images under the multiple viewpoints.
[0166] After obtaining the target UV feature map, it is possible that the target UV feature map still cannot meet the needs of the user. Therefore, the multi-view feature map obtained by rendering the target UV feature map can be restored into an image that can be displayed on the terminal interface based on the decoder, and then the user can adjust the image and perform denoising again to obtain a UV feature map that meets the user's needs. Furthermore, based on the UV feature map that meets the user's needs, a high-quality texture map can be generated.
[0167] In this way, images from multiple perspectives can be generated and visualized through the multi-perspective feature map, so that users can adjust the images of the multi-perspective feature map in real time, check the appearance of the model from different perspectives, ensure the image quality from each perspective, and users can personalize the display view according to their personal needs and preferences to obtain more expected results.
[0168] In combination with the above embodiments, Figure 3 A schematic diagram of a flow chart of an optional texture map generation method provided in an embodiment of the present application, such as Figure 3 As shown, taking the basic information as an image and the multi-view feature map as a six-view feature map as an example, the texture map generation method includes:
[0169] Step 1: Input the image, initialize the UV feature map, and extract the hierarchical features of the image.
[0170] Step 2: Extract the normal map corresponding to the initial six-view feature map, and use the diffusion model to predict the noise map of the six-view feature map according to the normal map, the hierarchical features and the initial six-view feature map, and then project the noise map into the UV space to generate a UV noise map, and fuse the UV noise map, and then use the fused UV noise map to denoise the UV feature map, and render the denoised UV feature map back to the six-view feature map, replace the initial six-view feature map with the six-view feature map, and continue to execute step 2. Step 2 can be iteratively executed at least once to achieve iterative denoising.
[0171] Step 3: Generate a UV map of the 3D model based on the UV feature map obtained in the last iteration.
[0172] In this way, in each round of denoising processing in the embodiment of the present application, the noise map is projected into the UV space for fusion processing, and then the UV noise map after the fusion processing is used to denoise the UV feature map, and the denoised UV feature map is used to render a new multi-view for the next round of denoising, and then the UV map of the three-dimensional model is generated based on the UV feature map obtained in the last round. Through the above process, high-quality UV maps with rich details and good visual effects can be generated, which solves the problem of map quality caused by incoordination between views, and greatly improves the efficiency and quality benchmark of three-dimensional content creation.
[0173] In the above embodiments, the texture map generation method provided by the embodiments of the present application is introduced. In order to realize the functions in the method provided by the embodiments of the present application, the electronic device as the execution subject may include a hardware structure and / or a software module, and the above functions are realized in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether one of the above functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.
[0174] For example, Figure 4 A structural diagram of a texture map generation device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the device 400 includes: an acquisition module 401, which is used to acquire basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles;
[0175] A prediction module 402 is used to predict noise maps of multiple perspectives using a trained diffusion model according to the basic information and the initial multi-perspective feature map;
[0176] A denoising module 403 is used to project the noise maps of the multiple viewing angles into a UV space and then perform fusion processing, and denoise the UV feature map according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-view feature map in the UV space;
[0177] The generating module 404 is used to generate a texture map of the three-dimensional model based on the target UV feature map.
[0178] Optionally, the denoising module 403 includes a processing unit and a denoising unit, and the processing unit is used to:
[0179] Projecting the noise maps of the multiple viewing angles into a UV space to obtain a UV noise map;
[0180] A weighted average operation is performed on the pixel values of the pixel points in the UV noise map to obtain a fused noise map.
[0181] Optionally, the denoising unit is used to:
[0182] The UV feature map is subtracted from the fused noise map to obtain a target UV feature map.
[0183] Optionally, the denoising unit is used to:
[0184] De-noising the UV feature map according to the fused noise map to obtain a first UV feature map;
[0185] Rendering the first UV feature map to obtain a first multi-view feature map;
[0186] According to the basic information and the first multi-view feature map, at least one round of denoising processing is performed based on the diffusion model to obtain a target UV feature map, wherein the target UV feature map is the first UV feature map obtained in the last round;
[0187] Among them, the denoising processing of each round includes: according to the characteristics of the basic information and the first multi-view feature map of the previous round, based on the diffusion model, obtaining the first noise maps of multiple viewpoints in the current round, projecting the first noise maps of the multiple viewpoints into the UV space and performing fusion processing, denoising the first UV feature map of the previous round according to the first noise map after the fusion processing to obtain the first UV feature map of the current round, and obtaining the first multi-view feature map of the current round according to the first UV feature map.
[0188] Optionally, the device 400 further includes an extraction module, wherein the extraction module is configured to:
[0189] Extracting a normal map corresponding to the initial multi-view feature map;
[0190] Accordingly, the prediction module 402 is specifically used for:
[0191] The normal map, the features of the basic information and the initial multi-view feature map are input into the diffusion model to predict noise maps of multiple viewpoints.
[0192] Optionally, the basic information is an initial image of the three-dimensional model at any viewing angle, and the device 400 further includes an extraction unit, which is used to:
[0193] The feature extraction is performed on the initial image of any perspective through a feature extraction model to obtain the features of the basic information, wherein the feature extraction model includes a plurality of feature extraction layers connected in sequence, which are used to extract and output features of different levels respectively.
[0194] Optionally, the device 400 further includes a model training module, wherein the model training module is used to:
[0195] Acquire multiple training sets, where each training set includes basic information corresponding to the three-dimensional model, an initial multi-view feature map, and manually processed noise maps of multiple viewpoints;
[0196] For one of the training sets, noise prediction is performed based on the initial diffusion model according to the basic information in the training set and the initial multi-view feature map, and the parameters of the initial diffusion model are adjusted according to the predicted noise maps of multiple views and the noise maps of multiple views in the training set to obtain a first diffusion model;
[0197] The remaining training sets are used to adjust the parameters of the first diffusion model one by one to obtain a trained diffusion model.
[0198] Optionally, the device 400 further includes an adjustment module, wherein the adjustment module is configured to:
[0199] Displaying the texture map through a terminal interface, and updating the texture map in response to a user's adjustment operation on the texture map; and / or,
[0200] According to the multi-view feature map rendered by the target UV feature map, images corresponding to the three-dimensional model at the multiple view angles are generated, the images corresponding to the multiple view angles are displayed through a terminal interface, and in response to a user's adjustment operation on the image corresponding to at least one view angle, the target UV feature map and the texture map are updated according to the adjusted image.
[0201] It should be noted that the specific implementation principle and effects of the above-mentioned texture map generation device can be found in the relevant descriptions and effects corresponding to the above-mentioned embodiments, and will not be elaborated here.
[0202] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device of this embodiment may include: at least one processor 501; and a memory 502 communicatively connected to the at least one processor 501; wherein the memory 502 stores instructions executable by the at least one processor 501, and the instructions are executed by the at least one processor 501 so that the electronic device executes the method described in any of the above embodiments.
[0203] Optionally, the memory 502 may be independent or integrated with the processor 501 .
[0204] The memory 502 and the processor 501 may be connected via a bus 503 .
[0205] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.
[0206] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.
[0207] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the above embodiments when executed by a processor.
[0208] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0209] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.
[0210] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include RAM (Random Access Memory), and may also include NVM (Non-Volatile Memory), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0211] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0212] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0213] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0214] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0215] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0216] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A texture map generation method, characterized in that: The method comprises: Acquire basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles; According to the basic information and the initial multi-view feature map, noise maps of multiple viewpoints are predicted using a trained diffusion model; The noise maps of the multiple viewing angles are projected into the UV space and then fused, and the UV feature map is denoised according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-view feature map in the UV space; A texture map of the three-dimensional model is generated based on the target UV feature map.
2. The method according to claim 1, characterized in that The noise maps of the multiple perspectives are projected into the UV space and then fused, including: Projecting the noise maps of the multiple viewing angles into a UV space to obtain a UV noise map; A weighted average operation is performed on the pixel values of the pixel points in the UV noise map to obtain a fused noise map.
3. The method according to claim 1, characterized in that: The UV feature map is denoised according to the fused noise map to obtain the target UV feature map, including: The UV feature map is subtracted from the fused noise map to obtain a target UV feature map.
4. The method according to claim 1, characterized in that The UV feature map is denoised according to the fused noise map to obtain the target UV feature map, including: De-noising the UV feature map according to the fused noise map to obtain a first UV feature map; Rendering the first UV feature map to obtain a first multi-view feature map; According to the basic information and the first multi-view feature map, at least one round of denoising processing is performed based on the diffusion model to obtain a target UV feature map, wherein the target UV feature map is the first UV feature map obtained in the last round; Among them, the denoising processing of each round includes: according to the characteristics of the basic information and the first multi-view feature map of the previous round, based on the diffusion model, obtaining the first noise maps of multiple viewpoints in the current round, projecting the first noise maps of the multiple viewpoints into the UV space and performing fusion processing, denoising the first UV feature map of the previous round according to the first noise map after the fusion processing to obtain the first UV feature map of the current round, and obtaining the first multi-view feature map of the current round according to the first UV feature map.
5. The method according to claim 1, characterized in that: The method further comprises: Extracting a normal map corresponding to the initial multi-view feature map; The method of predicting noise maps of multiple perspectives using a trained diffusion model according to the basic information and the initial multi-perspective feature map includes: The normal map, the features of the basic information and the initial multi-view feature map are input into the diffusion model to predict noise maps of multiple viewpoints.
6. The method according to claim 5, characterized in that The basic information is an initial image of the three-dimensional model at any viewing angle, and the method further includes: The feature extraction is performed on the initial image of any perspective through a feature extraction model to obtain the features of the basic information, wherein the feature extraction model includes a plurality of feature extraction layers connected in sequence, which are used to extract and output features of different levels respectively.
7. The method according to claim 1, characterized in that The training process of the diffusion model includes: Acquire multiple training sets, where each training set includes basic information corresponding to the three-dimensional model, an initial multi-view feature map, and manually processed noise maps of multiple viewpoints; For one of the training sets, noise prediction is performed based on the initial diffusion model according to the basic information in the training set and the initial multi-view feature map, and the parameters of the initial diffusion model are adjusted according to the predicted noise maps of multiple views and the noise maps of multiple views in the training set to obtain a first diffusion model; The remaining training sets are used to adjust the parameters of the first diffusion model one by one to obtain a trained diffusion model.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Displaying the texture map through a terminal interface, and updating the texture map in response to a user's adjustment operation on the texture map; and / or, According to the multi-view feature map rendered by the target UV feature map, images corresponding to the three-dimensional model at the multiple view angles are generated, the images corresponding to the multiple view angles are displayed through a terminal interface, and in response to a user's adjustment operation on the image corresponding to at least one view angle, the target UV feature map and the texture map are updated according to the adjusted image.
9. A texture map generating device, characterized in that: The device comprises: An acquisition module, used to acquire basic information and an initial multi-view feature map corresponding to the three-dimensional model to be processed, wherein the basic information is used to characterize the texture and / or shape of the three-dimensional model, and the initial multi-view feature map is a preset noise image under multiple viewing angles; A prediction module, configured to predict noise maps of multiple perspectives using a trained diffusion model based on the basic information and the initial multi-perspective feature map; A denoising module is used to project the noise maps of the multiple viewing angles into the UV space and then perform fusion processing, and denoise the UV feature map according to the fused noise map to obtain a target UV feature map, where the UV feature map is a feature map corresponding to the multi-view feature map in the UV space; A generation module is used to generate a texture map of the three-dimensional model based on the target UV feature map.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method according to any one of claims 1 to 8.