Model training method, rendering method and device thereof, and apparatus

By training a generative model and combining it with global information from the initial reference image, the material representation across different viewpoints is coordinated, solving the problems of consistency across multiple views and texture misalignment in 3D model reconstruction, and achieving high-quality 3D model rendering effects.

CN122223199APending Publication Date: 2026-06-16HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
Filing Date
2026-05-15
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In existing 3D model reconstruction technologies, it is difficult to guarantee the consistency of multiple views, and there are insufficient texture misalignment and material standardization output, making it difficult to meet high-quality requirements.

Method used

By acquiring an initial reference image and initial noisy images from multiple perspectives, a generative model is trained using a feature extraction module and a material diffusion module. By combining the global information of the initial reference image, the material representation between different perspectives is coordinated, reducing differences between perspectives, enhancing the consistency of materials across multiple perspectives, and reducing sensitivity to input noise.

Benefits of technology

It improves the consistency and realism of the generated results, enhances the consistency of materials from multiple perspectives, improves the stability and generalization ability of the model, reduces the complexity of training and inference, and improves the quality of 3D model rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223199A_ABST
    Figure CN122223199A_ABST
Patent Text Reader

Abstract

The present disclosure provides a model training method, a rendering image generation method and a device thereof. The method comprises: obtaining a target training sample; the target training sample comprises: an initial reference image, initial noise corresponding to each view angle in N view angles, and M initial material channel images determined based on an initial material of an initial three-dimensional model in each view angle in N view angles; performing fusion processing on the M initial material channel images in each view angle and the initial noise corresponding to each view angle to obtain an initial noisy image of each view angle; inputting the initial reference image and the initial noisy image of each view angle into a to-be-trained generative model to obtain M estimated material channel images in each view angle and estimated noise corresponding to each view angle; obtaining a target loss value based on the estimated material channel image and / or the estimated noise; and performing model training on the to-be-trained generative model using the target loss value to obtain a target generative model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a model training method, a method for generating effect diagrams, and the apparatus and equipment thereof. Background Technology

[0002] With the development of 3D model reconstruction technology, many reconstruction schemes have emerged in the field of 3D model reconstruction, such as implicit volume optimization, explicit volume optimization, and multimodal generation. However, existing reconstruction schemes still have the following problems: the consistency of the generated multiple views is difficult to guarantee, texture misalignment and insufficient material standardization output are difficult to meet the high-quality requirements of 3D model reconstruction. Summary of the Invention

[0003] This disclosure provides a model training method, a method for generating effect diagrams, and related apparatus and devices to solve or alleviate one or more technical problems in the prior art.

[0004] In a first aspect, this disclosure provides a model training method for a generative model used to reconstruct the material of a 3D model, including: Obtain target training samples; the target training samples include: an initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a 2D image of a reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; The M initial material channel images from each viewpoint are fused with the initial noise corresponding to each viewpoint to obtain the initial noisy image from each viewpoint. The initial reference image and the initial noisy images from each viewpoint are input into the generative model to be trained. The feature extraction module in the generative model to be trained extracts features from the initial reference image to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material. The material diffusion module in the generative model to be trained, combined with the features characterizing the geometric information of the reference 3D model and the material information of the initial material, and the initial noisy images from each viewpoint, obtains M estimated material channel images for each viewpoint and the estimated noise for each viewpoint. The target loss value is obtained based on the estimated material channel map and / or estimated noise; Using the target loss value, the feature extraction module and material diffusion module in the generative model to be trained are used to train the model, so as to obtain the target generative model.

[0005] Secondly, this disclosure provides a method for generating material renderings of a 3D model, including: Obtain the mesh data of the target 3D model and the target reference image, wherein the target reference image is a 2D image of the reference 3D model with the target material; the reference 3D model is structurally related to the target 3D model. The mesh data of the target 3D model is preprocessed to obtain noisy target images from each of the N viewpoints; Using a target generative model, and based on the target reference image and the noisy target images from each viewpoint, M predicted material channel images are obtained for each viewpoint; wherein, the target generative model is trained using the model training method described above. Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, a rendering effect map of the target 3D model with the target material is obtained.

[0006] Thirdly, this disclosure provides a model training apparatus for reconstructing a generative model of a 3D model material, comprising: A sample acquisition unit is used to acquire target training samples; the target training samples include: an initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a two-dimensional image of a reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; The model training unit is used to fuse M initial material channel maps from each viewpoint with the initial noise corresponding to each viewpoint to obtain initial noisy images from each viewpoint; input the initial reference map and the initial noisy images from each viewpoint into the generative model to be trained, so as to use the feature extraction module in the generative model to be trained to extract features from the initial reference map to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material; and use the material diffusion module in the generative model to be trained, combined with the features characterizing the geometric information of the reference 3D model and the material information of the initial material, and the initial noisy images from each viewpoint, to obtain M estimated material channel maps from each viewpoint and estimated noise corresponding to each viewpoint; obtain a target loss value based on the estimated material channel maps and / or estimated noise; and use the target loss value to train the feature extraction module and the material diffusion module in the generative model to be trained to obtain the target generative model.

[0007] Fourthly, this disclosure provides a device for generating material renderings of a three-dimensional model, comprising: The data input unit is used to acquire the mesh data of the target 3D model and the target reference image, wherein the target reference image is a 2D image of the reference 3D model with the target material; the reference 3D model is structurally related to the target 3D model. The material generation unit is used to preprocess the mesh data of the target 3D model to obtain the target noisy image of each view in N viewpoints; using the target generative model, and based on the target reference image and the target noisy image of each viewpoint, to obtain M predicted material channel images under each viewpoint; wherein, the target generative model is trained using the model training method described above. The rendering unit is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.

[0008] Fifthly, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0009] In a sixth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.

[0010] In a seventh aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.

[0011] In this disclosed scheme, the initial reference image is used as conditional information, and it, along with the initial noisy images from each of the N viewpoints, is used as input to the generative model to be trained. The feature extraction module learns the relevant features of the geometric and material information in the initial reference image. These learned features representing the geometric and material information are then used as input to the material diffusion module. Combined with the initial noisy images from each viewpoint, material diffusion is performed, resulting in M ​​estimated material channel maps for each viewpoint, along with the estimated noise for each viewpoint. This allows the model to more accurately understand and reconstruct material details, effectively enhancing the consistency and realism of the generated results. Furthermore, because the above learning process incorporates global information from the initial reference image, it helps to coordinate material representations across viewpoints, reducing differences between viewpoints and enhancing the consistency of materials across multiple viewpoints. This solves the problem of texture misalignment caused by independent material generation from each viewpoint, making the subsequently generated materials more realistic and coherent. Furthermore, because the material diffusion module of this disclosed solution incorporates global information from the initial reference image during the diffusion process, it can better guide denoising, reduce sensitivity to input noise, and thus improve the model's stability and generalization ability. Simultaneously, by obtaining global information from the initial reference image through pre-learning via the feature extraction module, redundant calculations during the diffusion process can be effectively avoided, improving overall training and inference efficiency.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.

[0014] Figure 1 This is a schematic flowchart of a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application; Figure 2 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 1 ; Figure 3 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 2 ; Figure 4This is an illustrative diagram illustrating the cross-module coupling connection between the feature extraction module and the material diffusion module according to an embodiment of this application. Figure 1 ; Figure 5 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 3 ; Figure 6 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 4 ; Figure 7 This is an illustrative schematic diagram of a feature extraction module according to an embodiment of this application; Figure 8 This is an illustrative schematic diagram of a material diffusion module according to an embodiment of this application; Figure 9 This is an illustrative diagram illustrating the cross-module coupling connection between the feature extraction module and the material diffusion module according to an embodiment of this application. Figure 2 ; Figure 10 This is an illustrative diagram illustrating the cross-module coupling connection between the feature extraction module and the material diffusion module according to an embodiment of this application. Figure 3 ; Figure 11 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 5 ; Figure 12 This is a schematic flowchart illustrating a method for generating a material effect diagram of a three-dimensional model according to an embodiment of this application; Figure 13 This is a flowchart illustrating a method for generating a material effect diagram of a three-dimensional model according to an embodiment of this application in one example; Figure 14 This is a schematic diagram of the structure of a model training device for reconstructing a generative model of a 3D model material according to an embodiment of this application; Figure 15 This is a schematic diagram of the structure of a three-dimensional model material effect drawing generation device according to an embodiment of this application; Figure 16 This is a block diagram of an electronic device used to implement the model training method for reconstructing a generative model of a three-dimensional model material or the method for generating a material effect map of a three-dimensional model according to the embodiments of this disclosure. Detailed Implementation

[0015] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0016] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0017] With the development of 3D model reconstruction technology, numerous reconstruction schemes have emerged in the field of 3D model reconstruction, such as implicit volume optimization, explicit volume optimization, and multimodal generation. Current 3D reconstruction schemes specifically include: Solution A: Implicit volume optimization solution, the core logic of which is: Taking text or an image as input, this method optimizes the implicit volume (i.e., neural radiance fields (NeRF) or signed distance function (SDF)) through score distillation to fit a 3D scene, and then obtains the target view through volume rendering. This allows for the generation of 3D models from natural language or single images, but this approach has high requirements for computational resources and training time.

[0018] Option B: Explicit structural optimization scheme, the core logic of which is: Optimization is achieved by employing explicit 3D structures (such as meshes (M), point clouds (Point), or Gaussian volumes) and incorporating scoring feedback from a diffusion model to iteratively improve the geometry and texture of the 3D structure, thereby directly generating editable 3D models. This improves later-stage usability, but the training stability and consistency of the approach remain limited.

[0019] Option C: A fast explicit representation scheme, whose core logic is as follows: This approach enables efficient reconstruction and real-time rendering by rapidly fitting lightweight structures (such as point clouds or Gaussian volumes) through multi-view supervision. Prioritizing speed and scalability, it is suitable for low-latency scenarios, such as real-time augmented reality (AR) or virtual reality (VR) modeling.

[0020] The above-mentioned 3D model reconstruction scheme still has significant limitations and cannot meet the high-quality requirements of 3D model reconstruction. Specifically, these limitations are as follows: (a) Long model training and inference time: When iterating and optimizing a single object, the 3D generation scheme based on score distillation sampling (SDS) has a long optimization time and is difficult to generate quickly.

[0021] (b) Multi-view Figure 1 Inconsistency issues: Multiple views generated from a single image or text are prone to inconsistency, leading to problems such as reconstruction tearing and texture misalignment.

[0022] (c) Semantic and style alignment is unstable: under textual or weak image conditions, detail style is prone to deviate from structural semantics.

[0023] (d) Insufficient standardized output: The generated 3D model lacks the complete Level of Detail (LOD), U-axis-V-axis mapping and material information required for cross-engine or pipeline docking, resulting in high implementation costs.

[0024] Based on this, the present disclosure provides a model training method for a generative model to reconstruct the material of a 3D model. This method uses an initial reference image, initial noisy images of each of the N viewpoints, and the initial noise corresponding to each viewpoint to train the feature extraction module and the material diffusion module in the generative model to be trained, thereby obtaining a target generative model that can directly output M estimated material channel maps under each viewpoint. The generative model obtained by the above training can standardize the output of multiple material channel maps under each viewpoint. Here, because the scheme fully learns the geometric and material information related features of the initial reference image with initial material during the training phase using the feature extraction module, and fully learns the related features of the M initial material channel maps corresponding to the initial material under different viewpoints using the material diffusion module, and combines the above two types of features for model training, it effectively solves the problem of texture misalignment caused by independent generation of materials from each viewpoint and enhances the consistency of materials from multiple viewpoints. Furthermore, because this scheme uses the initial reference image as a constraint, when faced with non-ideal inputs, such as mesh data of a 3D model, the model can still combine the structural and material prior information of the 3D model in the reference image to predict the distribution of materials on the 3D model. This effectively improves the model's generalization ability and lays the foundation for obtaining high-quality renderings of the reconstructed 3D model using the target generative model.

[0025] Specifically, Figure 1 This is a schematic flowchart illustrating a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application. This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0026] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes: Step S101: Obtain the target training sample.

[0027] Here, the target training samples include: an initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images (M is an integer greater than 1) determined based on the initial material in each of the N (N is an integer greater than 1) viewpoints of the initial 3D model.

[0028] Furthermore, the initial reference image is a two-dimensional image of a reference 3D model with the initial material, and the structure of the reference 3D model in the two-dimensional image is related to that of the initial 3D model. For example, in one example, the structure of the reference 3D model in the initial reference image is similar to that of the initial 3D model, or, further, the structure of the reference 3D model in the initial reference image is the same as that of the initial 3D model. This provides prior structural information about the 3D model for subsequent model diffusion, thereby enabling more thorough learning of material features from different perspectives.

[0029] It should be noted that the "material channel map" referred to in this disclosure refers to image data formed by breaking down the composite material properties of a 3D model surface into independent individual material properties. Each material channel map stores data corresponding to one material property, such as grayscale images, color images, or other data formats suitable for representing that material property. Furthermore, each material channel map corresponds to different material channels, such as primary color / diffuse reflection, normal, roughness, and metallicity channels in a physically based rendering (PBR) workflow, thus providing input data for the rendering calculations of the corresponding channels.

[0030] In other words, in one example, the different initial material channel maps in the M initial material channel maps of this disclosed scheme correspond to different material properties. In other words, different initial material channel maps carry different types of material property information. This provides rich data support for the model to fully learn the relationship between viewpoint, geometric information, and material information.

[0031] In one example, the M material properties include, but are not limited to: diffuse reflection, refraction, reflection, opacity, base color, metallicity, roughness, etc., and this disclosure does not specifically limit them.

[0032] Step S102: Fuse the M initial material channel images under each viewpoint with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.

[0033] It should be noted that, in one example, different viewpoints can correspond to a single initial noise. That is, the initial noise corresponding to each viewpoint mentioned above can be specifically understood as a single initial noise corresponding to each viewpoint. In other words, different initial material channel maps under the same viewpoint correspond to the same initial noise. In this case, N viewpoints correspond to N initial noises. Furthermore, it should be noted that the initial noise corresponding to different viewpoints can be the same or different. For example, some viewpoints may correspond to the same initial noise, or other viewpoints may correspond to different initial noises. This disclosure does not impose specific restrictions on this.

[0034] Alternatively, in another example, the initial noise corresponding to different material channel maps under the same viewpoint is different; or, some material channel maps under the same viewpoint correspond to the same initial noise, while other material channel maps correspond to another initial noise; or, the initial noise corresponding to different material channel maps under the same viewpoint is different from each other. In this case, for the same viewpoint, there are M initial noises, and further, for N viewpoints, there are N×M initial noises.

[0035] It should be noted that the above examples of initial noise are merely illustrative. In practical applications, the noise corresponding to the viewpoint can be determined based on the actual scenario requirements and actual training accuracy requirements. This disclosure does not impose any specific restrictions on this.

[0036] In addition, it should be noted that this disclosure does not specifically limit the method of generating the initial noise.

[0037] Furthermore, in one example, fusing the M initial material channel maps from each viewpoint with the single initial noise corresponding to each viewpoint can specifically include: For the initial material channel map to be processed under the current viewpoint, the initial noise corresponding to the current initial material channel map is determined, and the current initial material channel map and the determined initial noise are fused (e.g., through linear interpolation). This yields the initial noisy image corresponding to the current initial material channel map under the current viewpoint. Following this logic, for M initial material channel maps under the current viewpoint, M initial noisy images corresponding to the current viewpoint can be obtained. Furthermore, for N viewpoints, N×M initial noisy images can be obtained.

[0038] Step S103: Input the initial reference image and the initial noisy images of each viewpoint into the generative model to be trained, so as to use the feature extraction module in the generative model to extract features from the initial reference image to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material; and use the material diffusion module in the generative model to be trained, and combine the features characterizing the geometric information of the reference 3D model and the material information of the initial material, as well as the initial noisy images of each viewpoint, to obtain M estimated material channel maps under each viewpoint, and the estimated noise corresponding to each viewpoint.

[0039] Here, in one example, the above-described method of utilizing the material diffusion module in the generative model to be trained, combined with the geometric information representing the reference 3D model and the material information of the initial material, as well as the initial noisy images of each viewpoint, to obtain M estimated material channel maps and the estimated noise corresponding to each viewpoint, can specifically include: Using the material diffusion module in the generative model to be trained, and combining the geometric information representing the reference 3D model and the material information of the initial material, as well as the initial noisy images of each viewpoint, the estimated noise corresponding to the initial noisy image under each viewpoint is first estimated. For example, for a scene with M initial noisy images under each of N viewpoints, N×M estimated noises can be estimated. Then, using the estimated noise corresponding to the initial noisy image under each viewpoint, the initial noisy image corresponding to the estimated noise is denoised. In this way, M estimated material channel images under each viewpoint are obtained, for a total of N×M estimated material channel images.

[0040] In other words, in this example, the material diffusion module is first used to estimate the estimated noise corresponding to the initial noisy image under each viewpoint. Then, based on the estimated noise, the initial noisy image corresponding to the estimated noise is denoised to obtain the estimated material channel map corresponding to the initial noisy image. In this way, N×M estimated material channel maps are obtained.

[0041] Step S104: Obtain the target loss value based on the estimated material channel map and / or estimated noise.

[0042] Step S105: Using the target loss value, train the feature extraction module and material diffusion module in the generative model to be trained to obtain the target generative model.

[0043] In other words, the proposed solution first inputs the initial reference image and the initial noisy images from each viewpoint into the generative model to be trained. After being processed by the feature extraction module and the material diffusion module, M estimated material channel images and the estimated noise corresponding to each viewpoint are obtained. Then, the target loss value is obtained by using the M estimated material channel images and / or the estimated noise corresponding to each viewpoint. Finally, the parameters of the feature extraction module and the material diffusion module in the generative model to be trained are fine-tuned using the obtained target loss value to obtain the target generative model.

[0044] In this disclosed scheme, the initial reference image is used as conditional information, and it, along with the initial noisy images from each of the N viewpoints, is used as input to the generative model to be trained. The feature extraction module learns the relevant features of the geometric and material information in the initial reference image. These learned features representing the geometric and material information are then used as input to the material diffusion module. Combined with the initial noisy images from each viewpoint, material diffusion is performed, resulting in M ​​estimated material channel maps for each viewpoint, along with the estimated noise for each viewpoint. This allows the model to more accurately understand and reconstruct material details, effectively enhancing the consistency and realism of the generated results. Furthermore, because the above learning process incorporates global information from the initial reference image, it helps to coordinate material representations across viewpoints, reducing differences between viewpoints and enhancing the consistency of materials across multiple viewpoints. This solves the problem of texture misalignment caused by independent material generation from each viewpoint, making the subsequently generated materials more realistic and coherent. Furthermore, because the material diffusion module of this disclosed solution incorporates global information from the initial reference image during the diffusion process, it can better guide denoising, reduce sensitivity to input noise, and thus improve the model's stability and generalization ability. Simultaneously, by obtaining global information from the initial reference image through pre-learning via the feature extraction module, redundant calculations during the diffusion process can be effectively avoided, improving overall training and inference efficiency.

[0045] Furthermore, the disclosed solution can use M estimated material channel maps from each viewpoint and / or the estimated noise corresponding to each viewpoint to calculate the target loss value, so as to train the generative model to be trained, thereby obtaining a target generative model that can standardize and output material channel maps from multiple viewpoints. The low training complexity of the above training solution reduces the time and computing resources required for model training, which provides favorable support for efficient 3D reconstruction using the target generative model.

[0046] Moreover, since the generative model trained by this disclosed solution can output multiple material channel maps from different perspectives, such as outputting N×M estimated material channel maps, compared with the existing solution that uses 4 material channel maps for 3D reconstruction, this disclosed solution can significantly increase the number of material channel dimensions, thus enriching the dimensions of material attribute representation and effectively improving the quality of materials generated by the trained model.

[0047] Furthermore, since the initial reference image is used as a constraint in this disclosed solution, the model can still predict the distribution of materials on the initial 3D model by combining the structural and material priors of the reference 3D model in the reference image when facing non-ideal inputs. Therefore, this disclosed solution can fully combine the detailed style of materials with the structural semantics of the 3D model, thereby making the M material channel maps generated by the model from various perspectives more accurate, that is, improving the generalization ability of the model. This lays the foundation for subsequent 3D reconstruction using the target generative model and obtaining high-quality renderings of 3D models with materials.

[0048] Furthermore, in a specific example, the target loss value can be obtained in the following manner; specifically, the above-described method of obtaining the target loss value based on the estimated material channel map and / or estimated noise (e.g., step S104) can specifically include: Method 1: Based on M estimated material channel maps and M initial material channel maps from each viewpoint, the target loss value is obtained.

[0049] In other words, under this method 1, the M initial material channel maps (i.e., N×M initial material channel maps) from each viewpoint are used as label data, and the resulting N×M estimated material channel maps are used as the estimated values ​​output by the model. Then, the target loss value is obtained based on the difference information between the two.

[0050] Method 2: Obtain the target loss value based on the estimated noise and the initial noise corresponding to each viewpoint.

[0051] In other words, under this method 2, the initial noise (e.g., N×M initial noises) under each viewpoint is used as the label data, and the obtained estimated noise (e.g., N×M estimated noises) is used as the estimated value output by the model. Then, the target loss value is obtained based on the difference information between the two.

[0052] Method 3: Based on the loss value between the M estimated material channel maps and the M initial material channel maps under each viewpoint, and the loss value between the estimated noise and the initial noise corresponding to each viewpoint, the target loss value is obtained.

[0053] In other words, the first loss value can be obtained using method 1 described above, and the second loss value can be obtained using method 2 described above; then, the target loss value can be obtained by using the first and second loss values.

[0054] It should be noted that in practical applications, an appropriate loss calculation method can be selected according to actual needs (such as accuracy requirements). This disclosed solution does not impose specific restrictions on which loss calculation method to use.

[0055] Furthermore, the present disclosure can also obtain the target loss value (referred to as Method 4) based on M initial material channel maps from each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image from each viewpoint, specifically including: Step 1: Fuse the M initial material channel images from each viewpoint with the estimated noise corresponding to each viewpoint (e.g., linear interpolation) to obtain the estimated noisy image for each viewpoint.

[0056] For details on the linear interpolation process in step 1, please refer to the example above; it will not be repeated here.

[0057] Step 2: Based on the estimated noisy images from each viewpoint and the initial noisy images from each viewpoint, obtain the target loss value.

[0058] In other words, the estimated noise corresponding to each viewpoint obtained from the generative model to be trained is linearly interpolated with the M material channel images under each viewpoint to obtain the estimated noisy image of each viewpoint; then, using a loss function, such as the mean squared error loss function, and combining the image vector of the estimated noisy image of each viewpoint with the image vector of the initial noisy image of each viewpoint, the difference information between the two image vectors is calculated to obtain the target loss value.

[0059] It should be noted that, in calculating the target loss value, this disclosed solution may use method 4, any one of methods 1 and 2 mentioned above, or any two of methods 1, 2, and 4 mentioned above (such as method 1 and method 2 (i.e., method 3), or method 1 and method 4, or method 2 and method 4), or simultaneously use methods 1, 2, and 4 mentioned above to calculate the target loss value. This disclosed solution does not impose specific restrictions on which method is used to calculate the target loss value.

[0060] Thus, this disclosed solution provides a specific method for calculating the target loss value. This method is simple, practical, and highly interpretable, thereby giving the model training process great flexibility. It enables the model training to select the most suitable loss calculation method according to the specific needs of the scenario, thereby effectively avoiding the optimization bias or overfitting risk that may be caused by a single loss, and providing favorable support for subsequent training to obtain the target generative model.

[0061] Furthermore, in a specific example, the feature extraction module and the material diffusion module form a dual-branch structure; a cross-module coupling connection is provided between the feature extraction module and the material diffusion module. This cross-module coupling connection is used to transmit the features extracted by the feature extraction module to the material diffusion module, thereby constraining or modulating the diffusion process of the material diffusion module. This facilitates the integration of global information from the initial reference image extracted by the feature extraction module during the diffusion process, thus better guiding denoising, reducing sensitivity to input noise, and improving the model's stability and generalization ability.

[0062] In other words, in this example, the generative model to be trained contains two branches: a feature extraction module and a material diffusion module. The two are coupled together so that the global information extracted by the feature extraction module can be transmitted to the material diffusion module to constrain or modulate the diffusion process.

[0063] For example, such as Figure 2 As shown, the generative model to be trained includes a feature extraction module and a material diffusion module, and the feature extraction module and the material diffusion module form a dual-branch structure. For example, the feature extraction module and the material diffusion module are set up in parallel. In this case, the initial reference image of the "chair" with the initial material can be input into the feature extraction module to extract global information containing geometric and material information. At the same time, the initial noisy images of each of the N viewpoints are input into the material diffusion module, and combined with the global information containing geometric and material information input from the feature extraction module, the material diffusion process is constrained or modulated, thereby obtaining the estimated noise corresponding to each viewpoint and the M estimated material channel images corresponding to each viewpoint.

[0064] In this way, the proposed solution sets up the feature extraction module and the material diffusion module in parallel and connects them across modules. This allows the features (also known as global information) output by the feature extraction module, which contain geometric and material information, to be shared with the material diffusion module. This breaks down the information silos between modules and realizes information sharing between modules. As a result, the diffusion process of the model is given more accurate guidance or prior information, which provides rich feature information for the subsequent training of the generative model to be trained.

[0065] Furthermore, in a specific example, the feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks; that is, in this example, the feature extraction module and the material diffusion module correspond to each other in network structure design and have similar forms, belonging to two instances of the same network architecture.

[0066] For example, in one instance, the feature extraction module includes a first encoding stage and a first decoding stage in series; correspondingly, the material diffusion module includes a second encoding stage and a second decoding stage in series.

[0067] For example, continue with Figure 2 Taking the model structure shown as an example, as Figure 3 As shown, both the feature extraction module and the material diffusion module are networks containing encoding and decoding stages. For example, the feature extraction module includes a first encoding stage and a first decoding stage connected in series. Here, the first encoding stage is used to extract features from the initial reference image of the "chair" with the initial material, and the first decoding stage is used to recover (or decode) the features output by the first encoding stage. Further, as... Figure 3 As shown, the material diffusion module includes a second encoding stage and a second decoding stage connected in series. At this time, the second encoding stage is used to: extract features from the initial noisy images of each of the N viewpoints, or to extract features from the initial noisy images of each of the N viewpoints in combination with the features input by the feature extraction module. Correspondingly, the second decoding stage is used to: restore the features output by the second encoding stage, or to restore the features output by the second encoding stage in combination with the features input by the feature extraction module.

[0068] Furthermore, in one example, the feature extraction module and the material diffusion module in the generative model to be trained are structurally symmetrical isomorphic networks, which can specifically mean that the network structures of the first encoding stage and the second encoding stage are isomorphic, and / or the network structures of the first decoding stage and the second decoding stage are isomorphic.

[0069] In other words, in one example, the feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks, which can be specifically understood as: the network structures of the first encoding stage and the second encoding stage are isomorphic; or, in another example, the feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks, which can also be specifically understood as the network structures of the first decoding stage and the second decoding stage are isomorphic; or, in yet another example, the feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks, which can also be specifically understood as the network structures of the first encoding stage and the second encoding stage are isomorphic, as are the network structures of the first decoding stage and the second decoding stage.

[0070] Here, the "network structure isomorphism" referred to in this example means that neural networks are similar or identical in topological structure, such as the type and number of network layers, the arrangement of neurons, and the parameter configuration. In short, it means that they have similar or identical network skeletons.

[0071] Furthermore, in one example, where the feature extraction module includes a first encoding stage and a first decoding stage in series, and the material diffusion module includes a second encoding stage and a second decoding stage in series, the cross-module coupling connection includes at least one of the following connection methods: Connection Method 1: A cross-module coupling connection is established between the first encoding stage and the second encoding stage. For example, in one example, the features output by the first encoding stage are transmitted to the second encoding stage.

[0072] Connection Method 2: A cross-module coupling connection is established between the first decoding stage and the second decoding stage. For example, in one example, the features output by the first decoding stage are transmitted to the second decoding stage.

[0073] It should be noted that in practical applications, cross-module coupling connections may include any of the above connection methods, or both, and this disclosure does not impose any specific restrictions on this.

[0074] Continue with Figure 3 Taking the model structure shown as an example, in one example, such as Figure 4 As shown, the cross-module coupling connection between the feature extraction module and the material diffusion module can specifically include: setting a cross-module coupling connection between the first encoding stage in the feature extraction module and the second encoding stage in the material diffusion module; for example, the features output by the first encoding stage are transmitted to the second encoding stage (i.e., arrow ① in the figure). In other words, this example allows the intermediate features of the encoding stage of the feature extraction module to directly act on the encoding process of the material diffusion module. This makes the texture generated by the diffusion stage more consistent with the input conditions in structure, while improving the detail quality and training efficiency, and preventing the diffusion model from deviating from the expected texture structure when it "freely plays".

[0075] In another example, such as Figure 4 As shown, the cross-module coupling connection between the feature extraction module and the material diffusion module can specifically include: setting a cross-module coupling connection between the first decoding stage and the second decoding stage, for example, transmitting the features output by the first decoding stage to the second decoding stage (i.e., arrow ② in the figure). In other words, this example allows the features of the decoding stage of the feature extraction module to directly act on the decoding stage of the material diffusion module, realizing strong conditional guidance of structural information on material generation. In this way, the stability of the underlying geometry and contour is maintained, avoiding the destruction of the overall structure by detail diffusion, and the material synthesis can be adaptively refined based on accurate structural features, improving the structure-texture consistency, visual realism, and stability and efficiency of multi-stage training of the generated results, thereby significantly improving the quality and fidelity of material generation.

[0076] Furthermore, in another example, where the feature extraction module includes a first encoding stage and a first decoding stage connected in series, the connection method may include the following: Connection method 3: A cross-stage coupling connection is established between the first encoding stage and the first decoding stage. For example, in one example, the features output by the first encoding stage are transmitted to the first decoding stage.

[0077] Furthermore, in yet another example, where the material diffusion module includes a second encoding stage and a second decoding stage connected in series, the following connection method may also be included: Connection method 4: A cross-stage coupling connection established between the second encoding stage and the second decoding stage. For example, in one example, the features output by the second encoding stage are transmitted to the second decoding stage.

[0078] It should be noted here that, in one example, connection methods 3 and 4 can be arbitrarily combined with connection methods 1 and 2 mentioned above; for example, continuing with... Figure 3 Taking the model structure shown as an example, in another example, such as Figure 4 As shown, a cross-stage coupling connection is set between the first encoding stage and the first decoding stage, and a cross-module coupling connection is set between the first decoding stage and the second decoding stage. For example, the features output by the first encoding stage are transmitted to the first decoding stage (i.e., arrow ③ in the figure), and the features output by the first decoding stage are transmitted to the second decoding stage (i.e., arrow ② in the figure). In this way, setting a cross-stage coupling connection within the feature extraction module (i.e., the features of the first encoding stage are transmitted to the first decoding stage) helps to enhance the feature reuse and information integrity of the feature extraction module itself, thereby improving the accuracy of structure reconstruction. At the same time, a cross-module coupling connection is set between the feature extraction module and the material diffusion module (i.e., the features of the first decoding stage are transmitted to the second decoding stage). In this way, the optimized structural features can be directly used as conditions to input the material generation process, thereby ensuring that the material details are accurately aligned with the underlying structure, significantly improving the geometric-texture consistency, visual realism, and stability and efficiency of multi-stage collaborative training of the overall generation result.

[0079] Or, in yet another example, such as Figure 4As shown, a cross-stage coupling connection is set between the second encoding stage and the second decoding stage, and a cross-module coupling connection is set between the first decoding stage and the second decoding stage. For example, the features output by the second encoding stage are transmitted to the second decoding stage (i.e., arrow ④ in the figure), and the features output by the first decoding stage are transmitted to the second decoding stage (i.e., arrow ② in the figure). This improves the material generation capability. Thus, setting a cross-stage coupling connection within the material diffusion module (i.e., transmitting features from the second encoding stage to the second decoding stage) enhances the multi-scale feature fusion capability of the material diffusion module itself, improving the quality of material detail synthesis. Simultaneously, setting a cross-module coupling connection in the decoding stage between the feature extraction module and the material diffusion module (i.e., transmitting features from the first decoding stage to the second decoding stage) allows structural features to directly and stably guide the material generation process, ensuring accurate alignment between the material and the underlying geometry, and allowing detail synthesis to fully utilize multi-level information. This results in more realistic, structurally consistent, and high-quality generation results, and promotes the convergence and efficiency of the collaborative training of the two modules.

[0080] It should be noted that the coupling connections between the feature extraction module and the material diffusion module in the generative model to be trained are only an example. In practical applications, reasonable settings should be made according to the needs of the actual scenario, and this disclosure does not impose specific restrictions on them. For example, in one example, cross-stage coupling connections are set between the first encoding stage and the first decoding stage, and between the second encoding stage and the second decoding stage. At the same time, cross-module coupling connections are set between the first encoding stage and the second encoding stage, and between the first decoding stage and the second decoding stage.

[0081] In this way, the proposed solution breaks down the information silos of the dual-branch processing paths by setting cross-stage coupling connections and / or cross-module coupling connections between different stages, and realizes deep interaction and collaboration of features. This improves the robustness of feature representation, and enables the two modules to complement each other during training, thereby significantly improving the model's feature decoupling ability and material generation ability in material generation tasks.

[0082] Furthermore, in a specific example, the feature extraction module further includes a first bottleneck stage located between the first encoding stage and the first decoding stage. This first bottleneck stage is used to process the features output by the first encoding stage to obtain global prior features; and the material diffusion module further includes a second bottleneck stage located between the second encoding stage and the second decoding stage. This second bottleneck stage is used to process the features output by the second encoding stage to obtain global semantic features.

[0083] For example, continue with Figure 3 Taking the model structure shown as an example, as Figure 5 As shown, a first bottleneck stage is also set between the first encoding stage and the first decoding stage in the feature extraction module. For example, in one example, after inputting the initial reference image of the "chair" with the initial material into the first encoding stage and obtaining the features output by the first encoding stage, the features output by the first encoding stage can be input into the first bottleneck stage to obtain the global prior features for the initial reference image of the "chair". Further, the obtained global prior features are input into the first decoding stage for recovery processing.

[0084] Furthermore, a second bottleneck stage is set between the second encoding stage and the second decoding stage in the material diffusion module. For example, in one example, after inputting the features of the initial noisy images of each viewpoint and the initial reference image of the "chair" transmitted by the first encoding stage into the second encoding stage and obtaining the features output by the second encoding stage, the features output by the second encoding stage can also be input into the second bottleneck stage to obtain global semantic features. Further, the obtained global semantic features (or the obtained global semantic features and the features transmitted by the first decoding stage) are input into the second decoding stage for restoration processing.

[0085] Furthermore, in one example, the feature extraction module and the material diffusion module in the generative model to be trained are structurally symmetrical isomorphic networks, which may specifically include: the network structures of the first bottleneck stage and the second bottleneck stage are isomorphic.

[0086] Furthermore, in one example, the cross-module coupling connection may further include: a bottleneck stage cross-module coupling connection set between the first bottleneck stage and the second bottleneck stage. For example, in one example, features output from the first bottleneck stage are transmitted to the second bottleneck stage.

[0087] In other words, in one example, the cross-module coupling connection includes at least one of the following connection methods: Connection method 1: A cross-module coupling connection for the encoding stages is set between the first encoding stage and the second encoding stage.

[0088] Connection method 2: A cross-module coupling connection for the decoding stage is set between the first decoding stage and the second decoding stage.

[0089] Connection method 5: A bottleneck stage cross-module coupling connection is set between the first bottleneck stage and the second bottleneck stage.

[0090] It should be noted that in practical applications, cross-module coupling connections may include any of the above connection methods, or any two of the three, or all three. This disclosure does not impose any specific restrictions on this.

[0091] It should be noted that, in one example, the connection methods 1, 2 and 5 described above can also be combined with the connection methods 3 and 4 described above in any way.

[0092] For example, continue with Figure 5 Taking the model structure shown as an example, as Figure 6 As shown, when a first bottleneck stage is set between the first encoding stage and the first decoding stage, and a second bottleneck stage is set between the second encoding stage and the second decoding stage, the cross-module coupling connection set between the feature extraction module and the material diffusion module also includes: a bottleneck stage cross-module coupling connection set between the first bottleneck stage and the second bottleneck stage, and a decoding stage cross-module coupling connection set between the first decoding stage and the second decoding stage.

[0093] For example, for the feature extraction module, the initial reference image of the "chair" with the initial material is input into the first encoding stage to obtain the features output by the first encoding stage, such as the first encoded features. Further, the obtained first encoded features are input into the first bottleneck stage to obtain the global prior features for the initial reference image of the "chair". Then, the obtained global prior features are input into the first decoding stage to recover the first decoded features that represent the fusion of geometric and material information. Here, the obtained global prior features can be further transmitted to the second bottleneck stage through the cross-module coupling connection of the bottleneck stage (i.e., arrow ⑤ in the figure); the obtained first decoded features are input to the second decoding stage through the cross-module coupling connection of the decoding stage (i.e., arrow ② in the figure). Furthermore, for the material diffusion module, the initial noisy images from each viewpoint are input into the second encoding stage to obtain the second encoding features. The obtained second encoding features and the global prior features transmitted in the first bottleneck stage are then input into the second bottleneck stage to obtain the global semantic features. Further, the obtained global semantic features and the first decoding features transmitted in the first decoding stage are input into the second decoding stage to obtain the estimated noise corresponding to each viewpoint and the M estimated material channel maps corresponding to each viewpoint.

[0094] It should be noted that the above is only an example. In actual applications, connection method 5 can be combined with at least one of the connection methods 1 to 4. This disclosure does not impose specific restrictions on the specific combination methods.

[0095] In this way, the proposed solution sets bottleneck stages in both the feature extraction module and the material diffusion module, and achieves "hard sharing" of features between modules by constructing cross-module coupling connections for the bottleneck stages. Thus, "hard sharing" injects strong prior constraints into the material generation of the material diffusion module, effectively avoiding the deviation of local details during material diffusion, thereby significantly improving the reconstruction fidelity of the material, and laying the foundation for subsequent training of a target generative model that can output M material channel maps from various perspectives in a standardized manner.

[0096] Furthermore, in a specific example, the first encoding stage in the feature extraction module may specifically include multiple cascaded first encoding units (i.e., the output of the previous first encoding unit serves as the input of the next first encoding unit), which are used to perform multi-scale feature extraction on the input data. Furthermore, the first decoding stage in the feature extraction module may also specifically include multiple cascaded first decoding units (i.e., the output of the previous first decoding unit serves as the input of the next first decoding unit), which are used to perform step-by-step feature recovery processing.

[0097] Here, in one example, multiple cascaded first coding units can perform multi-scale extraction on the input initial reference map. In this way, high-frequency macroscopic spatial structure and microscopic material information in the initial reference map can be effectively captured, improving the model's ability to decouple features of the 3D model and materials.

[0098] For example, in one example, the feature extraction scale of some first coding units in multiple cascaded first coding units is different from that of other first coding units. In this way, the problem of information loss and over-smoothing in deep networks is effectively alleviated, and more feature information with different scales is provided for high-fidelity material generation.

[0099] Alternatively, in another example, the feature extraction scales of different first coding units are different, which can effectively capture more detailed local features in the image, such as low-frequency lighting information, providing rich and complementary underlying features for subsequent material generation.

[0100] Furthermore, in one example, multiple cascaded first decoding units can recover the features of the input step by step, thus accurately and coherently restoring the complex physical details in the initial reference diagram.

[0101] For example, in one example, the feature recovery scale of some first decoding units in multiple cascaded first decoding units is different from the feature recovery scale of other first decoding units. In this way, redundant calculations in feature decoding are effectively reduced and the computational efficiency of the feature decoding process is improved.

[0102] Alternatively, in another example, the feature recovery scales of different first decoding units are different, thus enabling lossless decoding of multi-dimensional material properties and improving the richness of feature information.

[0103] Furthermore, in one example, the cross-stage coupling connection established between the first encoding stage and the first decoding stage can specifically include the cross-unit coupling connection between the first encoding unit and the first decoding unit. That is, the cross-stage coupling connection established between the first encoding stage and the first decoding stage can be implemented through the cross-unit coupling connection between the first encoding unit and the first decoding unit.

[0104] It should be noted that, in one example, the cross-stage coupling connection set between the first encoding stage and the first decoding stage may include one or more cross-unit coupling connections.

[0105] Furthermore, in one example, the multiple cross-unit coupling connections included in the cross-stage coupling connection set between the first encoding stage and the first decoding stage can be implemented through the cross-unit coupling connection set between the first encoding unit and the first decoding unit.

[0106] For example, in one example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the first encoding stage and the first decoding stage can be implemented through cross-unit coupling connections between the same first encoding unit and the first decoding units in multiple different first decoding units; or, in another example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the first encoding stage and the first decoding stage can also be implemented through cross-unit coupling connections between the first encoding units in multiple different first encoding units and the same first decoding unit; or, in yet another example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the first encoding stage and the first decoding stage can also be implemented through cross-unit coupling connections between the first encoding units in multiple first encoding units and the first decoding units corresponding to the first encoding units in multiple first decoding units.

[0107] For example, such as Figure 7As shown, the first encoding stage of the feature extraction module in the generative model to be trained may include four first encoding units in series, and the feature extraction scale of each first encoding unit is different from that of the others; and the first decoding stage of the feature extraction module may include four first decoding units in series, and the feature recovery scale of each first decoding unit is different from that of the others; further, in one example, the cross-stage coupling connection set between the first encoding stage and the first decoding stage includes: a cross-unit coupling connection set between the first first encoding unit of the four first encoding units and the second first decoding unit of the four first decoding units, and a cross-unit coupling connection set between the second first encoding unit of the four first encoding units and the first first decoding unit of the four first decoding units.

[0108] In one example, an initial reference image of a "chair" with initial material can be input into the feature extraction module. Specifically, for the first encoding stage, this initial reference image is input into the first encoding unit to obtain the features output by the first encoding unit. The features output by the first encoding unit are then input into the second encoding unit to obtain the features output by the second encoding unit. Alternatively, the features output by the first encoding unit can be transmitted to the second decoding unit in the first decoding stage via a cross-unit coupling connection. Further, the features output by the second encoding unit are input into the third encoding unit to obtain the features output by the third encoding unit. Again, the features output by the second encoding unit can be transmitted to the first decoding unit in the first decoding stage via a cross-unit coupling connection. Finally, the features output by the third encoding unit are input into the fourth encoding unit to obtain the features output by the fourth encoding unit (i.e., the first encoded features).

[0109] Furthermore, for the first decoding stage, the features transmitted by the second first coding unit (i.e., the features output by the second first coding unit) and the output of the last first coding unit (e.g., the fourth first coding unit) in the first coding stage can be used as the input to the first first decoding unit in the first decoding stage. That is, the features transmitted by the second first coding unit (i.e., the features output by the second first coding unit) and the features output by the fourth first coding unit are input to the first first decoding unit in the first decoding stage to obtain the features output by the first first decoding unit. Further, the features output by the first first decoding unit and the features transmitted by the first first coding unit (i.e., the features output by the first first coding unit) are input to the second first decoding unit to obtain the features output by the second first decoding unit. Further, the features output by the second first decoding unit are input to the third first decoding unit to obtain the features output by the third first decoding unit. Finally, the features output by the third first decoding unit are input to the fourth first decoding unit to obtain the first decoded features.

[0110] Here, the number of first encoding units, the number of first decoding units, and the number of cross-unit coupling connections included in the cross-stage coupling connections set between the first encoding stage and the first decoding stage in the above example are only illustrative examples. In actual applications, they can be reasonably set according to specific scenario requirements. This disclosure does not impose any specific restrictions on them.

[0111] Furthermore, in one example, the first encoding unit in the first encoding stage and the first decoding unit in the first decoding stage can be specifically set symmetrically. In this case, the feature processing scales of the first encoding unit and the first decoding unit with symmetrical relationship are the same.

[0112] Furthermore, in one example, the multiple cross-unit coupling connections include multiple skip connections, such as cross-unit coupling connections between each of the multiple first coding units and the first decoding unit that is symmetrical to it in the multiple first decoding units. In other words, each first coding unit is provided with a cross-unit coupling connection between itself and the first decoding unit that has a symmetrical relationship with it.

[0113] Continue with Figure 7Taking the model structure shown as an example, the cross-stage coupling connection set between the first encoding stage and the first decoding stage can specifically include: the cross-unit coupling connection set between the first first encoding unit in the first encoding stage and the first decoding unit (such as the fourth first decoding unit) that has a symmetrical relationship with the first first encoding unit in the first decoding stage; and the cross-unit coupling connection set between the second first encoding unit in the first encoding stage and the first decoding unit (such as the third first decoding unit) that has a symmetrical relationship with the second first encoding unit in the first decoding stage.

[0114] Here, the number of cross-unit coupling connections included in the cross-stage coupling connection set between the first encoding stage and the first decoding stage in the above example is only an example. In actual applications, it can be reasonably set according to the specific scenario requirements. This disclosure does not impose any specific restrictions on this.

[0115] Thus, the present disclosure provides a specific structure for the first encoding stage and the first decoding stage. Specifically, the first encoding stage includes multiple cascaded first encoding units, and the first decoding stage includes multiple cascaded first decoding units. This endows the model with the ability to decouple features from the 3D model and the initial material, thereby effectively capturing the macroscopic distribution patterns and microscopic details of the material in the initial reference image. This effectively avoids the problem of detail loss at a single scale, thus providing richer feature information for subsequent training of the generative model to be trained.

[0116] Furthermore, in a specific example, the second encoding stage in the material diffusion module may include multiple cascaded second encoding units (i.e., the output of the previous second encoding unit serves as the input of the next second encoding unit), which are used to perform multi-scale feature extraction on the input data. Further, the second decoding stage in the material diffusion module may specifically include multiple cascaded second decoding units (i.e., the output of the previous second decoding unit serves as the input of the next second decoding unit), which are used to perform step-by-step feature recovery processing.

[0117] Here, in one example, multiple cascaded second coding units can perform multi-scale extraction on the initial noisy images of each of the N input viewpoints. In this way, deep feature representations rich in semantic information can be effectively captured, providing feature information for subsequent diffusion using materials.

[0118] For example, in one example, the feature extraction scale of some second coding units in multiple cascaded second coding units is different from that of other second coding units. In this way, the problem of information loss and over-smoothing in deep networks is effectively alleviated, and feature information at different scales is provided for material diffusion.

[0119] Alternatively, in another example, the feature extraction scales of different second coding units are different, which can effectively capture more detailed local features in the image, such as the geometric contours of the initial 3D model, providing a basis for subsequent material diffusion.

[0120] Furthermore, in one example, multiple cascaded second decoding units can recover the input features step by step, thus accurately and coherently restoring the geometric information of the initial 3D model and the material information of the initial material.

[0121] For example, in one example, the feature recovery scale of some of the second decoding units in multiple cascaded second decoding units is different from the feature recovery scale of other second decoding units. In this way, redundant calculations in feature decoding are effectively reduced and the computational efficiency of the feature decoding process is improved.

[0122] Alternatively, in another example, the feature recovery scales of different second decoding units are different, thus achieving refined feature reconstruction "from macroscopic outline to microscopic detail", thereby providing a basis for subsequent accurate estimation of noise.

[0123] Furthermore, in one example, the cross-stage coupling connection established between the second encoding stage and the second decoding stage can specifically include the cross-unit coupling connection between the second encoding unit and the second decoding unit. That is, the cross-stage coupling connection established between the second encoding stage and the second decoding stage can be implemented through the cross-unit coupling connection between the second encoding unit and the second decoding unit.

[0124] It should be noted that, in one example, the cross-stage coupling connection set between the second encoding stage and the second decoding stage may include one or more cross-unit coupling connections.

[0125] Furthermore, in one example, the multiple cross-unit coupling connections included in the cross-stage coupling connection set between the second encoding stage and the second decoding stage can be implemented through the cross-unit coupling connection set between the second encoding unit and the second decoding unit.

[0126] For example, in one example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the second encoding stage and the second decoding stage can be implemented through cross-unit coupling connections between the same second encoding unit and the second decoding units in multiple different second decoding units; or, in another example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the second encoding stage and the second decoding stage can be implemented through cross-unit coupling connections between the second encoding units in multiple different second encoding units and the same second decoding unit; or, in yet another example, the multiple cross-unit coupling connections included in the cross-stage coupling connection between the second encoding stage and the second decoding stage can also be implemented through cross-unit coupling connections between the second encoding units in multiple second encoding units and the second decoding units corresponding to the second encoding units in multiple second decoding units.

[0127] For example, such as Figure 8 As shown, the second encoding stage of the material diffusion module in the generative model to be trained may include four second encoding units in series, and the feature extraction scale of each second encoding unit is different; and the second decoding stage of the material diffusion module may include four second decoding units in series, and the feature recovery scale of each second decoding unit is different; further, in one example, the cross-stage coupling connection set between the second encoding stage and the second decoding stage includes: a cross-unit coupling connection set between the first second encoding unit of the four second encoding units and the first second decoding unit of the four second decoding units, and a cross-unit coupling connection set between the second second encoding unit of the four second encoding units and the fourth second decoding unit of the four second decoding units.

[0128] In one example, the initial noisy images of each of the N viewpoints, along with the features transmitted by the feature extraction module, can be input to the material diffusion module. Specifically, for the second encoding stage, the initial noisy images of each viewpoint (or, the initial noisy images of each viewpoint and the features transmitted by the feature extraction module) are input to the first second encoding unit in the second encoding stage to obtain the features output by the first second encoding unit. Then, the features output by the first second encoding unit (or, the features output by the first second encoding unit and the features transmitted by the feature extraction module) are input to the second second encoding unit to obtain the features output by the second second encoding unit. Alternatively, the features output by the first second encoding unit can be transmitted to the first second decoding unit in the second decoding stage via a cross-unit coupling connection. Further, the features output by the second second coding unit (or the features output by the second second coding unit and the features transmitted by the feature extraction module) are input to the third second coding unit to obtain the features output by the third second coding unit. Here, the features output by the second second coding unit can also be transmitted to the fourth second decoding unit in the second decoding stage via a cross-unit coupling connection. Further, the features output by the third second coding unit (or the features output by the third second coding unit and the features transmitted by the feature extraction module) are input to the fourth second coding unit to obtain the features output by the fourth second coding unit (i.e., the second coding features).

[0129] Furthermore, for the second decoding stage, the features transmitted by the first second coding unit (i.e., the features output by the first second coding unit) and the features output by the last second coding unit (e.g., the fourth second coding unit) in the first coding stage (or, the features transmitted by the first second coding unit, the features output by the last second coding unit in the first coding stage, and the features transmitted by the feature extraction module) can be used as input to the first second decoding unit in the second decoding stage. In other words, the features transmitted by the first second coding unit (i.e., the features output by the first second coding unit) and the features output by the fourth second coding unit (or, the features transmitted by the first second coding unit, the features output by the fourth second coding unit, and the features transmitted by the feature extraction module) are input to the first second decoding unit in the second decoding stage to obtain the features output by the first second decoding unit. Step by step, the features output by the first second decoding unit (or the features output by the first second decoding unit and the features transmitted by the feature extraction module) are input to the second second decoding unit to obtain the features output by the second second decoding unit. Further, the features output by the second second decoding unit (or the features output by the second second decoding unit and the features transmitted by the feature extraction module) are input to the third second decoding unit to obtain the features output by the third second decoding unit. Finally, the features transmitted by the second second encoding unit (i.e., the features output by the second second encoding unit) and the features output by the third second decoding unit (or the features output by the second second encoding unit, the features output by the third second decoding unit, and the features transmitted by the feature extraction module) are input to the fourth second decoding unit to obtain M estimated material channel maps under each viewpoint, as well as the estimated noise corresponding to each viewpoint.

[0130] Here, the number of second encoding units, the number of second decoding units, and the number of cross-unit coupling connections included in the cross-stage coupling connections set between the second encoding stage and the second decoding stage in the above example are only illustrative examples. In actual applications, they can be reasonably set according to specific scenario requirements. This disclosure does not impose any specific restrictions on them.

[0131] Furthermore, in one example, the second encoding unit in the second encoding stage and the second decoding unit in the second decoding stage can also be specifically set symmetrically. In this case, the feature processing scale of the second encoding unit and the second decoding unit with symmetrical relationship is the same.

[0132] Furthermore, in one example, the multiple cross-unit coupling connections include multiple skip connections, such as cross-unit coupling connections between each of the multiple second coding units and the second decoding unit that is symmetrical to it among the multiple second decoding units. In other words, each second coding unit has a cross-unit coupling connection between it and its symmetrically related second decoding unit.

[0133] Continue with Figure 8 Taking the model structure shown as an example, the cross-stage coupling connection set between the second encoding stage and the second decoding stage may also include: the cross-unit coupling connection set between the first second encoding unit in the second encoding stage and the second decoding unit (such as the fourth second decoding unit) that has a symmetrical relationship with the first second encoding unit in the second decoding stage; and the cross-unit coupling connection set between the second second encoding unit in the second encoding stage and the second decoding unit (such as the third second decoding unit) that has a symmetrical relationship with the second second encoding unit in the second decoding stage.

[0134] Here, the number of cross-unit coupling connections included in the cross-stage coupling connection set between the second encoding stage and the second decoding stage in the above example is only an example. In actual applications, it can be reasonably set according to the specific scenario requirements. This disclosure does not impose any specific restrictions on this.

[0135] Thus, this disclosure provides a specific structure for the second encoding stage and the second decoding stage. Specifically, the second encoding stage includes multiple cascaded second encoding units, and the second decoding stage includes multiple cascaded second decoding units. This effectively captures deep feature representations rich in semantic information in noisy images, providing a basis for the material diffusion module in the subsequent generative model to be trained to have diffusion processing capabilities.

[0136] Furthermore, in one example, the cross-module coupling connection between the first and second encoding stages can specifically include a cross-unit coupling connection between the first and second encoding units. In other words, the cross-module coupling connection between the first and second encoding stages can be implemented through a cross-unit coupling connection between the first and second encoding units.

[0137] It should be noted that, in one example, the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding stage may include one or more cross-unit coupling connections.

[0138] Furthermore, in one example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding stage can be implemented through the cross-unit coupling connection set between the first encoding unit and the second encoding unit.

[0139] For example, in one example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding stage can be implemented through cross-unit coupling connections between the same first encoding unit and the second encoding units in multiple different second encoding units; or, in another example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding stage can also be implemented through cross-unit coupling connections between the first encoding units in multiple different first encoding units and the same second encoding unit; or, in yet another example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding unit in multiple second encoding units can also be implemented through cross-unit coupling connections between the first encoding units in multiple first encoding units and the second encoding units corresponding to the first encoding units in multiple second encoding units.

[0140] For example, such as Figure 9 As shown, the cross-module coupling connections between the first and second encoding stages include: a cross-unit coupling connection between the first first encoding unit among the four first encoding units and the first second encoding unit among the four second encoding units, and a cross-unit coupling connection between the second first encoding unit among the four first encoding units and the third second encoding unit among the four second encoding units. Further, in one example, the first encoding units in the first encoding stage and the second encoding units in the second encoding stage can also be symmetrically arranged. In this case, the feature processing scales of the symmetrically related first and second encoding units are the same.

[0141] Furthermore, in one example, the multiple cross-unit coupling connections include cross-unit coupling connections between each of the multiple first coding units and the second coding unit symmetrical to it among the multiple second coding units. In other words, each first coding unit has a cross-unit coupling connection between it and its symmetrical second coding unit.

[0142] Continue with Figure 9Taking the model structure shown as an example, the cross-module coupling connection between the first encoding stage and the second encoding stage may also include: a cross-unit coupling connection between the first first encoding unit in the first encoding stage and a second encoding unit (such as the first second encoding unit) in the second encoding stage that has a symmetrical relationship with the first first encoding unit; and a cross-unit coupling connection between the second first encoding unit in the first encoding stage and a second encoding unit (such as the second second encoding unit) in the second encoding stage that has a symmetrical relationship with the second first encoding unit.

[0143] Here, the number of cross-unit coupling connections included in the cross-module coupling connection of the encoding stage set between the first encoding stage and the second encoding stage in the above example is only an example. In actual applications, it can be reasonably set according to the specific scenario requirements. This disclosure does not impose any specific restrictions on this.

[0144] Furthermore, in one example, the cross-module coupling connection between the first decoding stage and the second decoding stage can specifically include a cross-unit coupling connection between the first decoding unit and the second decoding unit. That is, the cross-module coupling connection between the first decoding stage and the second decoding stage can be implemented through a cross-unit coupling connection between the first decoding unit and the second decoding unit.

[0145] It should be noted that, in one example, the cross-module coupling connection of the decoding stage set between the first decoding stage and the second decoding stage may include one or more cross-unit coupling connections.

[0146] Furthermore, in one example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the decoding stage set between the first decoding stage and the second decoding stage can be implemented through the cross-unit coupling connection set between the first decoding unit and the second decoding unit.

[0147] For example, in one example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the decoding stage between the first decoding stage and the second decoding stage can be implemented through cross-unit coupling connections between the same first decoding unit and the second decoding units in multiple different second decoding units; or, in another example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the decoding stage between the first decoding stage and the second decoding stage can also be implemented through cross-unit coupling connections between the first decoding units in multiple different first decoding units and the same second decoding unit; or, in yet another example, the multiple cross-unit coupling connections included in the cross-module coupling connection of the decoding stage between the first decoding stage and the second decoding stage can also be implemented through cross-unit coupling connections between the first decoding units in multiple first decoding units and the second decoding units in multiple second decoding units corresponding to the first decoding units.

[0148] For example, such as Figure 10 As shown, the cross-module coupling connections between the first and second decoding stages include: a cross-unit coupling connection between the second of the four first decoding units and the second of the four second decoding units, and a cross-unit coupling connection between the third of the four first decoding units and the fourth of the four second decoding units. Further, in one example, the first decoding units in the first decoding stage and the second decoding units in the second decoding stage can also be symmetrically arranged. In this case, the feature processing scales of the symmetrically related second decoding units and the second decoding units are the same.

[0149] Furthermore, in one example, the multiple cross-unit coupling connections include cross-unit coupling connections between each of the multiple first decoding units and a second decoding unit that is symmetrical to it among the multiple second decoding units. In other words, each first decoding unit is provided with a cross-unit coupling connection between itself and a second decoding unit that has a symmetrical relationship with it.

[0150] Continue with Figure 10 Taking the model structure shown as an example, the cross-module coupling connection between the first decoding stage and the second decoding stage may also include: a cross-unit coupling connection between the first first decoding unit in the first decoding stage and a second decoding unit (such as the first second decoding unit) in the second decoding stage that has a symmetrical relationship with the first first decoding unit; and a cross-unit coupling connection between the second first decoding unit in the first decoding stage and a second decoding unit (such as the second second decoding unit) in the second decoding stage that has a symmetrical relationship with the second first decoding unit.

[0151] Here, the number of cross-unit coupling connections included in the cross-module coupling connection of the decoding stage set between the first decoding stage and the second decoding stage in the above example is only an example. In actual applications, it can be reasonably set according to the specific scenario requirements. This disclosure does not impose any specific restrictions on this.

[0152] Figure 11 This is a model structure diagram of the generative model to be trained in this disclosed scheme, such as... Figure 11 As shown, the generative model to be trained includes a feature extraction module and a material diffusion module, and the two modules are structurally symmetrical isomorphic networks. The feature extraction module includes a cascaded first encoding stage and a first decoding stage, and the material diffusion module includes a cascaded second encoding stage and a second decoding stage. Further, the first encoding stage of the feature extraction module contains four cascaded first encoding units, and the first decoding stage contains four cascaded first decoding units, with each first encoding unit and each first decoding unit being symmetrically arranged. Similarly, the second encoding stage of the material diffusion module contains four cascaded second encoding units, and the second decoding stage contains four cascaded second decoding units, with each second encoding unit and each second decoding unit also being symmetrically arranged. Furthermore, each first encoding unit and each second encoding unit are symmetrically arranged, and each first decoding unit and each second decoding unit are also symmetrically arranged.

[0153] In one example, for the first encoding and first decoding stages in the feature extraction module, each first encoding unit has a cross-unit coupling connection with a first decoding unit that has a symmetrical relationship with it. For example, the first first encoding unit among the four first encoding units has a cross-unit coupling connection with the first decoding unit that has a symmetrical relationship with the first first encoding unit (i.e., the fourth first decoding unit), so as to transmit the features output by the first first encoding unit to the first decoding unit that has a symmetrical relationship with the first first encoding unit. Similarly, for the second encoding and second decoding stages in the material diffusion module, each second encoding unit has a cross-unit coupling connection with a second decoding unit that has a symmetrical relationship with it. For example, the first second encoding unit among the four second encoding units has a cross-unit coupling connection with the second decoding unit that has a symmetrical relationship with the first second encoding unit (i.e., the fourth second decoding unit), so as to transmit the features output by the first second encoding unit to the second decoding unit that has a symmetrical relationship with the first second encoding unit. The first second encoding unit has a symmetrical relationship with the second decoding unit. For the encoding stage and decoding stage of each module, each first encoding unit is connected to a second encoding unit with which it has a symmetrical relationship. For example, the first first encoding unit among the four first encoding units is connected to a second encoding unit with which it has a symmetrical relationship (i.e., the first second encoding unit) to transmit the features output by the first first encoding unit to the second encoding unit corresponding to the first first encoding unit. Similarly, each first decoding unit is connected to a second decoding unit with which it has a symmetrical relationship. For example, the first first decoding unit among the four first decoding units is connected to a second decoding unit with which it has a symmetrical relationship (i.e., the first second decoding unit) to transmit the features output by the first first decoding unit to the second decoding unit corresponding to the first first decoding unit.

[0154] Figure 12 This is a schematic flowchart illustrating a method for generating material effect diagrams of a 3D model according to an embodiment of this application. This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0155] Furthermore, the method includes at least a portion of the following: For example... Figure 12 As shown, it includes: Step S1201: Obtain the mesh data of the target 3D model and the target reference image.

[0156] Here, the target reference image is a two-dimensional image of a reference three-dimensional model with the target material; the reference three-dimensional model is structurally related to the target three-dimensional model.

[0157] It should be noted that the relevant content regarding the target reference diagram can be referred to the example of the initial reference diagram mentioned above, and will not be repeated here.

[0158] Step S1202: Preprocess the mesh data of the target 3D model to obtain noisy target images from each of the N viewpoints.

[0159] Step S1203: Using the target generative model, and based on the target reference map and the target noisy images from each viewpoint, obtain M predicted material channel maps for each viewpoint.

[0160] Here, the target generative model is trained using the model training method of any of the above embodiments.

[0161] It should be noted that the relevant content regarding material channel maps can be found in the example above, and will not be repeated here.

[0162] Step S1204: Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, obtain the rendering effect map of the target 3D model with the target material.

[0163] In other words, in one example, the target reference image and the target noisy image from each viewpoint are input into the trained target generative model to predict M predicted material channel images from each viewpoint. Then, using the M predicted material channel images from each viewpoint and the mesh data of the target 3D model, the rendering effect image of the target 3D model with the target material can be obtained.

[0164] In this way, the disclosed solution utilizes a pre-trained target generative model, taking the target reference image as a priori and combining it with noisy target images from various perspectives, to predict M predicted material channel maps for each perspective, thereby obtaining a rendered image of the target 3D model using the target material. This effectively enhances the consistency of materials across multiple perspectives, solves the texture misalignment problem caused by independent material generation from each perspective, and improves the detail realism and overall coherence of the materials distributed on the target 3D model, resulting in a high-quality rendered image of the target 3D model with the target material.

[0165] Furthermore, since this disclosed solution can utilize the target generative model to standardize the output of multiple material channel maps required for material generation, compared with the existing solution that uses 4 material channel maps for reconstruction, this disclosed solution can output more material channel maps, thereby improving the representation accuracy of material properties. In this way, the quality of materials distributed on the target 3D model is further improved, thereby improving the quality of the rendered image of the reconstructed target 3D model.

[0166] Furthermore, in a specific example, preprocessing can be performed as follows; specifically, the preprocessing of the mesh data of the target 3D model described above to obtain noisy target images from each of the N viewpoints (e.g., step S1202) can specifically include: Step S1202-1: Based on the mesh data of the target 3D model, obtain the normal map and position map of the target 3D model in each of the N viewpoints.

[0167] In this example, the normal image (also known as a normal map) is used to simulate the bumps and shadows of the surface of the target 3D model; the position image (also known as a position map) is used to store the position information of each point on the surface of the target 3D model in 3D space.

[0168] Step S1202-2: Fusion processing is performed on the normal images, position images, and random noise corresponding to each viewpoint to obtain the target noisy image of each viewpoint.

[0169] For example, in one example, the normal images and position images under each viewpoint are fused together, such as the normal image and position image under the same viewpoint are fused together to obtain the fused image under each viewpoint; then, the random noise corresponding to each viewpoint is used to add noise to the fused image under each viewpoint to obtain the target noisy image under each viewpoint.

[0170] Thus, the present disclosure provides a preprocessing scheme for mesh data of a target 3D model. This scheme is simple, practical and highly interpretable. Moreover, on the one hand, it achieves high-dimensional complementarity and strong coupling of the inherent geometric properties of the target 3D model surface (such as local normal direction and global spatial coordinates). On the other hand, by introducing random noise corresponding to each viewpoint, it provides a generalized input for the target generative model that conforms to the distribution of real complex physical imaging, thereby ensuring that the subsequent prediction process can stably and accurately restore high-quality material information containing complete 3D geometric priors.

[0171] Further, in a specific example, the rendering effect of the target 3D model with the target material can be obtained in the following manner; specifically, the above-mentioned method of obtaining the rendering effect of the target 3D model with the target material based on the mesh data of the target 3D model and the M predicted material channel maps under each viewpoint (for example, step S1204) can specifically include: Step S1204-1: Map the M predicted material channel maps from each viewpoint from the two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps.

[0172] For example, in one example, the M predicted material channel maps from each viewpoint are mapped from a two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model. For example, the M predicted material channel maps from each viewpoint are mapped to the mesh occupied by the target three-dimensional model in the three-dimensional space to obtain the three-dimensional material data corresponding to the target three-dimensional model. Then, the three-dimensional material data corresponding to the target three-dimensional model is transformed from the three-dimensional space to the two-dimensional space to obtain M material unfolded maps.

[0173] Step S1204-2: Using the M material unfolded diagrams and the mesh data of the target 3D model, perform rendering processing to obtain a rendering effect diagram of the target 3D model with the target material.

[0174] For example, in one example, the renderer is invoked, and the M material unfolded maps and the mesh data of the target 3D model are combined to perform rendering processing, resulting in a rendered image of the target 3D model with the target material.

[0175] Thus, the present invention provides a rendering scheme that uses M predicted material channel maps from various perspectives to obtain the rendered image. In this way, material transfer for the target 3D model is achieved, and a high-quality rendered image of the 3D model with the target material is reconstructed.

[0176] Figure 13 This is a schematic diagram generated from the material rendering of the 3D model of the disclosed solution, such as... Figure 13 As shown, firstly, the white model data of the target 3D model (corresponding to the mesh data of the target 3D model mentioned above) is preprocessed to obtain normal images and position images of each viewpoint in N viewpoints. Then, the normal images, position images, and corresponding random noise of each viewpoint are fused to obtain the target noisy image of each viewpoint. Secondly, the reference image of the "chair" with the target material (corresponding to the target reference image mentioned above) and the target noisy images of each viewpoint are input into the target generative model to obtain N×M predicted material channel maps. Finally, through a baking algorithm, combined with the white model data of the target 3D model and the N×M predicted material channel maps, M material unfolded maps are obtained. Then, the renderer is called, and combined with the white model data of the target 3D model and the obtained M material unfolded maps, the rendering process is performed to obtain the rendered image of the "chair" with the target material. In this way, the disclosed solution effectively enhances the consistency of materials in multiple viewpoints, thereby improving the quality of materials distributed on the target 3D model, and thus improving the quality of the rendered image of the reconstructed target 3D model.

[0177] This disclosure provides a model training device for reconstructing generative models of 3D model materials, such as... Figure 14 As shown, it includes: The sample acquisition unit 1401 is used to acquire target training samples; the target training samples include: an initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; The model training unit 1402 is used to fuse M initial material channel maps from each viewpoint with the initial noise corresponding to each viewpoint to obtain initial noisy images from each viewpoint; input the initial reference map and the initial noisy images from each viewpoint into the generative model to be trained, so as to use the feature extraction module in the generative model to be trained to extract features from the initial reference map to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material; and use the material diffusion module in the generative model to be trained, combined with the features characterizing the geometric information of the reference 3D model and the material information of the initial material, and the initial noisy images from each viewpoint, to obtain M estimated material channel maps from each viewpoint and estimated noise corresponding to each viewpoint; obtain a target loss value based on the estimated material channel maps and / or estimated noise; and use the target loss value to train the feature extraction module and the material diffusion module in the generative model to be trained to obtain a target generative model.

[0178] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Based on M estimated material channel maps from each viewpoint and M initial material channel maps from each viewpoint, the target loss value is obtained; and / or, The target loss value is obtained based on the estimated noise and the initial noise corresponding to each viewpoint.

[0179] In a specific example of the scheme disclosed herein, the feature extraction module and the material diffusion module form a dual-branch structure; a cross-module coupling connection is provided between the feature extraction module and the material diffusion module for transmitting the features extracted by the feature extraction module to the material diffusion module, so as to constrain or modulate the diffusion process of the material diffusion module.

[0180] In a specific example of the scheme disclosed herein, the feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks; wherein, the feature extraction module includes a first encoding stage and a first decoding stage in series; the material diffusion module includes a second encoding stage and a second decoding stage in series; The cross-module coupling connection includes at least one of the following connection methods: A cross-module coupling connection for the encoding stages is set between the first encoding stage and the second encoding stage; A cross-module coupling connection for the decoding stage is set between the first decoding stage and the second decoding stage.

[0181] In a specific example of the scheme disclosed herein, the feature extraction module further includes: a first bottleneck stage located between the first encoding stage and the first decoding stage, used to process the features output by the first encoding stage to obtain global prior features; The material diffusion module further includes a second bottleneck stage located between the second encoding stage and the second decoding stage, used to process the features output by the second encoding stage to obtain global semantic features.

[0182] In a specific example of the scheme disclosed herein, the cross-module coupling connection may further include: A bottleneck stage cross-module coupling connection is set between the first bottleneck stage and the second bottleneck stage.

[0183] In a specific example of the scheme disclosed herein, the first encoding stage includes multiple cascaded first encoding units for multi-scale feature extraction of the input data; the first decoding stage includes multiple cascaded first decoding units for progressive feature recovery processing. The cross-stage coupling connection between the first encoding stage and the first decoding stage includes the cross-unit coupling connection between the first encoding unit and the first decoding unit.

[0184] In a specific example of the scheme disclosed herein, the second encoding stage includes multiple cascaded second encoding units for multi-scale feature extraction of the input data; the second decoding stage includes multiple cascaded second decoding units for progressive feature recovery processing. The cross-stage coupling connection between the second encoding stage and the second decoding stage includes the cross-unit coupling connection between the second encoding unit and the second decoding unit.

[0185] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0186] This disclosure also provides a device for generating material renderings of 3D models, such as... Figure 15 As shown, it includes: The data input unit 1501 is used to acquire the mesh data of the target 3D model and the target reference image, wherein the target reference image is a 2D image of the reference 3D model with the target material; the reference 3D model is structurally related to the target 3D model. The material generation unit 1502 is used to preprocess the mesh data of the target 3D model to obtain the target noisy image of each view in N viewpoints; using the target generative model, and based on the target reference image and the target noisy image of each viewpoint, to obtain M predicted material channel images under each viewpoint; wherein, the target generative model is trained using the model training method of any of the above embodiments. The rendering unit 1503 is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.

[0187] In a specific example of the disclosed solution, the material generation unit is specifically used for: Based on the mesh data of the target 3D model, the normal image of the target 3D model under each viewpoint in N viewpoints and the position image under each viewpoint in N viewpoints are obtained; The normal images, position images, and random noise corresponding to each viewpoint are fused together to obtain noisy target images for each viewpoint.

[0188] In a specific example of the scheme disclosed herein, the rendering unit is specifically used for: M predicted material channels from each viewpoint are mapped from a two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps. Using the unfolded diagrams of the M materials and the mesh data of the target 3D model, rendering processing is performed to obtain a rendered image of the target 3D model with the target materials.

[0189] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0190] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0191] Figure 16 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 16As shown, the electronic device includes a memory 1610 and a processor 1620. The memory 1610 stores a computer program that can run on the processor 1620. The number of memories 1610 and processors 1620 can be one or more. The memory 1610 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in the above-described method embodiments. The electronic device may also include a communication interface 1630 for communicating with external devices and performing data exchange and transmission.

[0192] If the memory 1610, processor 1620, and communication interface 1630 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 16 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0193] Optionally, in a specific implementation, if the memory 1610, processor 1620 and communication interface 1630 are integrated on a single chip, the memory 1610, processor 1620 and communication interface 1630 can communicate with each other through an internal interface.

[0194] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0195] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).

[0196] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.

[0197] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0198] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0199] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0200] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0201] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.

Claims

1. A model training method for a generative model used to reconstruct the material of a 3D model, comprising: Obtain the target training samples; The target training samples include: an initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; The M initial material channel images from each viewpoint are fused with the initial noise corresponding to each viewpoint to obtain the initial noisy image from each viewpoint. The initial reference image and the initial noisy images from each viewpoint are input into the generative model to be trained. The feature extraction module in the generative model to be trained extracts features from the initial reference image to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material. The material diffusion module in the generative model to be trained, combined with the features characterizing the geometric information of the reference 3D model and the material information of the initial material, and the initial noisy images from each viewpoint, obtains M estimated material channel images for each viewpoint and the estimated noise for each viewpoint. The target loss value is obtained based on the estimated material channel map and / or estimated noise; Using the target loss value, the feature extraction module and material diffusion module in the generative model to be trained are used to train the model, so as to obtain the target generative model.

2. The method according to claim 1, wherein, The process of obtaining the target loss value based on the estimated material channel map and / or estimated noise includes: Based on M estimated material channel maps from each viewpoint and M initial material channel maps from each viewpoint, the target loss value is obtained; and / or, The target loss value is obtained based on the estimated noise and the initial noise corresponding to each viewpoint.

3. The method according to claim 1 or 2, wherein, The feature extraction module and the material diffusion module form a dual-branch structure; a cross-module coupling connection is provided between the feature extraction module and the material diffusion module to transmit the features extracted by the feature extraction module to the material diffusion module, so as to constrain or modulate the diffusion process of the material diffusion module.

4. The method according to claim 3, wherein, The feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks; wherein, the feature extraction module includes a first encoding stage and a first decoding stage in series; the material diffusion module includes a second encoding stage and a second decoding stage in series. The cross-module coupling connection includes at least one of the following connection methods: A cross-module coupling connection for the encoding stages is set between the first encoding stage and the second encoding stage; A cross-module coupling connection for the decoding stage is set between the first decoding stage and the second decoding stage.

5. The method according to claim 4, wherein, The feature extraction module further includes: a first bottleneck stage located between the first encoding stage and the first decoding stage, used to process the features output by the first encoding stage to obtain global prior features; The material diffusion module further includes a second bottleneck stage located between the second encoding stage and the second decoding stage, used to process the features output by the second encoding stage to obtain global semantic features.

6. The method according to claim 5, wherein, The cross-module coupling connection may also include: A bottleneck stage cross-module coupling connection is set between the first bottleneck stage and the second bottleneck stage.

7. The method according to claim 4, wherein, The first encoding stage includes multiple cascaded first encoding units for multi-scale feature extraction of the input data; the first decoding stage includes multiple cascaded first decoding units for progressive feature recovery processing. The cross-stage coupling connection between the first encoding stage and the first decoding stage includes the cross-unit coupling connection between the first encoding unit and the first decoding unit.

8. The method according to claim 4, wherein, The second encoding stage includes multiple cascaded second encoding units for multi-scale feature extraction of the input data; the second decoding stage includes multiple cascaded second decoding units for step-by-step feature recovery processing. The cross-stage coupling connection between the second encoding stage and the second decoding stage includes the cross-unit coupling connection between the second encoding unit and the second decoding unit.

9. A method for generating material renderings of a 3D model, comprising: Obtain the mesh data of the target 3D model and the target reference image, wherein the target reference image is a 2D image of the reference 3D model with the target material; The reference 3D model is structurally related to the target 3D model; The mesh data of the target 3D model is preprocessed to obtain noisy target images from each of the N viewpoints; Using a target generative model, and based on the target reference image and the target noisy images from each viewpoint, M predicted material channel images are obtained from each viewpoint; wherein, the target generative model is trained using the model training method described in claims 1 to 8. Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, a rendering effect map of the target 3D model with the target material is obtained.

10. The method according to claim 9, wherein, The preprocessing of the mesh data of the target 3D model to obtain noisy target images from N viewpoints includes: Based on the mesh data of the target 3D model, the normal image of the target 3D model under each viewpoint in N viewpoints and the position image under each viewpoint in N viewpoints are obtained; The normal images, position images, and random noise corresponding to each viewpoint are fused together to obtain noisy target images for each viewpoint.

11. The method according to claim 9 or 10, wherein, The rendering effect of the target 3D model with the target material is obtained by using the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, including: M predicted material channels from each viewpoint are mapped from a two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps. Using the unfolded diagrams of the M materials and the mesh data of the target 3D model, rendering processing is performed to obtain a rendered image of the target 3D model with the target materials.

12. A model training device for reconstructing generative models of 3D model materials, comprising: The sample acquisition unit is used to acquire target training samples; The target training samples include: an initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; The model training unit is used to fuse M initial material channel maps from each viewpoint with the initial noise corresponding to each viewpoint to obtain initial noisy images from each viewpoint; input the initial reference map and the initial noisy images from each viewpoint into the generative model to be trained, so as to use the feature extraction module in the generative model to be trained to extract features from the initial reference map to obtain features that can characterize the geometric information of the reference 3D model and the material information of the initial material; and use the material diffusion module in the generative model to be trained, combined with the features characterizing the geometric information of the reference 3D model and the material information of the initial material, and the initial noisy images from each viewpoint, to obtain M estimated material channel maps from each viewpoint and estimated noise corresponding to each viewpoint; obtain a target loss value based on the estimated material channel maps and / or estimated noise; and use the target loss value to train the feature extraction module and the material diffusion module in the generative model to be trained to obtain the target generative model.

13. The apparatus according to claim 12, wherein, The feature extraction module and the material diffusion module form a dual-branch structure; a cross-module coupling connection is provided between the feature extraction module and the material diffusion module to transmit the features extracted by the feature extraction module to the material diffusion module, so as to constrain or modulate the diffusion process of the material diffusion module.

14. The apparatus according to claim 13, wherein, The feature extraction module and the material diffusion module are structurally symmetrical isomorphic networks; wherein, the feature extraction module includes a first encoding stage and a first decoding stage in series; the material diffusion module includes a second encoding stage and a second decoding stage in series. The cross-module coupling connection includes at least one of the following connection methods: A cross-module coupling connection for the encoding stages is set between the first encoding stage and the second encoding stage; A cross-module coupling connection for the decoding stage is set between the first decoding stage and the second decoding stage.

15. A device for generating material renderings of a three-dimensional model, comprising: The data input unit is used to acquire the mesh data of the target 3D model and the target reference image, wherein the target reference image is a 2D image of the reference 3D model with the target material; The reference 3D model is structurally related to the target 3D model; The material generation unit is used to preprocess the mesh data of the target 3D model to obtain the target noisy image of each view in N viewpoints; using the target generative model, and based on the target reference image and the target noisy image of each viewpoint, to obtain M predicted material channel images under each viewpoint; wherein, the target generative model is trained using the model training method described in claims 1 to 8. The rendering unit is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.

16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-8 or 9-11.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8 or 9-11.

18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8 or 9-11.