Model training method, rendering method and device thereof, and apparatus
Patent Information
- Application Number
- CN202610677162.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-05-15
AI Technical Summary
[0003]本公开提供了一种模型训练方法、效果图生成方法及其装置、设备,以解决或缓解现有技术中的一项或更多项技术问题
[0011] In this way, the disclosed solution can utilize the feature extraction module in the generative model to be trained to form multiple modal data (such as initial prompt text, initial reference image, and geometric information of the initial 3D model from each viewpoint (i.e., the contour information carried by the initial noisy image from each viewpoint)) in a unified representation space, thereby obtaining target semantic control features and target visual prior features corresponding to each viewpoint. This achieves effective fusion and deep correlation of different modal data, enhancing the controllability and generalization ability of the generated materials. Furthermore, by using the prediction head module and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, the solution predicts M estimated material channel images and estimated noise corresponding to each viewpoint. Based on this, the solution trains the model, enabling it to more accurately understand and reconstruct material details, effectively enhancing the consistency and realism of the generated results.
Smart Images

Figure CN122223200B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a model training method, a method for generating effect diagrams, and the apparatus and equipment thereof. Background Technology
[0002] With the development of 3D model reconstruction technology, many reconstruction schemes have emerged in the field of 3D model reconstruction, such as implicit volume optimization, explicit volume optimization, and multimodal generation. However, existing reconstruction schemes still have the following problems: the consistency of the generated multiple views is difficult to guarantee, texture misalignment and insufficient material standardization output are difficult to meet the high-quality requirements of 3D model reconstruction. Summary of the Invention
[0003] This disclosure provides a model training method, a method for generating effect diagrams, and related apparatus and devices to solve or alleviate one or more technical problems in the prior art.
[0004] In a first aspect, this disclosure provides a model training method for a generative model used to reconstruct the material of a 3D model, including: Obtain target training samples; the target training samples include initial prompt text, initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a 2D image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; the initial prompt text is used to describe the initial 3D model; The M initial material channel images from each viewpoint are fused with the initial noise corresponding to each viewpoint to obtain the initial noisy image from each viewpoint. The initial prompt text, the initial reference image, and the initial noisy images from each viewpoint are input into the generative model to be trained. The feature extraction module in the generative model is used to obtain target semantic control features and target visual prior features corresponding to each viewpoint. The prediction head module in the generative model, combined with the target semantic control features and the target visual prior features corresponding to each viewpoint, is used to obtain M estimated material channel maps for each viewpoint, and estimated noise for each viewpoint. The target semantic control features characterize the geometric structure of the initial 3D model from each viewpoint, and the target visual prior features corresponding to each viewpoint characterize the material distribution information of the initial material of the initial 3D model from each viewpoint. The target loss value is obtained based on the estimated material channel map and / or estimated noise; Using the target loss value, the feature extraction module and prediction head module in the generative model to be trained are used to train the model, so as to obtain the target generative model.
[0005] Secondly, this disclosure provides a method for generating material renderings of a 3D model, including: The process involves acquiring mesh data of a target 3D model, a target reference image, and target prompt text. The target reference image is a 2D image of a reference 3D model with the target material. The reference 3D model is structurally related to the target 3D model. The target prompt text is used to describe the target 3D model. The mesh data of the target 3D model is preprocessed to obtain noisy images of the target 3D model from each of the N viewpoints; Using a target generative model, and based on the target prompt text, the target reference image, and the target noisy image from each viewpoint, M predicted material channel images are obtained for each viewpoint; wherein, the target generative model is trained using the model training method described above; Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, a rendering effect map of the target 3D model with the target material is obtained.
[0006] Thirdly, this disclosure provides a model training apparatus for reconstructing a generative model of a 3D model material, comprising: A sample acquisition unit is used to acquire target training samples; the target training samples include initial prompt text, initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; the initial prompt text is used to describe the initial 3D model; The model training unit is used to fuse M initial material channel maps from each viewpoint with the initial noise corresponding to each viewpoint to obtain initial noisy images from each viewpoint. The initial prompt text, the initial reference image, and the initial noisy images from each viewpoint are input into the generative model to be trained. The feature extraction module in the generative model to be trained obtains target semantic control features and target visual prior features corresponding to each viewpoint. The prediction head module in the generative model to be trained, combined with the target semantic control features and the target visual prior features corresponding to each viewpoint, obtains M estimated material channel maps from each viewpoint and estimated noise corresponding to each viewpoint. The target semantic control features characterize the geometric structure of the initial 3D model from each viewpoint, and the target visual prior features corresponding to each viewpoint characterize the material distribution information of the initial material of the initial 3D model from each viewpoint. Based on the estimated material channel maps and / or estimated noise, a target loss value is obtained. Using the target loss value, the feature extraction module and prediction head module in the generative model to be trained are used to train the model to obtain the target generative model.
[0007] Fourthly, this disclosure provides a device for generating material renderings of a three-dimensional model, including: The data input unit is used to acquire mesh data of the target 3D model, a target reference image, and target prompt text. The target reference image is a 2D image of a reference 3D model with the target material. The reference 3D model is structurally related to the target 3D model. The target prompt text is used to describe the target 3D model. The material generation unit is used to preprocess the mesh data of the target 3D model to obtain the target noisy image of the target 3D model in N viewpoints; using the target generative model, and based on the target prompt text, the target reference image and the target noisy image of each viewpoint, M predicted material channel maps are obtained under each viewpoint; wherein, the target generative model is trained using the above-mentioned model training method. The rendering unit is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.
[0008] Fifthly, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0009] In a sixth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.
[0010] In a seventh aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.
[0011] In this way, the disclosed solution can utilize the feature extraction module in the generative model to be trained to form multiple modal data (such as initial prompt text, initial reference image, and geometric information of the initial 3D model from each viewpoint (i.e., the contour information carried by the initial noisy image from each viewpoint)) in a unified representation space, thereby obtaining target semantic control features and target visual prior features corresponding to each viewpoint. This achieves effective fusion and deep correlation of different modal data, enhancing the controllability and generalization ability of the generated materials. Furthermore, by using the prediction head module and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, the solution predicts M estimated material channel images and estimated noise corresponding to each viewpoint. Based on this, the solution trains the model, enabling it to more accurately understand and reconstruct material details, effectively enhancing the consistency and realism of the generated results.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0014] Figure 1 This is an illustrative flowchart of a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application. Figure 1 ; Figure 2 This is an illustrative flowchart of a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application. Figure 2 ; Figure 3 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 1 ; Figure 4This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 2 ; Figure 5 This is an illustrative diagram of a generative model to be trained according to an embodiment of this application. Figure 3 ; Figure 6 This is an illustrative schematic diagram of K cascaded diffusion transformation layers according to an embodiment of this application; Figure 7 This is a schematic flowchart illustrating a method for generating a material effect diagram of a three-dimensional model according to an embodiment of this application; Figure 8 This is a schematic diagram of a method for generating a material effect map of a three-dimensional model according to an embodiment of this application in one example; Figure 9 This is a schematic diagram of the structure of a model training device for reconstructing a generative model of a 3D model material according to an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a three-dimensional model material effect drawing generation device according to an embodiment of this application; Figure 11 This is a block diagram of an electronic device used to implement the model training method for reconstructing a generative model of a three-dimensional model material or the method for generating a material effect map of a three-dimensional model according to the embodiments of this disclosure. Detailed Implementation
[0015] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0016] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0017] With the development of 3D model reconstruction technology, numerous reconstruction schemes have emerged in the field of 3D model reconstruction, such as implicit volume optimization, explicit volume optimization, and multimodal generation. Current 3D reconstruction schemes specifically include: Solution A: Implicit volume optimization solution, the core logic of which is: Taking text or an image as input, this method optimizes the implicit volume (i.e., neural radiance fields (NeRF) or signed distance function (SDF)) through score distillation to fit a 3D scene, and then obtains the target view through volume rendering. This allows for the generation of 3D models from natural language or single images, but this approach has high requirements for computational resources and training time.
[0018] Option B: Explicit structural optimization scheme, the core logic of which is: Optimization is achieved by employing explicit 3D structures (such as meshes (M), point clouds (Point), or Gaussian volumes) and incorporating scoring feedback from a diffusion model to iteratively improve the geometry and texture of the 3D structure, thereby directly generating editable 3D models. This improves later-stage usability, but the training stability and consistency of the approach remain limited.
[0019] Option C: A fast explicit representation scheme, whose core logic is as follows: This approach enables efficient reconstruction and real-time rendering by rapidly fitting lightweight structures (such as point clouds or Gaussian volumes) through multi-view supervision. Prioritizing speed and scalability, it is suitable for low-latency scenarios, such as real-time augmented reality (AR) or virtual reality (VR) modeling.
[0020] Solution D: Generative multimodal solution (Text-to-3D / Multi-view Diffusion), its core logic is as follows: By integrating multimodal inputs such as text, images, and depth, we can achieve more semantically understood 3D model generation. However, for complex model structures, there are still discrepancies between the semantic and geometric mappings.
[0021] However, the above-mentioned 3D model reconstruction scheme still has significant limitations and cannot meet the high-quality requirements of 3D model reconstruction. Specifically, these limitations are as follows: (a) Long model training and inference time: When iterating and optimizing a single object, the 3D generation scheme based on score distillation sampling (SDS) has a long optimization time and is difficult to generate quickly.
[0022] (b) Multi-view Figure 1 Inconsistency issues: Multiple views generated from a single image or text are prone to inconsistencies, leading to problems such as reconstruction tearing and texture misalignment.
[0023] (c) Semantic and style alignment is unstable: under textual or weak image conditions, detail style is prone to deviate from structural semantics.
[0024] (d) Insufficient standardized output: The generated 3D model lacks the complete Level of Detail (LOD), U-axis-V-axis mapping and material information required for cross-engine or pipeline docking, resulting in high implementation costs.
[0025] (e) Low efficiency of multimodal fusion: The fusion methods for multimodal features such as text, image, and geometry are relatively loose, failing to form a unified representation space and resulting in poor generalization ability.
[0026] Based on this, the present disclosure provides a model training method for a generative model used to reconstruct the materials of a 3D model. This method uses initial prompt text, initial reference image, initial noisy images of each of N viewpoints, and initial noise corresponding to each viewpoint to train the feature extraction module and prediction head module in the generative model to be trained, thereby obtaining a target generative model that can output M estimated material channel maps under each viewpoint. The generative model obtained by the above training can standardize the output of multiple material channel maps under each viewpoint, thus effectively solving the problem of texture misalignment caused by independent generation of materials from each viewpoint and enhancing the consistency of materials across multiple viewpoints. Moreover, the present disclosure can fuse multiple input data such as initial reference image, initial prompt text, and geometry into a unified representation through the above feature extraction module, thereby enabling the model to accurately predict the distribution of materials on the initial 3D model, thus improving the model's generalization ability and laying the foundation for obtaining high-quality renderings of the reconstructed 3D model using the target generative model.
[0027] Specifically, Figure 1 This is an illustrative flowchart of a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0028] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes: Step S101: Obtain the target training sample.
[0029] Here, the target training samples include initial prompt text, initial reference image, initial noise corresponding to each of the N (N is an integer greater than 1) viewpoints, and M (M is an integer greater than 1) initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints.
[0030] Furthermore, the initial reference image is a two-dimensional image of a reference 3D model with the initial material, and the structure of the reference 3D model in the two-dimensional image is related to that of the initial 3D model. For example, in one example, the structure of the reference 3D model in the initial reference image is similar to that of the initial 3D model, or, further, the structure of the reference 3D model in the initial reference image is the same as that of the initial 3D model. In this way, prior structural information of the 3D model is provided for the material diffusion of the subsequent model, thereby fully learning the distribution of the material on the initial 3D model from different perspectives.
[0031] It should be noted that the "material channel map" referred to in this disclosure refers to image data formed by breaking down the composite material properties of a 3D model surface into independent individual material properties. Each material channel map stores data corresponding to one material property, such as grayscale images, color images, or other data formats suitable for representing that material property. Furthermore, each material channel map corresponds to different material channels, such as primary color / diffuse reflection, normal, roughness, and metallicity channels in a physically based rendering (PBR) workflow, thus providing input data for the rendering calculations of the corresponding channels.
[0032] In other words, in one example, the different initial material channel maps in the M initial material channel maps of this disclosed scheme correspond to different material properties. In other words, different initial material channel maps carry different types of material property information. This provides rich data support for the model to fully learn the relationship between viewpoint, geometric information, and material information.
[0033] In one example, the M material properties include, but are not limited to: diffuse reflection, refraction, reflection, opacity, base color, metallicity, roughness, etc., and this disclosure does not specifically limit them.
[0034] In other words, the target training samples in this example include not only the initial reference image of the reference 3D model with initial material in a single view, but also the initial reference images of the reference 3D model with initial material in each of the N viewpoints. In this way, the material generated by the model can be further guided to maintain consistency across different viewpoints.
[0035] Furthermore, in one example, the initial prompt text is used to describe the geometry and / or materials of the initial 3D model, such as the appearance, structure, and materials of the initial 3D model, thus providing sufficient prior information for subsequent model generation of the required materials.
[0036] Step S102: Fuse the M initial material channel images under each viewpoint with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.
[0037] It should be noted that, in one example, different viewpoints can correspond to a single initial noise. That is, the initial noise corresponding to each viewpoint mentioned above can be specifically understood as a single initial noise corresponding to each viewpoint. In other words, different initial material channel maps under the same viewpoint correspond to the same initial noise. In this case, N viewpoints correspond to N initial noises. Furthermore, it should be noted that the initial noise corresponding to different viewpoints can be the same or different. For example, some viewpoints may correspond to the same initial noise, or other viewpoints may correspond to different initial noises. This disclosure does not impose specific restrictions on this.
[0038] Alternatively, in another example, the initial noise corresponding to different material channel maps under the same viewpoint is different; or, some material channel maps under the same viewpoint correspond to the same initial noise, while other material channel maps correspond to another initial noise; or, the initial noise corresponding to different material channel maps under the same viewpoint is different from each other. In this case, for the same viewpoint, there are M initial noises, and further, for N viewpoints, there are N×M initial noises.
[0039] It should be noted that the above examples of initial noise are merely illustrative. In practical applications, the noise corresponding to the viewpoint can be determined based on the actual scenario requirements and actual training accuracy requirements. This disclosure does not impose any specific restrictions on this.
[0040] In addition, it should be noted that this disclosure does not specifically limit the method of generating the initial noise.
[0041] Furthermore, in one example, fusing the M initial material channel maps from each viewpoint with the single initial noise corresponding to each viewpoint can specifically include: For the initial material channel map to be processed under the current viewpoint, the initial noise corresponding to the current initial material channel map is determined, and it is fused with the determined initial noise. This yields the initial noisy image corresponding to the current initial material channel map under the current viewpoint. Following this logic, for M initial material channel maps under the current viewpoint, M initial noisy images corresponding to the current viewpoint can be obtained. Furthermore, for N viewpoints, N×M initial noisy images can be obtained.
[0042] Furthermore, in one example, the above-described process of fusing the M initial material channel images from each viewpoint with the initial noise corresponding to each viewpoint to obtain the initial noisy image for each viewpoint (e.g., step S102) can specifically include: The M initial material channel images under each viewpoint are linearly interpolated with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.
[0043] For example, in one example, linear interpolation of the M initial material channel maps under each viewpoint with the single initial noise corresponding to each viewpoint can specifically include: For the initial material channel map to be processed under the current viewpoint, the initial noise corresponding to the current initial material channel map is determined, and the current initial material channel map and the determined initial noise are fused (e.g., through linear interpolation). This yields the initial noisy image corresponding to the current initial material channel map under the current viewpoint. Following this logic, for M initial material channel maps under the current viewpoint, M initial noisy images corresponding to the current viewpoint can be obtained. Furthermore, for N viewpoints, N×M initial noisy images can be obtained.
[0044] Thus, the present invention provides an image noise-adding processing method, namely, using linear interpolation to add noise to the initial material channel map. This can construct a noisy image with an extremely smooth noise distribution, eliminating the random redundancy of the traditional noise-adding process. This provides favorable support for subsequent training to obtain a more generalizable target generative model.
[0045] Step S103: Input the initial prompt text, the initial reference image, and the initial noisy images of each viewpoint into the generative model to be trained, so as to obtain the target semantic control features and the target visual prior features corresponding to each viewpoint using the feature extraction module in the generative model to be trained; and obtain M estimated material channel images and estimated noise corresponding to each viewpoint by using the prediction head module in the generative model to be trained, combined with the target semantic control features and the target visual prior features corresponding to each viewpoint.
[0046] Here, in this example, the target semantic control features are used to characterize the geometric structure of the initial 3D model from various perspectives, that is, to characterize the geometric structure features from various perspectives.
[0047] Furthermore, the target visual prior features corresponding to each viewpoint are used to characterize the material distribution information of the initial material of the initial 3D model under each viewpoint. That is, for a specific viewpoint, the target visual prior features under that viewpoint can be specifically used to characterize the material distribution features of the initial material of the initial 3D model under that viewpoint.
[0048] Here, in one example, the above-described method of utilizing the prediction head module in the generative model to be trained, combined with the target semantic control features and the target visual prior features corresponding to each viewpoint, to obtain M estimated material channel maps under each viewpoint, and the estimated noise corresponding to each viewpoint, can specifically include: Using the prediction head module in the generative model to be trained, and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, the estimated noise corresponding to the initial noisy image under each viewpoint is first estimated. For example, for a scene with M initial noisy images under each of N viewpoints, N×M estimated noises can be estimated. Then, using the estimated noise corresponding to the initial noisy image under each viewpoint, the initial noisy image corresponding to the estimated noise is denoised. In this way, M estimated material channel images under each viewpoint are obtained, for a total of N×M estimated material channel images.
[0049] In other words, in this example, the prediction head module is first used to estimate the estimated noise corresponding to the initial noisy image under each viewpoint. Then, based on the estimated noise, the initial noisy image corresponding to the estimated noise is denoised to obtain the estimated material channel map corresponding to the initial noisy image. In this way, N×M estimated material channel maps are obtained.
[0050] Step S104: Obtain the target loss value based on the estimated material channel map and / or estimated noise.
[0051] Step S105: Using the target loss value, train the feature extraction module and prediction head module in the generative model to be trained to obtain the target generative model.
[0052] In other words, this disclosed solution inputs the initial prompt text, initial reference image, and initial noisy images from each viewpoint into the generative model to be trained. The feature extraction module within the generative model extracts features from these materials to obtain target semantic control features and target visual prior features corresponding to each viewpoint. Then, using the prediction head module within the generative model, combined with the target semantic control features and the target visual prior features, it predicts M estimated material channel maps for each viewpoint, as well as the estimated noise for each viewpoint. Finally, using the M estimated material channel maps and / or the estimated noise for each viewpoint, it obtains the target loss value. Finally, using the obtained target loss value, it fine-tunes the parameters of the feature extraction module and prediction head module in the generative model to obtain the target generative model.
[0053] In this way, the disclosed solution can utilize the feature extraction module in the generative model to be trained to form multiple modal data (such as initial prompt text, initial reference image, and geometric information of the initial 3D model from each viewpoint (i.e., the contour information carried by the initial noisy image from each viewpoint)) in a unified representation space, thereby obtaining target semantic control features and target visual prior features corresponding to each viewpoint. This achieves effective fusion and deep correlation of different modal data, enhancing the controllability and generalization ability of the generated materials. Furthermore, by using the prediction head module and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, the solution predicts M estimated material channel images and estimated noise corresponding to each viewpoint. Based on this, the solution trains the model, enabling it to more accurately understand and reconstruct material details, effectively enhancing the consistency and realism of the generated results.
[0054] Moreover, since the above learning process integrates global information from the initial reference image and material channel information from each viewpoint, it helps to coordinate the material representation between different viewpoints, reduce differences between viewpoints, enhance the consistency of materials from multiple viewpoints, and thus solve the problem of texture misalignment caused by independent generation of materials from each viewpoint, making the subsequently generated materials more realistic and coherent.
[0055] Furthermore, the disclosed solution can use M estimated material channel maps from each viewpoint and / or the estimated noise corresponding to each viewpoint to calculate the target loss value, so as to train the generative model to be trained, thereby obtaining a target generative model that can standardize and output material channel maps from multiple viewpoints. The low training complexity of the above training solution reduces the time and computing resources required for model training, which provides favorable support for efficient 3D reconstruction using the target generative model.
[0056] Moreover, since the generative model trained by this disclosed solution can output multiple material channel maps from different perspectives, such as outputting N×M estimated material channel maps, compared with the existing solution that uses 4 material channel maps for 3D reconstruction, this disclosed solution can significantly increase the number of material channel dimensions, thus enriching the dimensions of material attribute representation and effectively improving the quality of materials generated by the trained model.
[0057] In addition, the proposed solution utilizes a feature extraction module to achieve alignment of multimodal conditions in a unified feature space, thereby reducing the complexity of model training and the time and computational resources required for model training. This provides favorable support for efficient 3D reconstruction using target generative models.
[0058] Furthermore, in a specific example, the target loss value can be obtained as follows; specifically, the above-described method of obtaining the target loss value based on the estimated material channel map and / or estimated noise (e.g., step S103) can specifically include: Method 1: Based on the M estimated material channel maps under each viewpoint and the M initial material channel maps under each viewpoint, the target loss value is obtained.
[0059] In other words, under this method 1, the M initial material channel maps (i.e., N×M initial material channel maps) from each viewpoint are used as label data, and the resulting N×M estimated material channel maps are used as the estimated values output by the model. Then, the target loss value is obtained based on the difference information between the two.
[0060] Method 2: Based on the estimated noise corresponding to each viewpoint and the initial noise corresponding to each viewpoint, the target loss value is obtained.
[0061] In other words, under this method 2, the initial noise (e.g., N×M initial noises) from each perspective is used as the label data, and the estimated noise (e.g., N×M estimated noises) is used as the estimated value output by the model. Then, the target loss value is obtained based on the difference between the two.
[0062] Method 3: Based on M initial material channel maps under each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image of each viewpoint, the target loss value is obtained. For example, in one example, the above-mentioned method of obtaining the target loss value based on the M initial material channel maps under each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image of each viewpoint (i.e., Method 3 above) can specifically include: Step 1: Perform linear interpolation on the M initial material channel images under each viewpoint and the estimated noise corresponding to each viewpoint to obtain the estimated noisy image for each viewpoint.
[0063] For details on the linear interpolation process in step 1, please refer to the example above; it will not be repeated here.
[0064] Step 2: Based on the estimated noisy images from each viewpoint and the initial noisy images from each viewpoint, obtain the target loss value.
[0065] In other words, the estimated noise obtained from each viewpoint of the generative model to be trained is linearly interpolated with the M material channel images under each viewpoint to obtain the estimated noisy image for each viewpoint. Then, using a loss function, such as the mean squared error loss function, and combining the image vectors of the estimated noisy image and the image vectors of the initial noisy image for each viewpoint, the difference information between the two image vectors is calculated to obtain the target loss value. Thus, compared with existing schemes that directly use the estimated noise as the model optimization target, the scheme disclosed in this publication further enhances the stability and robustness of model training, providing favorable support for subsequent training to obtain the target generative model.
[0066] It should be noted that, in calculating the target loss value, this disclosed solution may use any one of the three methods described above, or any two of the three methods (such as method 1 and method 2, or method 1 and method 3, or method 2 and method 3), or simultaneously use all three methods to calculate the target loss value. This disclosed solution does not impose specific restrictions on which method is used to calculate the target loss value.
[0067] For example, in one instance, the present invention can obtain the target loss value based on the difference information between the M estimated material channel maps under each viewpoint and the M initial material channel maps under each viewpoint, and the difference information between the estimated noise corresponding to each viewpoint and the initial noise corresponding to each viewpoint. That is, the first loss value can be obtained using method 1 described above, and the second loss value can be obtained using method 2 described above; then, the target loss value can be obtained using the obtained first and second loss values.
[0068] Alternatively, in another example, the disclosed solution can obtain the target loss value based on the difference information between the M estimated material channel maps under each viewpoint and the M initial material channel maps under each viewpoint, the difference information between the estimated noise corresponding to each viewpoint and the initial noise corresponding to each viewpoint, and the difference information between the estimated noisy image under each viewpoint (that is, the noisy image obtained after linear interpolation of the M initial material channel maps under each viewpoint and the estimated noise corresponding to each viewpoint) and the initial noisy image under each viewpoint. In other words, the first loss value can be obtained using method 1 described above, the second loss value can be obtained using method 2 described above, and the third loss value can be obtained using method 3 described above; then, the target loss value can be obtained using the obtained first loss value, second loss value, and third loss value.
[0069] The above is merely an example. In practical applications, the appropriate loss calculation method can be selected according to actual needs (such as accuracy requirements). This disclosure does not impose specific restrictions on the method used to calculate the loss.
[0070] Thus, this disclosed solution provides a specific method for calculating the target loss value. This method is simple, practical, and highly interpretable, thereby giving the model training process great flexibility. It enables the model training to select the most suitable loss calculation method according to the specific needs of the scenario, thereby effectively avoiding the optimization bias or overfitting risk that may be caused by a single loss, and providing favorable support for subsequent training to obtain the target generative model.
[0071] Figure 2 This is an illustrative flowchart of a model training method for reconstructing a generative model of a 3D model material according to an embodiment of this application. Figure 2 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that... Figure 1 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0072] Furthermore, the method includes at least a portion of the following: For example... Figure 2 As shown, it includes: Step S201: Obtain the target training sample.
[0073] Here, the target training samples include initial prompt text, initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the reference 3D model is structurally related to the initial 3D model; the initial prompt text is used to describe the initial 3D model.
[0074] It should be noted that the information regarding the initial prompt text, initial reference image, and M initial material channel images from each viewpoint can be found in the example above, and will not be repeated here.
[0075] Step S202: Fuse the M initial material channel images under each viewpoint with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.
[0076] It should be noted that the processing logic of step S202 can be referred to the example of step S102 above, and will not be repeated here.
[0077] Step S203: Input the initial prompt text, the initial reference image, and the initial noisy images of each viewpoint into the generative model to be trained, so as to utilize the K (K is an integer greater than 1) cascaded diffusion transformation layers in the generative model to be trained, and combine the initial prompt text, the initial reference image, and the initial noisy images corresponding to each viewpoint to obtain the Kth semantic control feature output by the last diffusion transformation layer in the K diffusion transformation layers and the Kth visual prior feature corresponding to each viewpoint.
[0078] In other words, in this example, the feature extraction module includes K cascaded diffusion transformation layers. It should be noted that the "K cascaded diffusion transformation layers" in this disclosure can specifically refer to K diffusion transformation layers connected in series. That is, in the K cascaded diffusion transformation layers, the output of the previous diffusion transformation layer serves as the input of the next diffusion transformation layer. For the first (i.e., the 1st) diffusion transformation layer in the K cascaded diffusion transformation layers, its input can specifically be the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint. The output of the last diffusion transformation layer is: the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint.
[0079] For example, in one example, the i-th diffusion transform layer (where i is an integer greater than 0 and less than or equal to K) in a cascaded K diffusion transform layer is used to perform feature processing on the (i-1)-th semantic control feature output by the (i-1)-th diffusion transform layer and the (i-1)-th visual prior features corresponding to each viewpoint, to obtain the i-th semantic control feature and the i-th visual prior features corresponding to each viewpoint. Here, for the first diffusion transform layer, the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint are used for feature processing. The output of the last diffusion transform layer is the K-th semantic control feature and the K-th visual prior features corresponding to each viewpoint.
[0080] Step S204: Using the feature transformation layer in the generative model to be trained, and combining the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
[0081] In other words, in this example, the feature extraction module also includes a feature transformation layer for adjusting the semantic dimension and spatial structure of the features.
[0082] In one example, the diffusion transformation layer may be specifically a diffusion transformer (DiT), or it may be other network structures with diffusion processing capabilities. This disclosure does not impose any specific limitations on this.
[0083] In one example, the feature transformation layer may be specifically a linear and reshape layer, or it may be other network layers with the above processing capabilities. This disclosure does not impose any specific limitations on this.
[0084] Step S205: Using the prediction head module in the generative model to be trained, and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, obtain M estimated material channel maps under each viewpoint, and estimated noise corresponding to each viewpoint.
[0085] Here, the target semantic control features are used to characterize the geometric structure of the initial 3D model from each viewpoint, and the target visual prior features corresponding to each viewpoint are used to characterize the material distribution information of the initial material of the initial 3D model from each viewpoint.
[0086] It should be noted that the relevant content regarding obtaining the M estimated material channel maps from each viewpoint can be found in the example above, and will not be repeated here.
[0087] Step S206: Obtain the target loss value based on the estimated material channel map and / or estimated noise.
[0088] Step S207: Using the target loss value, train the feature extraction module and prediction head module in the generative model to be trained to obtain the target generative model.
[0089] For example, such as Figure 3As shown, the feature extraction module in the generative model to be trained includes K cascaded diffusion transformation layers (i.e., the first diffusion transformation layer, the second diffusion transformation layer, ..., the Kth diffusion transformation layer) and a feature transformation layer, and the prediction head module includes a multilayer perceptron (MLP) for prediction. At this point, the initial prompt text, initial reference image, and initial noisy images from each viewpoint can be input into the cascaded K diffusion transformation layers in the feature extraction module to obtain the Kth semantic control feature output by the Kth diffusion transformation layer and the Kth visual prior feature corresponding to each viewpoint. Further, the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint are input into the feature transformation module in the feature extraction module to obtain the target semantic control feature and the target visual prior feature corresponding to each viewpoint. Further, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are input into the MLP in the prediction head module to predict M estimated material channel maps under each viewpoint and the estimated noise corresponding to each viewpoint. Finally, using the estimated material channel maps and / or estimated noise, the target loss value is calculated, and the obtained target loss value is used to train the feature extraction module and prediction head module in the generative model to obtain the target generative model.
[0090] In this way, the proposed solution utilizes K cascaded diffusion transformation layers in the feature extraction module to perform layer-by-layer feature processing on the initial prompt text, initial reference image, and initial noisy images from various perspectives. This effectively captures complex semantic association information and cross-perspective visual dependence information in data from different modalities, enhancing the richness of features. This provides abundant feature information for model training. Furthermore, the feature transformation layers in the feature extraction model further enhance the expressive power of the features and strengthen the consistency and complementarity of visual features from multiple perspectives in a unified representation space. Thus, the robustness of model training and the realism of the generated results are effectively improved, thereby enhancing the model's generalization ability and laying the foundation for high-precision 3D reconstruction using the trained target generative model.
[0091] Furthermore, in a specific example, the feature extraction module further includes a normalization layer; specifically, after obtaining the Kth semantic control feature output by the last diffusion transformation layer of the K diffusion transformation layers and the Kth visual prior features output by each viewpoint using the K diffusion transformation layers in the generative model to be trained, combined with the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint (for example, after step S203), the module further includes: By utilizing the normalization layer and combining the Kth semantic control feature with the Kth visual prior feature corresponding to each viewpoint, the normalized Kth semantic control feature and the normalized Kth visual prior feature corresponding to each viewpoint are obtained.
[0092] Furthermore, the above-described method of utilizing the feature transformation layer, combined with the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint, to obtain the target semantic control feature and the target visual prior feature corresponding to each viewpoint (e.g., step S204) can specifically include: By utilizing the aforementioned feature transformation layer, and combining the normalized Kth semantic control feature with the normalized Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
[0093] For example, continue with Figure 3 Taking the model structure and model processing shown as an example, as Figure 4 As shown, a normalization layer is set between the cascaded K diffusion transformation modules and the feature transformation layer in the feature extraction module. At this time, after obtaining the Kth semantic control feature output by the Kth diffusion transformation layer in the cascaded K diffusion transformation modules and the Kth visual prior features corresponding to each viewpoint, the Kth semantic control feature and the Kth visual prior features corresponding to each viewpoint are input into the normalization layer for normalization processing to obtain the normalized Kth semantic control feature and the normalized Kth visual prior features corresponding to each viewpoint. Then, the normalized Kth semantic control feature and the normalized Kth visual prior features corresponding to each viewpoint are input into the feature transformation layer to obtain the target semantic control feature and the target visual prior features corresponding to each viewpoint.
[0094] In this way, the proposed solution utilizes a normalization layer to normalize the Kth semantic control feature output by the Kth diffusion transformation layer in the cascaded K diffusion transformation modules, as well as the Kth visual prior features corresponding to each viewpoint. This effectively eliminates the differences in numerical scale and distribution of features of different modalities, while suppressing the interference of abnormal information on subsequent calculations, thus providing favorable support for improving the convergence speed of model training and the stability of the generated results.
[0095] Further, in a specific example, the i-th diffusion transformation layer among the K diffusion transformation layers in the generative model to be trained may specifically include: an i-th fully connected layer for obtaining the weight matrices corresponding to each data path in the multi-path data, an i-th sequence concatenation block for performing sequence concatenation processing on each weight matrix, an i-th fully self-attention block for capturing global dependency information between all elements in the features after sequence concatenation processing, an i-th text linear transformer for extracting global semantic information from the features after fully self-attention processing, and an i-th image linear transformer for extracting global visual information from the features after fully self-attention processing.
[0096] It should be noted that in one example, the three types of data input to the generative model to be trained (i.e., the initial prompt text, the initial reference image, and the initial noisy image) correspond to different data transmission paths. In this example, the three types of input data correspond to three data paths: the first path corresponds to the initial prompt text, the second path corresponds to the initial reference image, and the third path corresponds to the initial noisy images from each viewpoint.
[0097] For example, continue with Figure 4 Taking the model structure and processing shown as an example, the feature extraction module in the generative model to be trained includes K cascaded diffusion transformation layers. At this point, as... Figure 5 As shown, the i-th diffusion transform layer in the K diffusion transform layers can specifically include: the i-th fully connected layer, the i-th sequence cascade block, the i-th fully self-attention block, the i-th text linear transformer, and the i-th image linear transformer.
[0098] Furthermore, for the i-th diffusion transform layer among the K diffusion transform layers, after obtaining the features output by the previous diffusion transform layer (such as the (i-1)-th semantic control features output by the (i-1)-th diffusion transform layer and the (i-1)-th visual prior features corresponding to each viewpoint), the obtained features are input into the i-th fully connected layer in the i-th diffusion transform layer to obtain multiple weight matrices corresponding to each data path. For example, in one example, the multiple weight matrices corresponding to each data path include the query matrix, key matrix, and value matrix corresponding to that data path. Further, the multiple weight matrices corresponding to each data path are input into the i-th sequence concatenation block to connect the weight matrices of the same matrix type corresponding to each data path, for example, connecting the query matrices corresponding to each data path. The process involves concatenating the key matrices corresponding to each feature path and concatenating the value matrices corresponding to each feature path to obtain the target weight matrix (e.g., obtaining the weight matrices corresponding to various matrix types, such as the total query matrix, the total key matrix, and the total value matrix after concatenation). The target weight matrix is then input into the i-th fully self-attention block to perform self-attention processing on the weight matrices of different matrix types, resulting in global context features. The global context features are then input into the i-th text linear transformer (e.g., the i-th text MLP) to obtain the i-th semantic control features, and the global context features are input into the i-th image linear transformer (e.g., the i-th image MLP) to obtain the i-th visual prior features corresponding to each viewpoint.
[0099] Here, for the first diffusion transform layer among the K diffusion transform layers, the input of the first diffusion transform layer is the initial prompt text, the initial reference image, and the initial noisy images of each viewpoint; the Kth semantic control feature output by the last diffusion transform layer is the output result of the Kth text linear transformer in the last diffusion transform layer, and the Kth visual prior features corresponding to each viewpoint output are the output result of the Kth image linear transformer in the last diffusion transform layer.
[0100] Thus, the present invention provides a specific structure for a diffusion transformation layer. This structure not only improves the depth and efficiency of multimodal feature interaction, but also ensures that the generated material is highly semantically consistent with the prompt text and maintains strict geometric and textural consistency across multiple viewpoints, thereby significantly improving the stability of model training and the quality of the final generated results.
[0101] Furthermore, in one example, the i-th fully connected layer includes multiple fully connected blocks arranged in parallel, wherein each fully connected block corresponds to one data path to obtain multiple weight matrices corresponding to a single data path.
[0102] It should be noted that the number of fully connected blocks included in the i-th fully connected layer matches the number of input data paths to the feature extraction module. For example, if the data input to the feature extraction module includes three data paths: initial prompt text, initial reference image, and initial noisy images from each viewpoint, then the i-th fully connected layer can specifically include three parallel fully connected blocks.
[0103] Furthermore, in one example, in order to accurately obtain the feature information corresponding to each data path and achieve a high degree of control and accurate expression of the material generation process, after the i-th visual prior feature corresponding to each viewpoint is output by the i-th diffusion transformation layer, the i-th material guidance feature and the i-th spatial structure feature corresponding to each viewpoint can also be obtained based on the i-th visual prior feature corresponding to each viewpoint. At this time, the three features output by the i-th diffusion transformation layer correspond to the three input data paths, that is, the i-th diffusion transformation layer can output: the i-th semantic control feature corresponding to the initial prompt text, the i-th material guidance feature corresponding to the initial reference image, and the i-th spatial structure feature corresponding to the initial noisy image.
[0104] For example, continue with Figure 5 Taking the structure of the i-th diffusion transformation layer as an example, as shown... Figure 6As shown, the i-th fully connected layer in the i-th diffusion transform layer includes three i-th fully connected blocks arranged in parallel. At this point, after the (i-1)-th diffusion transform layer in the cascaded K diffusion transform layers outputs the i-th semantic control feature corresponding to the initial prompt text, the i-th material guidance feature corresponding to the initial reference image (i.e., the (i-1)-th material guidance feature corresponding to each viewpoint), and the i-th spatial structure feature corresponding to the initial noisy image (i.e., the (i-1)-th spatial structure feature corresponding to each viewpoint), the features corresponding to each data path are input in parallel to the three i-th fully connected blocks in the i-th diffusion transform layer to obtain multiple weight matrices corresponding to each data path. Further, the multiple weight matrices corresponding to each data path are input to the i-th sequence cascade block to obtain the target weight matrix, which is then input to... The i-th fully self-attention block performs self-attention processing on weight matrices of different matrix types to obtain global context features. These global context features are then input into the i-th text linear transformer (e.g., the i-th text MLP) to obtain the i-th semantic control features. Similarly, the global context features are input into the i-th image linear transformer (e.g., the i-th image MLP) to obtain the i-th visual prior features corresponding to each viewpoint. Based on these i-th visual prior features, the i-th material guidance features and i-th spatial structure features corresponding to each viewpoint are obtained. Furthermore, the obtained i-th semantic control features, i-th material guidance features, and i-th spatial structure features are input in parallel into three parallel fully connected blocks in the (i+1)-th diffusion transform layer. Thus, through continuous feature interaction and feature decoupling, the expressive power of the feature information corresponding to each data path is improved, thereby achieving a high degree of control over the material generation process.
[0105] Here, for the first diffusion transform layer in the cascaded K diffusion transform layers, the initial prompt text, the initial reference image, and the initial noisy images of each viewpoint are input in parallel to the three first fully connected blocks in the first diffusion transform layer.
[0106] In this way, the proposed solution can utilize multiple fully connected blocks of the i-th generation to process the input multi-path data in parallel. This allows input data from different sources (such as initial prompt text, initial reference image, and initial noisy images from various perspectives) to independently learn the optimal weights, thereby accurately capturing key semantic and visual information in each path of data and effectively suppressing the interference of redundant or noisy features. At the same time, it also enhances the model's adaptability to complex multimodal inputs, providing favorable support for improving the stability and robustness of subsequent model training.
[0107] Figure 7 This is a schematic flowchart illustrating a method for generating material effect diagrams of a 3D model according to an embodiment of this application. This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0108] Furthermore, the method includes at least a portion of the following: For example... Figure 7 As shown, it includes: Step S701: Obtain the mesh data of the target 3D model, the target reference image, and the target prompt text.
[0109] Here, the target reference image is a two-dimensional image of a reference three-dimensional model with the target material; the reference three-dimensional model is structurally related to the target three-dimensional model; the target prompt text is used to describe the target three-dimensional model.
[0110] It should be noted that the target reference image, reference 3D model, and target prompt text in this example can all be found in the example above, and will not be repeated here.
[0111] Step S702: Preprocess the mesh data of the target 3D model to obtain the target noisy image of the target 3D model from each of the N viewpoints.
[0112] Step S703: Using the target generative model, and based on the target prompt text, the target reference image, and the target noisy image from each viewpoint, obtain M predicted material channel images for each viewpoint.
[0113] Here, the target generative model is trained using the model training method of any of the above embodiments.
[0114] It should be noted that the material channel map can be referred to the example above, and will not be repeated here.
[0115] Step S704: Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, obtain the rendering effect map of the target 3D model with the target material.
[0116] In this way, the disclosed solution can utilize a trained target generative model to obtain M predicted material channel maps from each viewpoint. This reduces differences between viewpoints, effectively enhances the consistency of materials across multiple viewpoints, and solves the texture misalignment or artifact problems caused by independent material generation from each viewpoint, making the subsequently generated materials more realistic and coherent. Furthermore, by using the M predicted material channel maps from each viewpoint and the mesh data of the target 3D model, a rendered image of the target 3D model with the target material is obtained. Thus, the combination of the target material and the target 3D model is achieved, resulting in a high-quality rendered image.
[0117] Furthermore, since the present invention can utilize the target generative model to standardize the output of multiple material channel maps required for material generation, compared with the existing schemes that use 4 material channel maps for reconstruction, the present invention can output more material channel maps, thereby improving the representation accuracy of material properties. This further improves the quality of materials distributed on the target 3D model, thereby improving the quality of the rendering effect map of the target material and the target 3D model.
[0118] Further, in a specific example, the target noisy images from each viewpoint can be obtained in the following manner; specifically, the preprocessing of the mesh data of the target 3D model described above to obtain the target noisy images of the target 3D model from each of the N viewpoints (e.g., step S702) can specifically include: Step S702-1: Based on the mesh data of the target 3D model, obtain the texture image of the target 3D model under each of the N viewpoints and the position image under each of the N viewpoints.
[0119] Here, in this example, the texture image (also known as a normal map) is used to simulate the bumps and shadows of the surface of the target 3D model; the position image (also known as a position map) is used to store the position information of each point on the surface of the target 3D model in 3D space.
[0120] Step S702-2: Fusion processing is performed on the texture images, position images and random noise corresponding to each viewpoint to obtain the target noisy image of each viewpoint.
[0121] For example, in one example, the texture image and the position image under each viewpoint are fused together, such as the texture image and the position image under the same viewpoint are fused together to obtain the fused image under each viewpoint; then, the random noise corresponding to each viewpoint is used to add noise to the fused image under each viewpoint to obtain the target noisy image under each viewpoint.
[0122] Thus, the present disclosure provides a preprocessing scheme for mesh data of a target 3D model. This scheme is simple, practical and highly interpretable. Moreover, on the one hand, it achieves high-dimensional complementarity and strong coupling of the inherent geometric properties of the target 3D model surface (such as local normal direction and global spatial coordinates). On the other hand, by introducing random noise corresponding to each viewpoint, it provides a generalized input for the target generative model that conforms to the distribution of real complex physical imaging, thereby ensuring that the subsequent prediction process can stably and accurately reconstruct high-quality material information on the target 3D model.
[0123] Furthermore, in a specific example, the rendering effect of the target 3D model with the target material can be obtained in the following manner; specifically, the above-mentioned method of obtaining the rendering effect of the target 3D model with the target material based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint (for example, step S704) can specifically include: Step S704-1: Map the M predicted material channel maps from each viewpoint from the two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps.
[0124] For example, in one example, the M predicted material channel maps from each viewpoint are mapped from a two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model. For example, the M predicted material channel maps from each viewpoint are mapped to the mesh occupied by the target three-dimensional model in the three-dimensional space to obtain the three-dimensional material data corresponding to the target three-dimensional model. Then, the three-dimensional material data corresponding to the target three-dimensional model is transformed from the three-dimensional space to the two-dimensional space to obtain M material unfolded maps.
[0125] Step S704-2: Using the M material unfolded diagrams and the mesh data of the target 3D model, perform rendering processing to obtain a rendering effect diagram of the target 3D model with the target material.
[0126] For example, in one example, the renderer is invoked, and the M material unfolded maps and the mesh data of the target 3D model are combined to perform rendering processing, resulting in a rendered image of the target 3D model with the target material.
[0127] Thus, this disclosed solution provides a specific method for obtaining a rendered image using M predicted material channel maps from various perspectives. In this way, material transfer for the target 3D model is achieved, and a high-quality rendered image of the 3D model with the target material is reconstructed.
[0128] Figure 8 This is a schematic diagram generated from the material rendering of the 3D model of the disclosed solution, such as... Figure 8As shown, firstly, the white model data of the target 3D model (corresponding to the mesh data of the target 3D model mentioned above) is preprocessed to obtain texture images and position images of each viewpoint in N viewpoints. Then, the texture images, position images, and random noise corresponding to each viewpoint are fused to obtain the target noisy image of each viewpoint. Secondly, the descriptive text (corresponding to the target prompt text mentioned above, such as "A delicate electronic alarmclock" in the illustration), the reference images of the "electronic alarm clock" with the target material in N viewpoints (corresponding to the target reference image mentioned above), and the target noisy image of each viewpoint are input into the target generative model to obtain N×M predicted material channel maps. Finally, through the baking algorithm, and combined with the white model data of the target 3D model and the N×M predicted material channel maps, M material unfolded maps are obtained. Furthermore, the renderer is called, and combined with the white model data of the target 3D model and the obtained M material unfolded maps, the rendering process is performed to obtain the rendered image of the "electronic alarm clock" with the target material. Thus, the disclosed solution effectively enhances the consistency of materials from multiple perspectives, thereby improving the quality of materials distributed on the target 3D model, and consequently improving the quality of the rendered image of the target 3D model with the target material.
[0129] This disclosure provides a model training device for reconstructing generative models of 3D model materials, such as... Figure 9 As shown, it includes: The sample acquisition unit 901 is used to acquire target training samples; the target training samples include initial prompt text, initial reference image, initial noise corresponding to each viewpoint in N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each viewpoint of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the structure of the reference 3D model is related to that of the initial 3D model; the initial prompt text is used to describe the initial 3D model; The model training unit 902 is used to fuse M initial material channel maps from each viewpoint with the initial noise corresponding to each viewpoint to obtain initial noisy images from each viewpoint; input the initial prompt text, the initial reference image, and the initial noisy images from each viewpoint into the generative model to be trained, so as to obtain target semantic control features and target visual prior features corresponding to each viewpoint using the feature extraction module in the generative model to be trained; use the prediction head module in the generative model to be trained, and combine the target semantic control features and the target visual prior features corresponding to each viewpoint to obtain M estimated material channel maps from each viewpoint and estimated noise corresponding to each viewpoint; wherein, the target semantic control features are used to characterize the geometric structure of the initial 3D model from each viewpoint, and the target visual prior features corresponding to each viewpoint are used to characterize the material distribution information of the initial material of the initial 3D model from each viewpoint; based on the estimated material channel maps and / or estimated noise, a target loss value is obtained; using the target loss value, the feature extraction module and prediction head module in the generative model to be trained are used to train the model to obtain the target generative model.
[0130] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: The M initial material channel images under each viewpoint are linearly interpolated with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.
[0131] In a specific example of the scheme disclosed herein, the model training unit is specifically configured to perform at least one of the following: Based on the M estimated material channel maps from each viewpoint and the M initial material channel maps from each viewpoint, the target loss value is obtained. Based on the estimated noise corresponding to each viewpoint and the initial noise corresponding to each viewpoint, the target loss value is obtained; Based on the M initial material channel images from each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image from each viewpoint, the target loss value is obtained.
[0132] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Linear interpolation is performed on the M initial material channel images under each viewpoint and the estimated noise corresponding to each viewpoint to obtain the estimated noisy image of each viewpoint. The target loss value is obtained based on the estimated noisy image from each viewpoint and the initial noisy image from each viewpoint.
[0133] In a specific example of the scheme disclosed herein, the feature extraction module includes K cascaded diffusion transformation layers and a feature transformation layer for adjusting the semantic dimension and spatial structure of the features; Specifically, the model training unit is used for: Using the K diffusion transform layers, and combining the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint, the Kth semantic control feature output by the last diffusion transform layer in the K diffusion transform layers and the Kth visual prior feature corresponding to each viewpoint are obtained. By utilizing the feature transformation layer and combining the Kth semantic control feature with the Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
[0134] In a specific example of the scheme disclosed herein, the feature extraction module further includes a normalization layer; wherein, the model training unit is further configured to: By utilizing the normalization layer and combining the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint, the normalized Kth semantic control feature and the normalized Kth visual prior feature corresponding to each viewpoint are obtained. By utilizing the aforementioned feature transformation layer, and combining the normalized Kth semantic control feature with the normalized Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
[0135] In a specific example of the scheme disclosed herein, the i-th diffusion transformation layer among the K diffusion transformation layers includes: an i-th fully connected layer for obtaining the weight matrices corresponding to each data path in the multi-path data; an i-th sequence concatenation block for performing sequence concatenation processing on each weight matrix; an i-th fully self-attention block for capturing global dependency information among all elements in the features after sequence concatenation processing; an i-th text linear transformer for extracting global semantic information from the features after fully self-attention processing; and an i-th image linear transformer for extracting global visual information from the features after fully self-attention processing.
[0136] In a specific example of the scheme disclosed herein, the i-th fully connected layer includes multiple fully connected blocks arranged in parallel, wherein each fully connected block corresponds to one data path, so as to obtain multiple weight matrices corresponding to a single data path.
[0137] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0138] This disclosure also provides a device for generating material renderings of 3D models, such as... Figure 10 As shown, it includes: The data input unit 1001 is used to acquire mesh data of the target 3D model, a target reference image, and target prompt text. The target reference image is a 2D image of a reference 3D model with the target material. The reference 3D model is structurally related to the target 3D model. The target prompt text is used to describe the target 3D model. The material generation unit 1002 is used to preprocess the mesh data of the target 3D model to obtain the target noisy image of the target 3D model in N viewpoints; using the target generative model, and based on the target prompt text, the target reference image and the target noisy image of each viewpoint, to obtain M predicted material channel maps under each viewpoint; wherein, the target generative model is trained using the model training method of any of the above embodiments; The rendering unit 1003 is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.
[0139] In a specific example of the disclosed solution, the material generation unit is specifically used for: Based on the mesh data of the target 3D model, the texture image and position image of the target 3D model under each of the N viewpoints are obtained; The texture images, position images, and random noise corresponding to each viewpoint are fused together to obtain noisy target images from each viewpoint.
[0140] In a specific example of the scheme disclosed herein, the rendering unit is specifically used for: The M predicted material channel maps from each viewpoint are mapped from the two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps. Using the M material unfolded diagrams and the mesh data of the target 3D model, rendering processing is performed to obtain a rendered image of the target 3D model with the target material.
[0141] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0142] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0143] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 11As shown, the electronic device includes a memory 1110 and a processor 1120. The memory 1110 stores a computer program that can run on the processor 1120. The number of memories 1110 and processors 1120 can be one or more. The memory 1110 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the method provided in the above-described method embodiments. The electronic device may also include a communication interface 1130 for communicating with external devices and performing data exchange and transmission.
[0144] If the memory 1110, processor 1120, and communication interface 1130 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0145] Optionally, in a specific implementation, if the memory 1110, processor 1120 and communication interface 1130 are integrated on a single chip, the memory 1110, processor 1120 and communication interface 1130 can communicate with each other through an internal interface.
[0146] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0147] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).
[0148] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.
[0149] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0150] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0151] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0152] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0153] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A model training method for a generative model used to reconstruct the material of a 3D model, comprising: Obtain the target training samples; The target training samples include initial prompt text, initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; the reference 3D model is structurally related to the initial 3D model; the initial prompt text is used to describe the initial 3D model; The M initial material channel images from each viewpoint are fused with the initial noise corresponding to each viewpoint to obtain the initial noisy image from each viewpoint. The initial prompt text, the initial reference image, and the initial noisy images from each viewpoint are input into the generative model to be trained, so as to use the feature extraction module in the generative model to obtain the target semantic control features and the target visual prior features corresponding to each viewpoint. Using the prediction head module in the generative model to be trained, and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, M estimated material channel maps and estimated noise corresponding to each viewpoint are obtained. The target semantic control features characterize the geometric structure of the initial 3D model in each viewpoint, and the target visual prior features corresponding to each viewpoint characterize the material distribution information of the initial material of the initial 3D model in each viewpoint. The feature extraction module includes K cascaded diffusion transformation layers and a feature transformation layer for adjusting the semantic dimension and spatial structure of the features. The K diffusion transformation layers process the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint to obtain the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint output by the last diffusion transformation layer. The feature transformation layer processes the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint to obtain the target semantic control feature and the target visual prior feature corresponding to each viewpoint. The target loss value is obtained based on the estimated material channel map and / or estimated noise; Using the target loss value, the feature extraction module and prediction head module in the generative model to be trained are used to train the model, so as to obtain the target generative model.
2. The method according to claim 1, wherein, The process of fusing the M initial material channel images from each viewpoint with the corresponding initial noise from each viewpoint to obtain the initial noisy image for each viewpoint includes: The M initial material channel images under each viewpoint are linearly interpolated with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint.
3. The method according to claim 1, wherein, The method of obtaining the target loss value based on the estimated material channel map and / or estimated noise includes at least one of the following: Based on the M estimated material channel maps from each viewpoint and the M initial material channel maps from each viewpoint, the target loss value is obtained. Based on the estimated noise corresponding to each viewpoint and the initial noise corresponding to each viewpoint, the target loss value is obtained; Based on the M initial material channel images from each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image from each viewpoint, the target loss value is obtained.
4. The method according to claim 3, wherein, The target loss value is obtained based on M initial material channel maps from each viewpoint, the estimated noise corresponding to each viewpoint, and the initial noisy image from each viewpoint, including: Linear interpolation is performed on the M initial material channel images under each viewpoint and the estimated noise corresponding to each viewpoint to obtain the estimated noisy image of each viewpoint. The target loss value is obtained based on the estimated noisy image from each viewpoint and the initial noisy image from each viewpoint.
5. The method according to any one of claims 1-4, wherein, The step of using the feature extraction module in the generative model to be trained to obtain target semantic control features and target visual prior features corresponding to each viewpoint includes: Using the K diffusion transform layers, and combining the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint, the Kth semantic control feature output by the last diffusion transform layer in the K diffusion transform layers and the Kth visual prior feature corresponding to each viewpoint are obtained. By utilizing the feature transformation layer and combining the Kth semantic control feature with the Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
6. The method according to claim 5, wherein, The feature extraction module further includes a normalization layer; the method further includes: By utilizing the normalization layer and combining the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint, the normalized Kth semantic control feature and the normalized Kth visual prior feature corresponding to each viewpoint are obtained. The step of utilizing the feature transformation layer, combined with the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint, to obtain the target semantic control feature and the target visual prior feature corresponding to each viewpoint includes: By utilizing the aforementioned feature transformation layer, and combining the normalized Kth semantic control feature with the normalized Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
7. The method according to claim 5, wherein, The i-th diffusion transform layer in the K diffusion transform layers includes: an i-th fully connected layer for obtaining the weight matrix corresponding to each data path in the multi-path data; an i-th sequence concatenation block for performing sequence concatenation processing on each weight matrix; an i-th fully self-attention block for capturing global dependency information between all elements in the features after sequence concatenation processing; an i-th text linear transformer for extracting global semantic information from the features after fully self-attention processing; and an i-th image linear transformer for extracting global visual information from the features after fully self-attention processing.
8. The method according to claim 7, wherein, The i-th fully connected layer includes multiple fully connected blocks arranged in parallel, wherein each fully connected block corresponds to one data path, so as to obtain multiple weight matrices corresponding to a single data path.
9. A method for generating material renderings of a 3D model, comprising: Obtain the mesh data of the target 3D model, the target reference image, and the target prompt text, wherein the target reference image is a 2D image of the reference 3D model with the target material; The reference 3D model is structurally related to the target 3D model; the target prompt text is used to describe the target 3D model; The mesh data of the target 3D model is preprocessed to obtain noisy images of the target 3D model from each of the N viewpoints; Using a target generative model, and based on the target prompt text, the target reference image, and the target noisy image from each viewpoint, M predicted material channel images are obtained from each viewpoint; wherein, the target generative model is trained using the model training method described in any one of claims 1 to 8; Based on the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, a rendering effect map of the target 3D model with the target material is obtained.
10. The method according to claim 9, wherein, The preprocessing of the mesh data of the target 3D model to obtain noisy images of the target 3D model from N viewpoints includes: Based on the mesh data of the target 3D model, the texture image and position image of the target 3D model under each of the N viewpoints are obtained; The texture images, position images, and random noise corresponding to each viewpoint are fused together to obtain noisy target images from each viewpoint.
11. The method according to claim 9 or 10, wherein, The rendering effect of the target 3D model with the target material is obtained by using the mesh data of the target 3D model and the M predicted material channel maps from each viewpoint, including: The M predicted material channel maps from each viewpoint are mapped from the two-dimensional plane to the three-dimensional space represented by the mesh data of the target three-dimensional model to obtain M material unfolded maps. Using the M material unfolded diagrams and the mesh data of the target 3D model, rendering processing is performed to obtain a rendered image of the target 3D model with the target material.
12. A model training device for reconstructing generative models of 3D model materials, comprising: The sample acquisition unit is used to acquire target training samples; The target training samples include initial prompt text, initial reference image, initial noise corresponding to each of the N viewpoints, and M initial material channel images of the initial 3D model determined based on the initial material in each of the N viewpoints; the initial reference image is a two-dimensional image of the reference 3D model with the initial material; The reference 3D model is structurally related to the initial 3D model; the initial prompt text is used to describe the initial 3D model; The model training unit is used to fuse the M initial material channel images under each viewpoint with the initial noise corresponding to each viewpoint to obtain the initial noisy image of each viewpoint; the initial prompt text, the initial reference image and the initial noisy image of each viewpoint are input into the generative model to be trained, so as to use the feature extraction module in the generative model to be trained to obtain the target semantic control features and the target visual prior features corresponding to each viewpoint. Using the prediction head module in the generative model to be trained, and combining the target semantic control features and the target visual prior features corresponding to each viewpoint, M estimated material channel maps and estimated noise corresponding to each viewpoint are obtained. The target semantic control features are used to characterize the geometric structure of the initial 3D model in each viewpoint, and the target visual prior features corresponding to each viewpoint are used to characterize the material distribution information of the initial material of the initial 3D model in each viewpoint. The feature extraction module includes K cascaded diffusion transformation layers and a feature transformation layer for adjusting the semantic dimension and spatial structure of the features. The transformation layer processes the initial prompt text, initial reference image, and initial noisy images corresponding to each viewpoint to obtain the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint output by the last diffusion transformation layer. The feature transformation layer processes the Kth semantic control feature and the Kth visual prior feature corresponding to each viewpoint to obtain the target semantic control feature and the target visual prior feature corresponding to each viewpoint. Based on the estimated material channel map and / or estimated noise, the target loss value is obtained. Using the target loss value, the feature extraction module and prediction head module in the generative model to be trained are used to train the model to obtain the target generative model.
13. The apparatus according to claim 12, wherein, The model training unit is specifically used for: Using the K diffusion transform layers, and combining the initial prompt text, the initial reference image, and the initial noisy image corresponding to each viewpoint, the Kth semantic control feature output by the last diffusion transform layer in the K diffusion transform layers and the Kth visual prior feature corresponding to each viewpoint are obtained. By utilizing the feature transformation layer and combining the Kth semantic control feature with the Kth visual prior feature corresponding to each viewpoint, the target semantic control feature and the target visual prior feature corresponding to each viewpoint are obtained.
14. The apparatus according to claim 13, wherein, The i-th diffusion transform layer in the K diffusion transform layers includes: an i-th fully connected layer for obtaining the weight matrix corresponding to each data path in the multi-path data; an i-th sequence concatenation block for performing sequence concatenation processing on each weight matrix; an i-th fully self-attention block for capturing global dependency information between all elements in the features after sequence concatenation processing; an i-th text linear transformer for extracting global semantic information from the features after fully self-attention processing; and an i-th image linear transformer for extracting global visual information from the features after fully self-attention processing.
15. The apparatus according to claim 14, wherein, The i-th fully connected layer includes multiple fully connected blocks arranged in parallel, wherein each fully connected block corresponds to one data path, so as to obtain multiple weight matrices corresponding to a single data path.
16. A device for generating material renderings of a three-dimensional model, comprising: The data input unit is used to acquire the mesh data of the target 3D model, the target reference image, and the target prompt text, wherein the target reference image is a 2D image of the reference 3D model with the target material; The reference 3D model is structurally related to the target 3D model; the target prompt text is used to describe the target 3D model; The material generation unit is used to preprocess the mesh data of the target 3D model to obtain the target 3D model with noise from each of the N viewpoints. Using a target generative model, and based on the target prompt text, the target reference image, and the target noisy image from each viewpoint, M predicted material channel images are obtained from each viewpoint; wherein, the target generative model is trained using the model training method described in any one of claims 1 to 8; The rendering unit is used to obtain a rendering effect diagram of the target 3D model with the target material based on the mesh data of the target 3D model and M predicted material channel maps from each viewpoint.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-8 or 9-11.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8 or 9-11.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8 or 9-11.