Data processing method and device, computer equipment and readable storage medium

By combining multi-view image stitching and texture control data, fine material images are generated, which solves the problem of insufficient control of spatial objects in the prior art, and improves the accuracy of spatial objects after baking in multi-view angles.

CN120219218APending Publication Date: 2025-06-27SHENZHEN TENCENT COMP SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510287273.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art lacks fine control of spatial objects in material generation, resulting in reduced accuracy of spatial objects after texture mapping.

Method used

By obtaining geometric images of spatial objects at multiple perspectives, image stitching is performed to generate control images, and denoising the noise image with texture control data, generating predicted images and material images, and finally performing multi-view baking.

Benefits of technology

Improve the accuracy of spatial objects after baking in multiple perspectives, so that their textures and texture control data are exactly matched, and enhance the multi-view consistency of material images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219218A_ABST
    Figure CN120219218A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, computer equipment and a readable storage medium, and the method comprises the steps: obtaining a geometric image of a space object at each of S visual angles, carrying out the image splicing of the S geometric images, and obtaining a control image; obtaining texture control data, and performing denoising processing on the noise image according to the texture control data and the control image to obtain a prediction image corresponding to the space object; the prediction image comprises a sub-prediction image of the space object at each view angle; performing denoising processing on the prediction image to obtain a material image corresponding to the space object; the material image comprises a sub-material image of the space object at each view angle; performing multi-view baking on the space object according to the material image to obtain the space object after multi-view baking; the spatial object after multi-view baking has the texture indicated by the texture control data. According to the invention, the accuracy of the space object after multi-view baking can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, computer device, and readable storage medium. Background Art

[0002] Existing material generation solutions can obtain texture control data for guiding the generation of PBR (Physically-Based Rendering) materials (for example, text data for guiding the generation of PBR materials), denoise a noise image according to the texture control data to obtain a material image corresponding to a three-dimensional model (for example, a spatial object), and perform texture mapping on the spatial object according to the material image so that the texture-mapped spatial object has the texture indicated by the texture control data. For example, the texture control data can be "a bright ax", the spatial object can be "an ax", and the texture-mapped spatial object can be "a bright ax". However, the texture control data lacks fine control over the spatial object, so the material image generated according to the texture control data cannot fully match the spatial object, thereby reducing the accuracy of the texture-mapped spatial object. Summary of the Invention

[0003] Embodiments of this application provide a data processing method, apparatus, computer device, and readable storage medium, which can improve the accuracy of a spatial object after multi-view baking.

[0004] One aspect of an embodiment of this application provides a data processing method, including:

[0005] Obtain geometric images of a spatial object at each of S perspectives, and splice the S geometric images to obtain a control image; S is an integer greater than 1;

[0006] Obtain texture control data, and denoise a noise image according to the texture control data and the control image to obtain a prediction image corresponding to the spatial object; the prediction image includes sub-prediction images of the spatial object at each perspective;

[0007] Denoise the prediction image to obtain a material image corresponding to the spatial object; the material image includes sub-material images of the spatial object at each perspective;

[0008] Perform multi-view baking on the spatial object according to the material image to obtain a multi-view baked spatial object; the multi-view baked spatial object has the texture indicated by the texture control data.

[0009] One aspect of an embodiment of this application provides a data processing apparatus, including:

[0010] An image acquisition module, configured to acquire geometric images of a spatial object from each of S perspectives, perform image stitching on the S geometric images to obtain a control image; S is an integer greater than 1;

[0011] A first denoising module, configured to acquire texture control data, perform denoising processing on a noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object; the predicted image includes sub-predicted images of the spatial object from each perspective;

[0012] A second denoising module, configured to perform denoising processing on the predicted image to obtain a material image corresponding to the spatial object; the material image includes sub-material images of the spatial object from each perspective;

[0013] A multi-view baking module, configured to perform multi-view baking on the spatial object according to the material image to obtain a multi-view baked spatial object; the multi-view baked spatial object has the texture indicated by the texture control data.

[0014] Among them, the first denoising module is specifically configured to acquire a texture control feature corresponding to the texture control data, a control image feature corresponding to the control image, and a noise image feature corresponding to the noise image;

[0015] The first denoising module is specifically configured to perform denoising processing on the noise image feature according to the texture control feature and the control image feature in a first diffusion model to obtain a denoised image feature;

[0016] The first denoising module is specifically configured to perform decoding processing on the denoised image feature to obtain a predicted image corresponding to the spatial object.

[0017] Among them, the first denoising module is specifically configured to acquire a denoising time step feature corresponding to the denoising time step of the noise image feature;

[0018] The first denoising module is specifically configured to input the noise image feature, the denoising time step feature, and the texture control feature into a denoising network of the first diffusion model, input the noise image feature, the denoising time step feature, the texture control feature, and the control image feature into a control network of the first diffusion model, and predict the noise in the noise image feature through the denoising network and the control network;

[0019] The first denoising module is specifically configured to determine the denoised image feature according to the noise in the noise image feature, the noise image feature, and the noise intensity coefficient of the denoising network.

[0020] Among them, the second denoising module is specifically configured to acquire a denoised image feature corresponding to the predicted image;

[0021] The second denoising module is specifically configured to perform denoising processing on the denoised image feature in a second diffusion model to obtain a material denoised image feature;

[0022] The second denoising module is specifically configured to decode the material denoised image features to obtain the material image corresponding to the spatial object.

[0023] Among them, the number of material images is K, where K is a positive integer. The K material images respectively correspond to different physical materials. The material image corresponding to each physical material includes the sub-material images of the spatial object at each viewing angle. Each material image is used to perform multi-view baking on the spatial object to obtain the multi-view baked spatial object for the physical material corresponding to each material image; the K material images include at least one of a color material image, a metal material image, or a rough material image.

[0024] Among them, the multi-view baking module is specifically configured to randomly initialize the texture map of the spatial object to obtain the rendered texture map corresponding to the spatial object;

[0025] The multi-view baking module is specifically configured to sample the rendered texture map according to S viewing angles to obtain the rendered images at each viewing angle;

[0026] The multi-view baking module is specifically configured to obtain the viewing angle weights corresponding to the S viewing angles respectively, and determine the viewing angle loss values corresponding to the S viewing angles respectively according to the S rendered images, the S sub-material images in the material images, and the S viewing angle weights;

[0027] The multi-view baking module is specifically configured to perform a summation operation on the S viewing angle loss values to obtain the image loss value corresponding to the rendered texture map;

[0028] The multi-view baking module is specifically configured to adjust the rendered texture map according to the image loss value to obtain the material texture map corresponding to the spatial object; the material texture map matches the physical material corresponding to the material image;

[0029] The multi-view baking module is specifically configured to perform texture mapping on the spatial object according to the material texture map to obtain the multi-view baked spatial object.

[0030] Among them, the S viewing angles include viewing angle G i , where i is a positive integer less than or equal to S;

[0031] The multi-view baking module is specifically configured to obtain L object vertices of the spatial object; L is a positive integer greater than 1;

[0032] The multi-view baking module is specifically configured to determine the vertex weights of the L object vertices respectively for the viewing angle G i corresponding to the camera optical center position and the normal vectors corresponding to the L object vertices respectively according to the vertex coordinates corresponding to the L object vertices, the viewing angle G i ;

[0033] The multi-view baking module is specifically used to perform a summation operation on L vertex weights to obtain view G i The corresponding view weights.

[0034] Among them, the S views include view G i , where i is a positive integer less than or equal to S;

[0035] The multi-view baking module is specifically used for the rendered image under view G i and the sub-material image under view G i to perform a difference operation to obtain the difference image corresponding to view G i ;

[0036] The multi-view baking module is specifically used to perform an absolute value operation on the difference image corresponding to view G i to obtain the absolute value image corresponding to view G i ;

[0037] The multi-view baking module is specifically used to perform a multiplication operation on the view weight corresponding to view G i and the absolute value image corresponding to view G i to obtain the view loss value corresponding to view G i ;

[0038] Among them, the device further includes:

[0039] The first training module is used to obtain a second initial diffusion model, sample texture control data, and sample geometric images of a sample object under each of the S views, and perform image stitching on the S sample geometric images to obtain a sample control image;

[0040] The first training module is used to perform denoising processing on a noise image according to the sample texture control data and the sample control image to obtain a sample prediction image corresponding to the sample object; the sample prediction image includes sub-sample prediction images of the sample object under each view;

[0041] The first training module is used to input the sample prediction image features corresponding to the sample prediction image into the second initial diffusion model, and perform denoising processing on the sample prediction image through the second initial diffusion model to obtain a sample material image corresponding to the sample object; the sample material image includes sub-sample material images of the sample object under each view;

[0042] The first training module is used to perform light estimation on the sample prediction image to obtain the image light information of the sample prediction image, and perform image rendering on the image light information and the sample material image to obtain a sample rendering image corresponding to the sample object; the sample rendering image includes sub-sample rendering images of the sample object under each view;

[0043] The first training module is used to adjust the parameters of the second initial diffusion model according to the sample rendering image and the sample prediction image, so as to obtain the second diffusion model.

[0044] Wherein, the device further includes:

[0045] The second training module is used to obtain an initial denoising network, sample texture control data, a sample image, and a set of time steps; the set of time steps includes at least two time steps; the sample image has the texture indicated by the sample texture control data;

[0046] The second training module is used to perform random uniform sampling on the set of time steps to obtain the noisy time steps among at least two time steps;

[0047] The second training module is used to determine the noisy image features corresponding to the noisy time steps according to the sample image features corresponding to the sample image, the noisy Gaussian noise corresponding to the sample image features, and the noisy time steps; the noisy Gaussian noise follows a normal distribution;

[0048] The second training module is used to input the noisy image features, the sample texture control features corresponding to the sample texture control data, and the noisy time step features corresponding to the noisy time steps into the initial denoising network, predict the noise in the noisy image features through the initial denoising network, and determine the noise in the noisy image features as the first noise;

[0049] The second training module is used to adjust the parameters of the initial denoising network according to the first noise and the noisy Gaussian noise to obtain the denoising network.

[0050] Wherein, the device further includes:

[0051] The third training module is used to obtain the sample geometric images of the sample object from each of the S perspectives, and splice the S sample geometric images to obtain the sample control image;

[0052] The third training module is used to determine the initial control network according to the denoising network, input the noisy image features, the sample texture control features, and the noisy time step features into the denoising network, input the noisy image features, the sample texture control features, the noisy time step features, and the sample control image features corresponding to the sample control image into the initial control network, predict the noise in the noisy image features through the denoising network and the initial control network, and determine the noise in the noisy image features as the second noise;

[0053] The third training module is used to adjust the parameters of the initial control network according to the second noise and the noisy Gaussian noise to obtain the control network.

[0054] On the one hand, an embodiment of the present application provides a computer device, including: a processor and a memory;

[0055] The processor is connected to a memory, where the memory is used to store a computer program. When the computer program is executed by the processor, the computer device is caused to execute the method provided by the embodiments of the present application.

[0056] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiments of the present application.

[0057] On the one hand, an embodiment of the present application provides a computer program product including a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to execute the method provided by the embodiments of the present application.

[0058] Embodiments of the present application can generate geometric images for each of the S perspectives for the geometric information of the spatial object, splice the S geometric images into a multi-perspective image (i.e., a control image) for the geometric information, perform denoising processing on the noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object, and further perform denoising processing on the predicted image to obtain a material image corresponding to the spatial object. Perform multi-perspective baking on the spatial object according to the material image to obtain a multi-perspective baked spatial object, where the multi-perspective baked spatial object has the texture indicated by the texture control data. Therefore, the present application can introduce geometric images for describing the geometric information of the spatial object. Since the geometric information can achieve fine control of the spatial object, the material image generated according to the texture control data and the geometric image can perfectly match the spatial object, thereby improving the accuracy of the multi-perspective baked spatial object. In addition, the present application can use the geometric images for each of the S perspectives simultaneously instead of using the geometric image under a single perspective. The geometric images for each perspective can learn from each other, thereby avoiding the perspective inconsistency (e.g., the inconsistency in the overlapping area) when generating sub-material images for each of the S perspectives, and ensuring the multi-perspective consistency of the sub-material images for each perspective in the material image. Description of the Drawings

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for the description of the embodiments or the related art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0060] Figure 1It is a schematic structural diagram of a network architecture provided by an embodiment of the present application;

[0061] Figure 2 It is a schematic diagram of a scenario for data interaction provided by an embodiment of the present application;

[0062] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0063] Figure 4 It is a schematic diagram of a scenario for perspective control provided by an embodiment of the present application;

[0064] Figure 5 It is a schematic structural diagram of a texture model provided by an embodiment of the present application;

[0065] Figure 6 It is a schematic diagram of a scenario for texture baking provided by an embodiment of the present application;

[0066] Figure 7 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0067] Figure 8 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0068] Figure 9 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0069] Figure 10 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0071] Among them, the solutions provided by the embodiments of the present application mainly involve natural language processing (NLP) technology, machine learning (ML) technology, and computer vision (CV) technology of artificial intelligence (AI).

[0072] Specifically, please refer to Figure 1 , Figure 1It is a schematic structural diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a server 2000 and a cluster of terminal devices. Among them, the cluster of terminal devices may specifically include one or more terminal devices, and the number of terminal devices in the cluster of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include terminal device 3000a, terminal device 3000b, terminal device 3000c, …, terminal device 3000n; terminal device 3000a, terminal device 3000b, terminal device 3000c, …, terminal device 3000n may be directly or indirectly network-connected to the server 2000 through wired or wireless communication methods, so that each terminal device can perform data interaction with the server 2000 through this network connection.

[0073] Among them, the terminal devices in the cluster of terminal devices may include: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances (such as smart TVs), wearable devices, vehicle-mounted terminals, aircraft and other intelligent terminals with data processing functions. Among them, the vehicle-mounted terminal may be a terminal device in the intelligent transportation scenario and the assisted driving scenario. For ease of understanding, an embodiment of the present application may Figure 1 select one terminal device from the multiple terminal devices shown as the target terminal device. For example, an embodiment of the present application may use Figure 1 the terminal device 3000a shown as the target terminal device.

[0074] Among them, the server 2000 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0075] It can be understood that the data processing method provided by the embodiments of the present application can be executed by a computer device. Given a three-dimensional model with or without texture, the computer device can use the texture control data input by the user as control information (or prompt information, control signal), and generate texture information (referred to as texture for short) for the three-dimensional model through a diffusion model (a generative model), so as to achieve multi-view baking of the three-dimensional model. Among them, the texture of the present application can be the texture of a PBR material (the PBR material refers to a material designed and optimized for physically based rendering, aiming to accurately simulate the physical behavior of light sources and materials to achieve a realistic effect), and the multi-view baking of the three-dimensional model can include the white film coloring function of the three-dimensional model (that is, generating a texture for a three-dimensional model without texture. Here, it is illustrated by taking the three-dimensional model without texture as white) or the texture replacement function of the three-dimensional model (that is, generating a texture for a three-dimensional model with texture, and the original texture of the three-dimensional model is different from the texture generated for the three-dimensional model). In other words, the embodiments of the present application can be executed by the server 2000 (that is, the computer device can be the server 2000), or can be executed by the target terminal device (that is, the computer device can be the target terminal device), or can also be jointly executed by the server 2000 and the target terminal device.

[0076] Among them, the three-dimensional model is used to represent real-world or fictional objects in three-dimensional space. The geometric basis of the three-dimensional model is a mesh, and the mesh is used to define the surface structure of the three-dimensional model (that is, the mesh is used to represent the geometric shape of the three-dimensional model). The present application does not limit the shape of the mesh (for example, the shape of the mesh can be triangular), and the present application does not limit the specific object of the three-dimensional model (for example, the three-dimensional model can be an animal, a building, a mechanical device, a game prop). For the convenience of understanding, the embodiments of the present application can refer to the three-dimensional model here as a spatial object or a sample object. The spatial object can be the three-dimensional model used by the diffusion model in the inference stage, and the sample object can be the three-dimensional model used by the diffusion model in the training stage. It should be understood that the present application does not limit the model type of the diffusion model. For example, the diffusion model can be a StableDiffusion (abbreviated as SD) model, a Latent Diffusion model. For the convenience of understanding, the embodiments of the present application are illustrated by taking the diffusion model as the Stable Diffusion model.

[0077] Among them, the texture of a 3D model refers to a 2D image mapped on the surface of the 3D model, which is used to enhance the visual effect of the 3D model and make it look more real and detailed. For example, the texture can include attributes such as color, smoothness, and reflectivity. The multi-view baking of a 3D model refers to attaching the texture to the 3D model through a texture image or other means during 3D modeling and rendering. Multi-view baking can include texture baking and texture mapping. Texture baking refers to pre-computing complex material effects and storing them as 2D texture images (i.e., material texture maps, texture maps, texture information). Texture mapping refers to the process of pasting a 2D texture image onto the surface of a 3D model. Among them, texture mapping requires the use of texture coordinates (i.e., UV coordinates, 2D coordinates) to determine how each pixel on the texture image is associated with the vertices of the 3D model. The texture coordinates are stored together with the vertices of the 3D model and are used for interpolation and calculation of texture coordinates during the rendering process.

[0078] It should be understood that the above network framework can be applied to multi-view baking scenarios. The specific services of multi-view baking scenarios can include game development services, architectural design services, animation production services, virtual reality services, etc. Here, the specific services will not be listed one by one. For example, in game development services, multi-view baking can create the appearance of various scenes, characters, and props, adding rich details to the virtual world in the game (e.g., brick texture of the wall, grass texture of the ground, skin texture of the character), enabling game players to obtain a more realistic visual experience. In architectural design services, multi-view baking can help designers quickly create the appearance effect of buildings, showing details such as walls, roofs, and windows of different materials, allowing customers to more intuitively view the final effect of the design plan. In animation production services, multi-view baking can add rich details to characters and scenes (e.g., adding skin texture and clothing texture to the model of an animated character), improving the quality and visual effect of the animation, making the animated character more vivid and enhancing the expressiveness of the animation. In virtual reality services, multi-view baking can add various realistic textures to the objects in the virtual scene, enabling users to feel a more real visual experience in the virtual world.

[0079] It should be understood that the collection and processing of relevant data (such as spatial objects, sample objects, texture control data, sample texture control data) in this application should strictly comply with the requirements of relevant laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.

[0080] For easy understanding, please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario for data interaction provided by an embodiment of this application. As Figure 2The server 20a shown can be the server 2000 in the corresponding embodiment described above, such as Figure 1 The terminal device 20b shown can be the target terminal device in the corresponding embodiment described above. The user corresponding to the terminal device 20b can be the user 20c, that is, the user 20c can be the user using the terminal device 20b. For ease of understanding, the embodiment of the present application is described by taking the data processing method as being executed by the server 20a as an example. Figure 2 As shown, when the user 20c needs to implement multi-view baking of a 3D model, the user 20c can perform a selection operation and an input operation on the terminal device 20b. The selection operation represents selecting a 3D model (here, the 3D model selected by the user 20c on the terminal device 20b is used as an example of a spatial object for illustration). The input operation represents inputting texture control data (the texture control data is used to indicate the texture generated for the spatial object. The embodiment of the present application does not limit the type of the texture control data. Here, the texture control data can be text data, audio data, or image data). In this way, the terminal device 20b can respond to the selection operation and the input operation performed by the user 20c, obtain the spatial object selected by the selection operation and the texture control data input by the input operation, and send the spatial object and the texture control data to the server 20a. Figure 1 As shown, the server 20a can receive the spatial object and the texture control data sent by the terminal device 20b, and obtain the geometric images of the spatial object under each of the S perspectives. Here, S can be an integer greater than 1, and the S perspectives can cover as much visible area as possible. For example, when S is equal to 4, the S perspectives can be front, back, left, and right. When S is equal to 6, the S perspectives can be front, back, left, right, up (i.e., the top), and down (i.e., the bottom). For ease of understanding, the embodiment of the present application is described by taking S equal to 6 as an example. The 6 geometric images can include the geometric image 21a (i.e., the geometric image under the front perspective), the geometric image 21b (i.e., the geometric image under the back perspective), the geometric image 21c (i.e., the geometric image under the left perspective), the geometric image 21d (i.e., the geometric image under the right perspective), the geometric image 21e (i.e., the geometric image under the up perspective), and the geometric image 21f (i.e., the geometric image under the down perspective).

[0081] Such as Figure 2 As shown, when the user 20c needs to implement multi-view baking of a 3D model, the user 20c can perform a selection operation and an input operation on the terminal device 20b. The selection operation represents selecting a 3D model (here, the 3D model selected by the user 20c on the terminal device 20b is used as an example of a spatial object for illustration). The input operation represents inputting texture control data (the texture control data is used to indicate the texture generated for the spatial object. The embodiment of the present application does not limit the type of the texture control data. Here, the texture control data can be text data, audio data, or image data). In this way, the terminal device 20b can respond to the selection operation and the input operation performed by the user 20c, obtain the spatial object selected by the selection operation and the texture control data input by the input operation, and send the spatial object and the texture control data to the server 20a.

[0082] Such as Figure 2 As shown, the server 20a can receive the spatial object and the texture control data sent by the terminal device 20b, and obtain the geometric images of the spatial object under each of the S perspectives. Here, S can be an integer greater than 1, and the S perspectives can cover as much visible area as possible. For example, when S is equal to 4, the S perspectives can be front, back, left, and right. When S is equal to 6, the S perspectives can be front, back, left, right, up (i.e., the top), and down (i.e., the bottom). For ease of understanding, the embodiment of the present application is described by taking S equal to 6 as an example. The 6 geometric images can include the geometric image 21a (i.e., the geometric image under the front perspective), the geometric image 21b (i.e., the geometric image under the back perspective), the geometric image 21c (i.e., the geometric image under the left perspective), the geometric image 21d (i.e., the geometric image under the right perspective), the geometric image 21e (i.e., the geometric image under the up perspective), and the geometric image 21f (i.e., the geometric image under the down perspective).

[0083] Furthermore, the server 20a can perform image stitching on S geometric images (i.e., geometric image 21a, geometric image 21b, geometric image 21c, geometric image 21d, geometric image 21e, and geometric image 21f) (image stitching refers to stitching the S geometric images in the vertical dimension, and the order of stitching the S geometric images is not limited in this application) to obtain a control image. The control image includes the geometric images of the spatial object at each perspective, that is, the S geometric images are stitched in the vertical dimension to obtain the control image.

[0084] Furthermore, the server 20a can, in the first diffusion model, perform denoising processing on a noise image (the noise image is a randomly sampled Gaussian noise image that follows a normal distribution, and the noise image includes random noise that follows a normal distribution) according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object. The predicted image includes sub-predicted images of the spatial object at each perspective. The predicted image is an RGB (Red, Green, Blue) image of S perspectives. The predicted image conforms to both the texture control data and the control image at the same time, that is, the S sub-predicted images are stitched in the vertical dimension to obtain the predicted image, and the S sub-predicted images are associated with the S geometric images.

[0085] Furthermore, the server 20a can, in the second diffusion model, perform denoising processing on the predicted image to obtain a material image corresponding to the spatial object. The material image includes sub-material images of the spatial object at each perspective. The material image is a PBR material image of S perspectives. The material image is used to represent the PBR material, that is, the S sub-material images are stitched in the vertical dimension to obtain the material image, and the S sub-material images are associated with the S geometric images.

[0086] Furthermore, the server 20a can perform multi-view baking on the spatial object according to the material image (that is, baking the material image back to the 3D model) to obtain the spatially object after multi-view baking, and return the spatially object after multi-view baking to the terminal device 20b so that the terminal device 20b can display the spatial object and the spatially object after multi-view baking at the same time. The spatially object after multi-view baking has the texture indicated by the texture control data.

[0087] It can be seen that the embodiments of the present application can introduce a geometric image for representing the geometric information of a spatial object on the basis of texture control data, and jointly control the generation of a material image according to the texture control data and the geometric image. Among them, the texture control data can be used for macroscopic control of the material image, and the geometric image can be used for fine control of the material image, so that the material image can completely match the spatial object. In this way, when multi-view baking is performed on the spatial object according to the material image, the accuracy of the spatial object after multi-view baking can be improved. In addition, the present application can obtain the geometric images of the spatial object under each of the S views at the same time, rather than the geometric image under a single view. The geometric images under multiple views can learn from each other to achieve accurate control of the spatial object and improve the multi-view consistency of the sub-material images in the material image.

[0088] Further, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. This data processing method can be executed by a computer device, and the computer device can be the server 20a in the corresponding embodiment of the above Figure 2 or the terminal device 20b in the corresponding embodiment of the above Figure 2 . Among them, this data processing method can include steps S101-S104:

[0089] Step S101, obtain the geometric images of the spatial object under each of the S views, and perform image stitching on the S geometric images to obtain a control image;

[0090] Among them, S here can be an integer greater than 1. The S views refer to the views respectively corresponding to S cameras, and the positions of the S cameras are different from each other. The geometric image under each view is determined by the camera respectively corresponding to each view; the geometric image (abbreviated as geometric map) is used to represent the geometric information of the spatial object, and the geometric images of the spatial object under each of the S views are used to represent the geometric information of the spatial object under each of the S views (for example, depth, position, normal). Different geometric images can guide the generation of different spatial objects after multi-view rendering, so that the spatial object after multi-view rendering meets the geometric information represented by the geometric image.

[0091] It should be understood that the present application does not limit the type of geometric images. For example, the geometric image can be a depth map, an XYZ map, or a normal map. The geometric information represented by the depth map is depth, the geometric information represented by the XYZ map is position, and the geometric information represented by the normal map is normal. In other words, the depth map is used to represent the distance (depth value) of a pixel relative to the camera in three-dimensional space, the XYZ map is used to represent the actual position of a pixel in three-dimensional space (i.e., the X coordinate, Y coordinate, and Z coordinate), and the normal map is used to represent the surface normal direction of a pixel in three-dimensional space (usually represented by a unit vector).

[0092] In other words, the computer device can obtain the depth map of the spatial object at each of the S viewpoints, perform image stitching on the S depth maps to obtain a control image; or, the computer device can obtain the XYZ map of the spatial object at each of the S viewpoints, perform image stitching on the S XYZ maps to obtain a control image; or, the computer device can obtain the normal map of the spatial object at each of the S viewpoints, perform image stitching on the S normal maps to obtain a control image.

[0093] Step S102: Obtain texture control data, and perform denoising processing on the noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object;

[0094] Specifically, the computer device can obtain texture control data, and obtain the texture control features corresponding to the texture control data, the control image features corresponding to the control image, and the noise image features corresponding to the noise image. Among them, the texture control features are obtained by encoding the texture control data, the control image features are obtained by encoding the control image, and the noise image features are obtained by encoding the noise image. The embodiments of the present application do not limit the encoder used for the encoding process. For example, when the texture control data is image data, the encoder for encoding the texture control data can be the Image Encoder of the CLIP (Contrastive Language-Image Pre-Training) model; when the texture control data is text data, the encoder for encoding the texture control data can be the Text Encoder of the CLIP model; the encoder for encoding the control image and the encoder for encoding the noise image can be the encoder of the Variational Autoencoder (VAE). Further, the computer device can perform denoising processing on the noise image features according to the texture control features and the control image features in the first diffusion model to obtain denoised image features. Further, the computer device can perform decoding processing on the denoised image features to obtain a predicted image corresponding to the spatial object. Among them, the predicted image includes sub-predicted images of the spatial object at each perspective; the embodiments of the present application do not limit the decoder used for the decoding process. For example, the decoder for decoding the denoised image features can be the decoder of the variational autoencoder, and the decoder of the variational autoencoder and the encoder of the variational autoencoder are a matching encoder and decoder.

[0095] Optionally, the computer device can obtain the geometric image features corresponding to S geometric images respectively, and perform splicing processing on the S geometric image features to obtain control image features. Among them, the geometric image features are obtained by encoding the geometric images. The embodiments of the present application do not limit the encoder used for the encoding process. For example, the encoder for encoding the geometric images can be the encoder of the variational autoencoder.

[0096] For ease of understanding, please refer to Figure 4 , Figure 4 which is a schematic diagram of a perspective control scenario provided by the embodiments of the present application. As Figure 4As shown, image 40a is a control image (i.e., control image 40a), and image 40b is a predicted image (i.e., predicted image 40b). For ease of understanding, here, a geometric image (or control image 40a) is taken as an example of a depth map for explanation. Here, it is taken as an example that the S geometric images in the control image 40a are not stitched in the longitudinal dimension. Here, it is taken as an example that the S sub-predicted images in the predicted image 40b are not stitched in the longitudinal dimension. In fact, the S geometric images are stitched in the control image 40a through the longitudinal dimension, and the S sub-predicted images are stitched in the predicted image 40b through the longitudinal dimension. The longitudinal dimension refers to the token dimension in the latent (i.e., the features of the latent space). Among them, since the control image 40a is a depth map, the edge of the predicted image 40b conforms to the edge of the depth map, and the details of the predicted image 40b conform to the change trend of the depth map.

[0097] Among them, as Figure 4 shown, here, it is taken as an example that S is equal to 6 for explanation. The S sub-predicted images are associated with the S geometric images. The 6 geometric images can include geometric image 41a, geometric image 41b, geometric image 41c, geometric image 41d, geometric image 41e, and geometric image 41f. The 6 sub-predicted images can include sub-predicted image 42a, sub-predicted image 42b, sub-predicted image 42c, sub-predicted image 42d, sub-predicted image 42e, and sub-predicted image 42f. Geometric image 41a and sub-predicted image 42a can be images under the front view angle (i.e., geometric image 41a and sub-predicted image 42a are associated). Geometric image 41b and sub-predicted image 42b (i.e., geometric image 41b and sub-predicted image 42b are associated) can be images under the right view angle. Geometric image 41c and sub-predicted image 42c (i.e., geometric image 41c and sub-predicted image 42c are associated) can be images under the rear view angle. Geometric image 41d and sub-predicted image 42d (i.e., geometric image 41d and sub-predicted image 42d are associated) can be images under the left view angle. Geometric image 41e and sub-predicted image 42e (i.e., geometric image 41e and sub-predicted image 42e are associated) can be images under the lower view angle. Geometric image 41f and sub-predicted image 42f (i.e., geometric image 41f and sub-predicted image 42f are associated) can be images under the upper view angle.

[0098] It can be understood that the computer device can perform M denoising processes on the noise image features in the first diffusion model in sequence according to the texture control features and the control image features to obtain the denoised image features, and perform decoding processing on the denoised image features obtained from the Mth denoising process to obtain the predicted image corresponding to the spatial object. Among them, M here can be a positive integer. For ease of understanding, this application takes M greater than 1 as an example for illustration. The first diffusion model can include M denoising networks and M control networks. One denoising network corresponds to one control network. One denoising network and one control network are used to implement one denoising process. The M denoising networks are the same, and the M control networks are the same. This application embodiment does not limit the network types of the denoising network and the control network. For example, the denoising network can be a U-Net (U Network) network, and the control network can be a ControlNet network; the denoising time steps (i.e., time steps, time step lengths) corresponding to the M denoising processes continuously decrease. The denoising time step is a positive integer less than or equal to the first time step threshold. The first time step threshold is used to determine the number of denoising processes performed by the first diffusion model (for example, the denoising time steps corresponding to the M denoising processes are adjacent, and the number of denoising processes is equal to the first time step threshold). This application embodiment does not limit the specific value of the first time step threshold (for example, the first time step threshold can be equal to 1000). The denoising time step feature corresponding to the denoising time step is used to determine the proportion of noise removed (i.e., the amplitude of the predicted noise); the noise intensity coefficient of the denoising network is used to balance between retaining the content of the input image and introducing randomness. The noise intensity coefficient of the denoising network is a coefficient predefined in relation to the denoising time step that determines the noise intensity, and the noise intensity coefficient of the denoising network decreases as the denoising time step increases.

[0099] It should be understood that when the denoising process of the first diffusion model is the first denoising process, the computer device can obtain the denoising time step feature corresponding to the denoising time step of the noise image features (i.e., the denoising time step of the first denoising process), input the noise image features, the denoising time step feature, and the texture control features into the denoising network of the first diffusion model (i.e., the first denoising network), input the noise image features, the denoising time step feature, the texture control features, and the control image features into the control network of the first diffusion model (i.e., the first control network), and predict the noise in the noise image features through the denoising network (i.e., the first denoising network) and the control network (i.e., the first control network). Further, the computer device can determine the denoised image features (i.e., the denoised image features obtained from the first denoising process) according to the noise in the noise image features, the noise image features, and the noise intensity coefficient of the denoising network. Optionally, the computer device can perform a subtraction operation on the noise image features and the noise in the noise image features to obtain the denoised image features obtained from the first denoising process.

[0100] Among them, the noise intensity coefficients of the denoising network include a first noise intensity coefficient corresponding to the denoising time step of the noise image features and a second noise intensity coefficient corresponding to the denoising time step of the denoised image features obtained by the first denoising process. It can be understood that the computer device can perform a subtraction operation on the default noise intensity coefficient (for example, the default noise intensity coefficient can be equal to 1) and the first noise intensity coefficient to obtain a first difference intensity coefficient; the computer device can perform a subtraction operation on the default noise intensity coefficient and the second noise intensity coefficient to obtain a second difference intensity coefficient. Further, the computer device can perform a square root operation on the first noise intensity coefficient, the second noise intensity coefficient, the first difference intensity coefficient, and the second difference intensity coefficient respectively to obtain a first square root noise intensity coefficient corresponding to the first noise intensity coefficient, a second square root noise intensity coefficient corresponding to the second noise intensity coefficient, a first square root difference intensity coefficient corresponding to the first difference intensity coefficient, and a second square root difference intensity coefficient corresponding to the second difference intensity coefficient. Among them, the square root operation means taking the square root. Further, the computer device can perform a multiplication operation on the first square root difference intensity coefficient and the noise in the noise image features to obtain a first difference noise; the computer device can perform a multiplication operation on the second square root difference intensity coefficient and the noise in the noise image features to obtain a second difference noise. Further, the computer device can perform a subtraction operation on the noise image features and the first difference noise to obtain a first candidate difference noise. Further, the computer device can perform a division operation on the first candidate difference noise and the first square root noise intensity coefficient to obtain a second candidate difference noise. Further, the computer device can perform a multiplication operation on the second candidate difference noise and the second square root noise intensity coefficient to obtain a third candidate difference noise. Further, the computer device can perform an addition operation on the third candidate difference noise and the second difference noise to obtain the denoised image features obtained by the first denoising process.

[0101] Among them, optionally, the noise intensity coefficient of the denoising network includes a first noise intensity coefficient corresponding to the denoising time step of the noise image feature and noise intensity coefficients corresponding to time steps less than the denoising time step of the noise image feature. For example, the denoising time step of the noise image feature can be equal to 5, and the time steps less than the denoising time step of the noise image feature can be 4, 3, 2, and 1. It can be understood that the computer device can perform a multiplication operation on the first noise intensity coefficient corresponding to the denoising time step of the noise image feature and the noise intensity coefficients corresponding to time steps less than the denoising time step of the noise image feature to obtain a combined noise intensity coefficient corresponding to the denoising time step of the noise image feature; the computer device can perform a subtraction operation on the default noise intensity coefficient and the first noise intensity coefficient to obtain a first difference intensity coefficient; the computer device can perform a square root operation on the first noise intensity coefficient to obtain a first square root noise intensity coefficient. Further, the computer device can perform a subtraction operation on the default noise intensity coefficient and the combined noise intensity coefficient to obtain a third difference intensity coefficient, and perform a square root operation on the third difference intensity coefficient to obtain a third square root difference intensity coefficient. Further, the computer device can perform a division operation on the first difference intensity coefficient and the third square root difference intensity coefficient to obtain a first candidate intensity coefficient, and perform a multiplication operation on the first candidate intensity coefficient and the noise in the noise image feature to obtain a fourth candidate difference noise; the computer device can perform a division operation on the default noise intensity coefficient and the first square root noise intensity coefficient to obtain a second candidate intensity coefficient. Further, the computer device can perform a subtraction operation on the noise image feature and the fourth candidate difference noise to obtain a fifth candidate difference noise, and perform a multiplication operation on the second candidate intensity coefficient and the fifth candidate difference noise to obtain a sixth candidate difference noise. Further, the computer device can perform an addition operation on the sixth candidate difference noise and the random noise to obtain the denoised image feature obtained by the first denoising process. Among them, the random noise is obtained by performing a multiplication operation on the noise sampled from the standard normal distribution and the standard deviation of the noise.

[0102] It should be understood that when the denoising process of the first diffusion model is the u-th denoising process (where u can be a positive integer less than or equal to M and u is not equal to 1), the computer device can obtain the denoising time step feature corresponding to the denoising time step of the denoised image feature obtained from the (u - 1)-th denoising process (i.e., the denoising time step of the u-th denoising process), input the denoised image feature, denoising time step feature, and texture control feature obtained from the (u - 1)-th denoising process into the denoising network of the first diffusion model (i.e., the u-th denoising network), input the denoised image feature, denoising time step feature, texture control feature, and control image feature obtained from the (u - 1)-th denoising process into the control network of the first diffusion model (i.e., the u-th control network), and predict the noise in the denoised image feature obtained from the (u - 1)-th denoising process through the denoising network (i.e., the u-th denoising network) and the control network (i.e., the u-th control network). Further, the computer device can determine the denoised image feature (i.e., the denoised image feature obtained from the u-th denoising process) based on the noise in the denoised image feature obtained from the (u - 1)-th denoising process, the denoised image feature obtained from the (u - 1)-th denoising process, and the noise intensity coefficient of the denoising network. The specific process of the computer device determining the denoised image feature based on the noise in the denoised image feature obtained from the (u - 1)-th denoising process, the denoised image feature obtained from the (u - 1)-th denoising process, and the noise intensity coefficient of the denoising network can refer to the description of determining the denoised image feature based on the noise in the noise image feature, the noise image feature, and the noise intensity coefficient of the denoising network above, which will not be elaborated here.

[0103] Among them, the first diffusion module can combine the features learned by the denoising network and the features learned by the control network through a skip connection (i.e., a jump connection, which means directly transmitting the features in the front to the layers in the back by skipping some layers), so as to not only retain the generation ability of the denoising network, but also introduce the control of external conditions; the denoising network can include a residual module and an attention module, the control network can include a residual module and an attention module, the residual module can be a Residual Network (ResNet), and the attention module can include a Self-Attention module and a Cross-Attention module. The Self-Attention module can obtain the receptive field containing the entire input (i.e., denoising image features and / or control image features) and capture the relationships between different positions in the features. The Cross-Attention module can fuse the inputs of different modalities (i.e., denoising image features and / or control image features, texture control features); the residual module can be used to process the denoising time step features, and the attention module can be used to process the texture control features. Among them, the Self-Attention module of the control network can be associated with S geometric images (or share the calculation of S geometric images), and the geometric images in each perspective can fuse the geometric images in other perspectives to ensure multi-perspective consistency.

[0104] It can be understood that the execution logic of the attention module can refer to the following formula (1):

[0105]

[0106] Among them, Q can represent the query vector, K can represent the key vector, V can represent the value vector, d can represent the dimension of the key vector. In the Self-Attention module, Q, K, and V are obtained by linearly transforming the features associated with the control image features and / or denoising image features through weight matrices (query weight matrix, key weight matrix, and value weight matrix). The features associated with the control image features and / or denoising image features refer to the features generated by the denoising network and the control network based on the denoising image features, denoising time step features, texture control features, and control image features obtained from the (u - 1)-th denoising process, or the features generated by the denoising network and the control network based on the noise image features, denoising time step features, texture control features, and control image features. Optionally, in the Cross-Attention module, Q is obtained by linearly transforming the features associated with the control image features and / or denoising image features through a weight matrix (query weight matrix), and K and V are obtained by linearly transforming the texture control features through weight matrices (key weight matrix and value weight matrix).

[0107] Step S103, perform denoising processing on the predicted image to obtain the material image corresponding to the spatial object;

[0108] Specifically, the computer device can obtain the denoised image features corresponding to the predicted image (i.e., the denoised image features obtained by the first diffusion model for the M-th denoising process, where the M-th denoising process is the last denoising process of the first diffusion model). Further, the computer device can perform denoising processing on the denoised image features in the second diffusion model to obtain the material denoised image features. Further, the computer device can perform decoding processing on the material denoised image features to obtain the material image corresponding to the spatial object. Wherein, the material image includes sub-material images of the spatial object at each viewing angle; the embodiments of the present application do not limit the decoder used for the decoding process. For example, the decoder for performing decoding processing on the material denoised image features can be the decoder of a variational autoencoder.

[0109] It can be understood that the computer device can perform Q denoising processes on the denoised image features in the second diffusion model in sequence to obtain the material denoised image features, and perform decoding processing on the material denoised image features obtained by the Q-th denoising process to obtain the material image corresponding to the spatial object. Wherein, Q here can be a positive integer. For ease of understanding, the present application takes Q greater than 1 as an example for illustration. The second diffusion model can include Q denoising networks. One denoising network is used to implement one denoising process, and the Q denoising networks are the same. The embodiments of the present application do not limit the network type of the denoising network. For example, the denoising network can be a U-Net network; the denoising time steps corresponding to the Q denoising processes continuously decrease. The denoising time step refers to a positive integer less than or equal to the second time step threshold, and the second time step threshold is used to determine the number of denoising processes performed by the second diffusion model (for example, the denoising time steps corresponding to the Q denoising processes are not adjacent, and the number of denoising processes is equal to the number of denoising time steps selected from 1 to the second time step threshold, so as to reduce the sampling steps through non-adjacent denoising time steps and improve the efficiency of generating the material image. For example, the denoising time steps selected from 1 to the second time step threshold can be 1000, 900, …, 100, and the number of denoising processes is equal to 10 times). The embodiments of the present application do not limit the specific value of the second time step threshold (for example, the second time step threshold can be equal to 1000). The denoising time step feature corresponding to the denoising time step is used to determine the proportion of noise removal (i.e., the amplitude of the predicted noise); the noise intensity coefficient of the denoising network is used to balance between retaining the content of the input image and introducing randomness. The noise intensity coefficient of the denoising network is a predefined coefficient related to the denoising time step that determines the noise intensity, and the noise intensity coefficient of the denoising network decreases as the denoising time step increases.

[0110] It should be understood that when the denoising process of the second diffusion model is the first denoising process, the computer device can obtain the denoising time step feature corresponding to the denoising time step of the denoised image feature (i.e., the denoising time step of the first denoising process), input the denoised image feature and the denoising time step feature into the denoising network of the second diffusion model (i.e., the first denoising network), and predict the noise in the denoised image feature through the denoising network (i.e., the first denoising network). Further, the computer device can determine the material denoised image feature (i.e., the material denoised image feature obtained by the first denoising process) according to the noise in the denoised image feature, the denoised image feature, and the noise intensity coefficient of the denoising network. The specific process of the computer device determining the material denoised image feature according to the noise in the denoised image feature, the denoised image feature, and the noise intensity coefficient of the denoising network can refer to the description of determining the denoised image feature according to the noise in the noise image feature, the noise image feature, and the noise intensity coefficient of the denoising network above, and will not be elaborated here.

[0111] It should be understood that when the denoising process of the second diffusion model is the e-th denoising process (where e can be a positive integer less than or equal to Q and e is not equal to 1), the computer device can obtain the denoising time step feature corresponding to the denoising time step of the material denoised image feature obtained by the (e - 1)-th denoising process (i.e., the denoising time step of the e-th denoising process), input the material denoised image feature obtained by the (e - 1)-th denoising process and the denoising time step feature into the denoising network of the second diffusion model (i.e., the e-th denoising network), and predict the noise in the material denoised image feature obtained by the (e - 1)-th denoising process through the denoising network (i.e., the e-th denoising network). Further, the computer device can determine the material denoised image feature (i.e., the material denoised image feature obtained by the e-th denoising process) according to the noise in the material denoised image feature obtained by the (e - 1)-th denoising process, the material denoised image feature obtained by the (e - 1)-th denoising process, and the noise intensity coefficient of the denoising network. The specific process of the computer device determining the material denoised image feature according to the noise in the material denoised image feature obtained by the (e - 1)-th denoising process, the material denoised image feature obtained by the (e - 1)-th denoising process, and the noise intensity coefficient of the denoising network can refer to the description of determining the denoised image feature according to the noise in the noise image feature, the noise image feature, and the noise intensity coefficient of the denoising network above, and will not be elaborated here.

[0112] Among them, the number of material images is K. The K material images can be called physical material components (i.e., PBR components). Here, K can be a positive integer. The K material images respectively correspond to different physical materials (i.e., PBR materials). The material image corresponding to each physical material includes the sub-material images of the spatial object at each viewing angle. Each material image is used to perform multi-view baking on the spatial object to obtain the spatially object after multi-view baking corresponding to the physical material corresponding to each material image. The spatially objects after multi-view baking corresponding to each physical material can be combined into the final spatially object after multi-view baking. The K physical materials include at least one of a color material (i.e., Albedo material, where Albedo refers to the base color under no light conditions and represents the proportion of reflected diffuse light), a metallic material (i.e., Metallic material, where Metallic is a scalar value used to describe metallic properties and determines whether it is a metal or a non-metal), or a roughness material (i.e., Roughness material, where Roughness is a scalar value used to describe smoothness and determines the distribution of reflected light). The K material images include at least one of a color material image (i.e., the material image corresponding to the color material), a metallic material image (i.e., the material image corresponding to the metallic material), or a roughness material image (i.e., the material image corresponding to the roughness material). The color material image includes the sub-material images of the spatial object at each viewing angle (i.e., the sub-material images corresponding to the color material). The metallic material image includes the sub-material images of the spatial object at each viewing angle (i.e., the sub-material images corresponding to the metallic material). The roughness material image includes the sub-material images of the spatial object at each viewing angle (i.e., the sub-material images corresponding to the roughness material). Therefore, the computer device can perform component decomposition on the predicted image to obtain K material images. Component decomposition means decomposing the predicted image into K components (i.e., K material images).

[0113] Therefore, in the embodiment of the present application, the noise image can be denoised according to the texture control data and the control image in the first diffusion model to obtain the predicted image corresponding to the spatial object, and then the predicted image can be denoised in the second diffusion model to obtain the material image corresponding to the spatial object. Thus, the material image corresponding to the spatial object is output through two diffusion models instead of directly outputting the material image corresponding to the spatial object through one diffusion model, so that the accuracy of the material image can be improved through the two-stage first diffusion model and second diffusion model.

[0114] Step S104, perform multi-view baking on the spatial object according to the material image to obtain the spatially object after multi-view baking.

[0115] Specifically, the computer device can randomly initialize the texture map (i.e., texture image) of the spatial object to obtain the rendered texture map corresponding to the spatial object (i.e., the rendered texture map refers to the texture map obtained by randomly initializing the spatial object). Sample the rendered texture map according to S perspectives (i.e., sample the rendered texture map according to S cameras) to obtain the rendered images under each perspective (i.e., the pixels of the rendered image are sampled from the rendered texture map according to the position of the camera). Among them, the S cameras (or S perspectives) share the rendered texture map, and the S perspectives include perspective G i , perspective G i can be any one of the S perspectives, and here i can be a positive integer less than or equal to S. Further, the computer device can obtain the perspective weights corresponding to the S perspectives respectively, and determine the perspective loss values corresponding to the S perspectives respectively according to the S rendered images, the S sub-material images in the material image, and the S perspective weights. Further, the computer device can perform a summation operation on the S perspective loss values to obtain the image loss value corresponding to the rendered texture map, and perform image adjustment on the rendered texture map according to the image loss value to obtain the material texture map corresponding to the spatial object (i.e., the optimized texture map). Among them, the image adjustment can be repeated iteratively multiple times (for example, 200 times) to optimize the rendered texture map; the computer device can perform multi-view baking on the spatial object according to K material images respectively, and the material texture map matches the physical material corresponding to the material image. The number of material texture maps is K, and the K material texture maps include at least one of the material texture map corresponding to the color material (i.e., albedo texture map), the material texture map corresponding to the metallic material (i.e., metallic texture map), or the material texture map corresponding to the rough material (i.e., roughness texture map). Further, the computer device can perform texture mapping on the spatial object according to the material texture map to obtain the spatial object after multi-view baking. Among them, the spatial object after multi-view baking has the texture indicated by the texture control data.

[0116] It should be understood that the computer device can obtain L object vertices of the spatial object (i.e., the vertices of the mesh of the three-dimensional model). Among them, here L can be a positive integer greater than 1, and the L object vertices include object vertex P j , here j can be a positive integer less than or equal to L, and object vertex P j can be any one of the L object vertices. Further, the computer device can determine the L object vertices respectively for perspective G according to the vertex coordinates corresponding to the L object vertices respectively, the position of the camera optical center corresponding to perspective G (i.e., the position of the optical center of the camera corresponding to perspective G i ), and the normal vectors (i.e., normal vectors) corresponding to the L object vertices respectively i ), and the normal vectors (i.e., normal vectors) corresponding to the L object vertices respectively iVertex weights. Among them, the normal vectors corresponding to the L object vertices are obtained by performing normal calculations on the L vertices respectively, that is, performing normal calculations on each object vertex can obtain the normal vector. Further, the computer device can perform a summation operation on the L vertex weights (that is, the vertex weights corresponding to the L object vertices respectively for the viewing angle G i of the vertex weights) to obtain the viewing angle G i corresponding viewing angle weight.

[0117] Among them, the computer device can perform a subtraction operation on the vertex coordinates corresponding to the L object vertices respectively (for example, the vertex coordinates corresponding to the object vertex P j and the camera optical center position corresponding to the viewing angle G i to obtain the line-of-sight directions corresponding to the L object vertices respectively (for example, the line-of-sight direction corresponding to the object vertex P j ). Further, the computer device can perform a multiplication operation on the line-of-sight directions corresponding to the L object vertices respectively and the normal vectors corresponding to the L object vertices respectively (for example, the line-of-sight direction corresponding to the object vertex P j and the normal vector corresponding to the object vertex P j ) to obtain L included angle values (for example, the included angle value corresponding to the object vertex P j ); the computer device can perform a multiplication operation on the absolute values of the line-of-sight directions corresponding to the L object vertices respectively and the absolute values of the normal vectors corresponding to the L object vertices respectively (for example, the absolute value of the line-of-sight direction corresponding to the object vertex P j and the absolute value of the normal vector corresponding to the object vertex P j ) to obtain L vector magnitudes (for example, the vector magnitude corresponding to the object vertex P j ). Further, the computer device can perform a division operation on the L included angle values (for example, the included angle value corresponding to the object vertex P j ) and the L vector magnitudes (for example, the vector magnitude corresponding to the object vertex P j ) to obtain the cosine value of the included angle between the line-of-sight direction and the normal vector (for example, the cosine value of the included angle between the line-of-sight direction corresponding to the object vertex P j and the normal vector corresponding to the object vertex P j ). Further, the computer device can determine the absolute value of the cosine value of the included angle as the vertex weights corresponding to the L object vertices respectively for the viewing angle G i (that is, the absolute value of the cos value of the line-of-sight direction and the normal vector) (for example, the vertex weight of the object vertex P j for the viewing angle G i ). Optionally, the computer device can enhance the weight discrimination of the vertex weights corresponding to the L object vertices respectively for the viewing angle G i through the exp power (that is, e as the base and the absolute value of the cosine value of the included angle as the exponent).

[0118] It can be understood that the larger the vertex weight is, the more positive the object vertex is (i.e., the more positive the object vertex is relative to the camera) and the more credible it is. The computer device obtains the vertex weights of L object vertices respectively for the viewing angle G i The specific process of the vertex weights can be seen in the following formula (2):

[0119]

[0120] where i can represent the object vertex (i.e., any one of the L object vertices, for example, the object vertex P j ), p i can represent the vertex coordinates corresponding to the object vertex, c can represent the position of the camera optical center, n i can represent the normal vector, p i -c can represent the line-of-sight direction, abs can represent the absolute value, exp represents the exponential function, and w i represents the vertex weight of the object vertex.

[0121] It should be understood that the computer device can perform a difference operation on the rendered image under the viewing angle G i and the sub-material image under the viewing angle G i to obtain the difference image corresponding to the viewing angle G i . Among them, the sub-material image represents the ground truth map, and the rendered image represents the rendered map. Further, the computer device can perform an absolute value operation on the difference image corresponding to the viewing angle G i to obtain the absolute value image corresponding to the viewing angle G i . Further, the computer device can perform a multiplication operation on the viewing angle weight corresponding to the viewing angle G i and the absolute value image corresponding to the viewing angle G i to obtain the viewing angle loss value corresponding to the viewing angle G i .

[0122] It can be understood that the specific process of the computer device obtaining the image loss value can be seen in the following formula (3):

[0123]

[0124] where i can represent the viewing angle G i , n can represent the number S of viewing angles, can represent the sub-material image under the viewing angle G i , (or ) can represent the rendered image under the viewing angle G i , X can represent the rendered texture map, and T i represents the position of the camera corresponding to the viewing angle G i , Wi can represent viewing angle G i the corresponding viewing angle weight can represent viewing angle G i the corresponding absolute value image can represent viewing angle G i the corresponding viewing angle loss value can represent the image loss value

[0125] For ease of understanding, please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a texture model provided by an embodiment of the present application. As Figure 5 shown, Stage 1 can correspond to the above-mentioned step S102, Stage 2 can correspond to the above-mentioned step S103, and Stage 3 can correspond to the above-mentioned step S104. Stage 1 is used to implement multi-view image generation based on geometric prior information, and Stage 2 is used to implement PBR material generation based on multi-view images

[0126] As Figure 5 shown, in Stage 1, the computer device can perform multi-view rendering on the spatial object to obtain the control image 50a corresponding to the spatial object. In the first diffusion model 51a, the noise image is denoised according to the texture control data and the control image 50a to obtain the predicted image 50b corresponding to the spatial object. Further, in Stage 2, the computer device can perform denoising on the predicted image 50b in the second diffusion model 51b to obtain the material images corresponding to the spatial object (i.e., material images 50c, 50d, and 50e). Further, in Stage 3, the computer device can perform multi-view baking on the spatial object according to the material images (i.e., material images 50c, 50d, and 50e) to obtain the spatially object 50g after multi-view baking

[0127] For ease of understanding, please refer to Figure 6 , Figure 6 which is a schematic diagram of a texture baking scenario provided by an embodiment of the present application. As Figure 6 shown, the user can provide a spatial object (e.g., spatial object 60a) through a selection operation and provide texture control data (e.g., the texture control data can be text data, and the text data can be "a colorful axe" or "a bright hammer") through an input operation. The computer device can generate a spatially object after multi-view baking (e.g., spatial object 60b) according to the spatial object and the texture control data. Among them, the spatial object 60a is a three-dimensional model without PBR material texture, and the spatial object 60b is a three-dimensional model with PBR material texture

[0128] It can be seen that the embodiments of the present application can generate geometric images for each of the S perspectives for the geometric information of the spatial object, splice the S geometric images into a multi-perspective image (i.e., the control image) for the geometric information, perform denoising processing on the noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object, and further perform denoising processing on the predicted image to obtain a material image corresponding to the spatial object, and perform multi-perspective baking on the spatial object according to the material image to obtain a multi-perspective baked spatial object, where the multi-perspective baked spatial object has the texture indicated by the texture control data. Therefore, the present application can introduce geometric images for describing the geometric information of the spatial object. Since the geometric information can achieve fine control of the spatial object, the material image generated according to the texture control data and the geometric image can perfectly match the spatial object, thereby improving the accuracy of the multi-perspective baked spatial object. In addition, the present application can use the geometric images for each of the S perspectives simultaneously instead of using the geometric image in a single perspective. The geometric images for each perspective can learn from each other, thereby avoiding the perspective inconsistency (e.g., the inconsistency in the overlapping area) when generating sub-material images for each of the S perspectives, and ensuring the multi-perspective consistency of the sub-material images for each perspective in the material image.

[0129] Further, please refer to Figure 7 , Figure 7 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. The data processing method can be executed by a computer device, and the computer device can be the server 20a in the corresponding embodiment of the above Figure 2 or the terminal device 20b in the corresponding embodiment of the above Figure 2 . Among them, the data processing method can include steps S201 - step S205:

[0130] Step S201, obtain a second initial diffusion model, sample texture control data, and sample geometric images of the sample object for each of the S perspectives, and perform image splicing on the S sample geometric images to obtain a sample control image;

[0131] Among them, the sample texture control data is equivalent to the texture control data, the sample object is equivalent to the spatial object, the sample geometric image is equivalent to the geometric image, the sample control image is equivalent to the control image, the sample texture control data, the sample object, the sample geometric image, and the sample control image are the relevant data of the diffusion model in the training stage, and the texture control data, the spatial object, the geometric image, and the control image are the relevant data of the diffusion model in the inference stage.

[0132] Among them, for the specific process of the computer device obtaining the sample geometric images of the sample object from each of the S perspectives, reference can be made to the description of obtaining the geometric images of the spatial object from each of the S perspectives above, and details will not be elaborated here; for the specific process of the computer device performing image stitching on the S sample geometric images to obtain the sample control image, reference can be made to the description of performing image stitching on the S geometric images to obtain the control image above, and details will not be elaborated here.

[0133] Step S202: Denoise the noise image according to the sample texture control data and the sample control image to obtain the sample prediction image corresponding to the sample object;

[0134] Among them, the sample prediction image includes the sub-sample prediction images of the sample object from each perspective. The sample prediction image is equivalent to the prediction image, the sub-sample prediction image is equivalent to the sub-prediction image. The sample prediction image and the sub-sample prediction image are relevant data in the training stage of the diffusion model, and the prediction image and the sub-prediction image are relevant data in the inference stage of the diffusion model.

[0135] Among them, for the specific process of the computer device denoising the noise image according to the sample texture control data and the sample control image to obtain the sample prediction image corresponding to the sample object, reference can be made to the description of denoising the noise image according to the texture control data and the control image to obtain the prediction image corresponding to the spatial object above, and details will not be elaborated here.

[0136] Step S203: Input the sample prediction image features corresponding to the sample prediction image into the second initial diffusion model, and denoise the sample prediction image through the second initial diffusion model to obtain the sample material image corresponding to the sample object;

[0137] Among them, the sample material image includes the sub-sample material images of the sample object from each perspective. The number of sample material images is K, and the K sample material images can be referred to as physical material components (i.e., PBR components); the sample prediction image features corresponding to the sample prediction image are equivalent to the denoised image features corresponding to the prediction image, the sample material image is equivalent to the material image, and the sub-sample material image is equivalent to the sub-material image. The sample prediction image features corresponding to the sample prediction image, the sample material image, and the sub-sample material image are relevant data in the training stage of the diffusion model, and the denoised image features corresponding to the prediction image, the material image, and the sub-material image are relevant data in the inference stage of the diffusion model.

[0138] Among them, for the specific process of the computer device denoising the sample prediction image through the second initial diffusion model to obtain the sample material image corresponding to the sample object, reference can be made to the description of denoising the prediction image through the second diffusion model to obtain the material image corresponding to the spatial object above, and details will not be elaborated here.

[0139] Step S204: Perform illumination estimation on the sample prediction image to obtain the image illumination information of the sample prediction image, and perform image rendering on the image illumination information and the sample material image to obtain the sample rendering image corresponding to the sample object;

[0140] Among them, the sample rendering image includes sub-sample rendering images of the sample object at each viewing angle. Illumination estimation means inferring the illumination conditions from the sample prediction image (i.e., realizing image de-illumination), and image rendering means fusing the image illumination information and the physical material components to achieve illumination rendering on the physical material components.

[0141] Step S205: Adjust the parameters of the second initial diffusion model according to the sample rendering image and the sample prediction image to obtain the second diffusion model.

[0142] Specifically, the computer device can determine the model loss value of the second initial diffusion model according to the sample rendering image and the sample prediction image. Further, the computer device can adjust the parameters of the second initial diffusion model according to the model loss value of the second initial diffusion model to obtain the second diffusion model. Among them, the parameter adjustment can be repeated iteratively to optimize the second diffusion model, and the computer device can adjust the parameters of the initial denoising network in the second initial diffusion model according to the model loss value of the second initial diffusion model to obtain the denoising network in the second initial diffusion model.

[0143] It can be understood that the sample rendering image represents the rendered image (i.e., rgb_render), and the sample prediction image represents the input image (i.e., rgb_input). The computer device can perform loss constraint on the rendered image and the input image (i.e., loss constraint), that is, the computer device can calculate the model loss value (i.e., loss) of the second initial diffusion model through the loss function, the sample rendering image, and the sample prediction image. This application does not limit the loss constraint here, that is, this application does not limit the function type of the loss function. For example, the loss function can be the L1 loss function (loss = ||rgb input -rgb render ||), the L2 loss function (loss = ||rgb input -rgb render || 2 )), the Perception loss function.

[0144] Among them, when the loss function is a perceptual loss function, the computer device can respectively extract features from the sample rendered image and the sample predicted image through the neural network model, and obtain N rendered image features with different sizes corresponding to the sample rendered image and N predicted image features with different sizes corresponding to the sample predicted image. Here, N can be a positive integer, and the application does not limit the model type of the neural network model. For example, the neural network model can be a VGG (Visual Geometry Group) model. Further, the computer device can obtain the rendered image features and predicted image features with the same size from the N rendered image features with different sizes and the N predicted image features with different sizes, and obtain the feature loss between the rendered image features and predicted image features with the same size (for example, the computer device can obtain the feature loss between the rendered image features and predicted image features with the same size through the L1 loss function), and obtain the feature losses corresponding to N different sizes respectively. Further, the computer device can perform an addition operation on the N feature losses to obtain the model loss value corresponding to the perceptual loss function.

[0145] Optionally, the computer device can determine the sub-loss values corresponding to S perspectives according to the S sub-sample rendered images and the S sub-sample predicted images (that is, determine the sub-loss values according to the sub-sample rendered images and sub-sample predicted images with the same perspective), and perform an addition operation on the S sub-loss values to obtain the model loss value of the second initial diffusion model. The specific process of the computer device determining the sub-loss values corresponding to S perspectives according to the S sub-sample rendered images and the S sub-sample predicted images can refer to the description of determining the model loss value of the second initial diffusion model according to the sample rendered image and the sample predicted image, and will not be elaborated here.

[0146] For ease of understanding, please refer to Figure 5 , as Figure 5 shown, Stage 1 can correspond to step S202 above, and Stage 2 can correspond to steps S203 - S205 above. Stage 1 is used to implement multi-view image generation based on geometric prior information, and Stage 2 is used to implement PBR material generation based on multi-view images.

[0147] As Figure 5As shown in the figure, in the first stage, the computer device can perform multi-view rendering on the sample object to obtain a sample control image 50a corresponding to the sample object. In the first diffusion model 51a, the noise image is denoised according to the sample texture control data and the sample control image 50a to obtain a sample prediction image 50b corresponding to the sample object. Further, in the second stage, the computer device can perform denoising on the sample prediction image 50b in the second initial diffusion model 51b to obtain sample material images corresponding to the sample object (i.e., sample material image 50c, sample material image 50d, and sample material image 50e). Further, the computer device can perform light estimation on the sample prediction image 50b to obtain the image light information of the sample prediction image 50b, and perform image rendering on the image light information and the sample material images (i.e., sample material image 50c, sample material image 50d, and sample material image 50e) to obtain a sample rendering image 50f corresponding to the sample object. Further, the computer device can determine the model loss value (i.e., the rendering loss) of the second initial diffusion model 51b according to the sample rendering image 50f and the sample prediction image 50b, and adjust the parameters of the second initial diffusion model 51b according to the model loss value of the second initial diffusion model 51b to obtain the second diffusion model 51b.

[0148] Among them, the second initial diffusion model and the second diffusion model can be collectively referred to as the second generalization network model. The second initial diffusion model and the second diffusion model are the names of the second generalization network model at different times. In the training stage, the second generalization network model can be called the second initial diffusion model. In the prediction stage (i.e., the inference stage), the second generalization network model can be called the second diffusion model.

[0149] It can be seen that in the embodiment of the present application, the noise image can be denoised according to the sample texture control data and the sample control image in the first diffusion model to obtain a sample prediction image corresponding to the sample object. Further, in the second initial diffusion model, the sample prediction image can be denoised to obtain a sample material image corresponding to the sample object. Further, the embodiment of the present application can estimate a lighting model for the sample prediction image (i.e., the image light information obtained by the lighting estimation), perform image rendering on the image light information and the sample material images (i.e., fuse the lighting model and the PBR material image) to re-render the sample prediction image to obtain a restored sample rendering image (i.e., the rendered sample rendering image). Further, the second initial diffusion model is trained through the input sample prediction image and the restored sample rendering image to obtain the second diffusion model, thereby improving the accuracy of the material image output by the second diffusion model during inference.

[0150] Further, please refer to Figure 8 , Figure 8It is a schematic flowchart of a data processing method provided by an embodiment of the present application. This data processing method can be executed by a computer device, which can be the server 20a in the corresponding embodiment above, or the terminal device 20b in the corresponding embodiment above. Among them, this data processing method can include step S301-step S303: Figure 2 and can also be Figure 2 in the corresponding embodiment of the terminal device 20b above. Among them, this data processing method can include steps S301-S303:

[0151] Step S301: Adjust the parameters of the initial denoising network in the first initial diffusion model to obtain a denoising network;

[0152] Specifically, the computer device can obtain the initial denoising network (i.e., the initial denoising network in the first initial diffusion model), sample texture control data, a sample image, and a set of time steps. Among them, the set of time steps includes at least two time steps. The embodiment of the present application does not limit the number of time steps in the set of time steps. The number of time steps in the set of time steps is equal to the time step threshold (i.e., the first time step threshold). For example, the time step threshold can be equal to 1000, and the set of time steps can include 1000 time steps. The 1000 time steps can include time step 1, …, time step 1000; the sample image has the texture indicated by the sample texture control data, and the sample texture control data is equivalent to the texture control data. Further, the computer device can perform random uniform sampling on the set of time steps to obtain the noisy time steps among at least two time steps. Among them, each time step in the set of time steps has the same probability of being sampled. For example, the noisy time step can be time step 200. Further, the computer device can determine the noisy image feature corresponding to the noisy time step according to the sample image feature corresponding to the sample image, the noisy Gaussian noise corresponding to the sample image feature, and the noisy time step. Among them, the noisy Gaussian noise follows a normal distribution. The noisy time step is used to indicate the amplitude of the added noise. The sample image feature is obtained by encoding the sample image. The embodiment of the present application does not limit the encoder used for the encoding process. For example, the encoder for encoding the sample image can be the encoder of a variational autoencoder. Further, the computer device can input the noisy image feature, the sample texture control feature corresponding to the sample texture control data, and the noisy time step feature corresponding to the noisy time step into the initial denoising network, predict the noise in the noisy image feature through the initial denoising network, and determine the noise in the noisy image feature as the first noise. Among them, the noisy time step feature corresponding to the noisy time step is used to determine the proportion of the added noise. Further, the computer device can adjust the parameters of the initial denoising network according to the first noise and the noisy Gaussian noise to obtain a denoising network (i.e., the denoising network in the first diffusion model). Among them, the first noise represents the noise added by the initial denoising network predicted for the noisy time step on the sample image feature, and the noisy Gaussian noise represents the actually added noise on the sample image feature.

[0153] Among them, the parameter adjustment can be repeated iteratively for multiple times to optimize the denoising network. The computer device can determine the loss value of the initial denoising network according to the first noise and the added Gaussian noise, and perform parameter adjustment (or iterative adjustment) on the initial denoising network according to the loss value of the initial denoising network to obtain the denoising network. Among them, the computer device can determine the loss value of the initial denoising network by means of mean square error, root mean square error, mean absolute error, etc.

[0154] It can be understood that the computer device can take the square root of the combined noise intensity coefficient corresponding to the noise addition time step to obtain the first square root coefficient; the computer device can perform a subtraction operation on the default noise intensity coefficient and the combined noise intensity coefficient corresponding to the noise addition time step to obtain the noise intensity difference, and take the square root of the noise intensity difference to obtain the second square root coefficient. Further, the computer device can perform a multiplication operation on the first square root coefficient and the sample image features corresponding to the sample image to obtain candidate image features; the computer device can perform a multiplication operation on the second square root coefficient and the added Gaussian noise corresponding to the sample image features to obtain candidate Gaussian noise. Further, the computer device can perform an addition operation on the candidate image features and the candidate Gaussian noise to obtain the added image features corresponding to the noise addition time step.

[0155] Step S302: Adjust the parameters of the initial control network in the first initial diffusion model to obtain a control network;

[0156] Specifically, the computer device can obtain the sample geometric images of the sample object from each of the S perspectives, perform image stitching on the S sample geometric images to obtain a sample control image. Among them, the sample object is equivalent to the spatial object, the sample geometric image is equivalent to the geometric image, and the sample control image is equivalent to the control image. Further, the computer device can determine an initial control network according to the denoising network (i.e., the initial control network in the first initial diffusion model), input the noisy image feature, the sample texture control feature, and the noisy time step feature into the denoising network, input the noisy image feature, the sample texture control feature, the noisy time step feature, and the sample control image feature corresponding to the sample control image into the initial control network, predict the noise in the noisy image feature through the denoising network and the initial control network, and determine the noise in the noisy image feature as the second noise. Among them, the network structure of the initial control network is determined by the network structure of the denoising network, and the initial control network includes a zero convolution layer and the encoder in the denoising network. Further, the computer device can adjust the parameters of the initial control network according to the second noise and the noisy Gaussian noise to obtain a control network (i.e., the control network in the first diffusion model). Among them, the second noise represents the noise added to the sample image feature at the noisy time step predicted by the denoising network and the initial control network, and the noisy Gaussian noise represents the noise actually added to the sample image feature.

[0157] Among them, the parameter adjustment can be repeated iteratively multiple times to optimize the control network. The computer device can determine the loss value of the initial control network according to the second noise and the noisy Gaussian noise, and adjust the parameters of the initial control network (or iterative adjustment) according to the loss value of the initial control network to obtain a control network. Among them, the computer device can determine the loss value of the initial control network by means of mean square error, root mean square error, mean absolute error, etc.

[0158] It can be understood that the specific process of the computer device determining the loss value of the initial control network according to the second noise and the noisy Gaussian noise can refer to the following formula (4):

[0159]

[0160] Among them, ∈ represents the noisy Gaussian noise, that is represents the Gaussian noise, represents the normal distribution, 0 represents the mean of the normal distribution, 1 represents the variance of the normal distribution, ∈ θ (z t ,t,c t ,c f ) represents the noise in the noisy image feature predicted by the denoising network and the initial control network, t represents the noisy time step, z t represents the noisy image feature, c t represents the sample texture control feature, cf represents the sample control image feature, L represents the loss value of the initial control network, ∈ θ represents the denoising network and the initial control network in the second initial diffusion model, and θ represents the trainable parameters in the denoising network and the initial control network.

[0161] Step S303, determining the denoising network and the control network as the first diffusion model.

[0162] Among them, the first initial diffusion model and the first diffusion model can be collectively referred to as the first generalization network model. The first initial diffusion model and the first diffusion model are the names of the first generalization network model at different times. In the training stage, the first generalization network model can be called the first initial diffusion model. In the prediction stage (i.e., the inference stage), the first generalization network model can be called the first diffusion model.

[0163] It can be seen that the embodiments of the present application can separately train the initial denoising network in the first initial diffusion model and the initial control network in the first initial diffusion model. First, train the initial denoising network, and then train the initial control network. The training of the initial control network does not affect the denoising network. Determine the trained denoising network and control network as the first diffusion model. Since the denoising network and the control network can fuse texture control data and control images, the present application can improve the accuracy of the predicted image output by the first diffusion model during inference, and further improve the accuracy of the material image output by the second diffusion model during inference, and improve the accuracy of the spatial object after multi-view baking.

[0164] Further, please refer to Figure 9 , Figure 9 is a schematic structural diagram of a data processing device provided by an embodiment of the present application. The data processing device 1 may include: an image acquisition module 11, a first denoising module 12, a second denoising module 13, and a multi-view baking module 14; further, the data processing device 1 may also include: a first training module 15, a second training module 16, and a third training module 17;

[0165] The image acquisition module 11 is configured to acquire geometric images of the spatial object at each of the S perspectives, and perform image stitching on the S geometric images to obtain a control image; S is an integer greater than 1;

[0166] The first denoising module 12 is configured to acquire texture control data, and perform denoising processing on the noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object; the predicted image includes sub-predicted images of the spatial object at each perspective;

[0167] Among them, the first denoising module 12 is specifically used to obtain the texture control features corresponding to the texture control data, the control image features corresponding to the control image, and the noise image features corresponding to the noise image;

[0168] The first denoising module 12 is specifically used to denoise the noise image features according to the texture control features and the control image features in the first diffusion model to obtain denoised image features;

[0169] The first denoising module 12 is specifically used to decode the denoised image features to obtain a predicted image corresponding to the spatial object.

[0170] Among them, the first denoising module 12 is specifically used to obtain the denoising time step features corresponding to the denoising time step of the noise image features;

[0171] The first denoising module 12 is specifically used to input the noise image features, the denoising time step features, and the texture control features into the denoising network of the first diffusion model, and input the noise image features, the denoising time step features, the texture control features, and the control image features into the control network of the first diffusion model, and predict the noise in the noise image features through the denoising network and the control network;

[0172] The first denoising module 12 is specifically used to determine the denoised image features according to the noise in the noise image features, the noise image features, and the noise intensity coefficient of the denoising network.

[0173] The second denoising module 13 is used to denoise the predicted image to obtain a material image corresponding to the spatial object; the material image includes sub-material images of the spatial object at each viewing angle;

[0174] Among them, the second denoising module 13 is specifically used to obtain the denoised image features corresponding to the predicted image;

[0175] The second denoising module 13 is specifically used to denoise the denoised image features in the second diffusion model to obtain material denoised image features;

[0176] The second denoising module 13 is specifically used to decode the material denoised image features to obtain a material image corresponding to the spatial object.

[0177] Among them, the number of material images is K, where K is a positive integer. The K material images respectively correspond to different physical materials. Each material image corresponding to a physical material includes sub-material images of the spatial object at each viewing angle. Each material image is used to perform multi-view baking on the spatial object to obtain a multi-view baked spatial object for each physical material corresponding to the material image; the K material images include at least one of a color material image, a metal material image, or a rough material image.

[0178] The multi-view baking module 14 is used to perform multi-view baking on a spatial object according to a material image to obtain a multi-view baked spatial object; the multi-view baked spatial object has the texture indicated by the texture control data.

[0179] Among them, the multi-view baking module 14 is specifically configured to randomly initialize the texture map of the spatial object to obtain a rendered texture map corresponding to the spatial object;

[0180] The multi-view baking module 14 is specifically configured to sample the rendered texture map according to S views to obtain a rendered image under each view;

[0181] The multi-view baking module 14 is specifically configured to obtain view weights respectively corresponding to the S views, and determine view loss values respectively corresponding to the S views according to the S rendered images, S sub-material images in the material image, and the S view weights;

[0182] The multi-view baking module 14 is specifically configured to perform a summation operation on the S view loss values to obtain an image loss value corresponding to the rendered texture map;

[0183] The multi-view baking module 14 is specifically configured to perform image adjustment on the rendered texture map according to the image loss value to obtain a material texture map corresponding to the spatial object; the material texture map matches the physical material corresponding to the material image;

[0184] The multi-view baking module 14 is specifically configured to perform texture mapping on the spatial object according to the material texture map to obtain a multi-view baked spatial object.

[0185] Among them, the S views include view G i , where i is a positive integer less than or equal to S;

[0186] The multi-view baking module 14 is specifically configured to obtain L object vertices of the spatial object; L is a positive integer greater than 1;

[0187] The multi-view baking module 14 is specifically configured to determine, according to the vertex coordinates respectively corresponding to the L object vertices, the camera optical center position corresponding to view G i and the normal vectors respectively corresponding to the L object vertices, the vertex weights of the L object vertices respectively for view G i ;

[0188] The multi-view baking module 14 is specifically configured to perform a summation operation on the L vertex weights to obtain the view weight corresponding to view G i ;

[0189] Among them, the S views include view G i , where i is a positive integer less than or equal to S;

[0190] The multi-view baking module 14 is specifically used for the perspective G i of the rendered image and the perspective G i of the sub-material image to perform a difference operation to obtain the difference image corresponding to the perspective G i ;

[0191] The multi-view baking module 14 is specifically used for the perspective G i corresponding difference image to perform an absolute value operation to obtain the absolute value image corresponding to the perspective G i ;

[0192] The multi-view baking module 14 is specifically used for the perspective G i corresponding perspective weight and the perspective G i corresponding absolute value image to perform a multiplication operation to obtain the perspective loss value corresponding to the perspective G i ;

[0193] Optionally, the first training module 15 is used to obtain a second initial diffusion model, sample texture control data, and sample geometric images of the sample object at each of the S perspectives, perform image stitching on the S sample geometric images to obtain a sample control image;

[0194] The first training module 15 is used to denoise the noise image according to the sample texture control data and the sample control image to obtain a sample prediction image corresponding to the sample object; the sample prediction image includes sub-sample prediction images of the sample object at each perspective;

[0195] The first training module 15 is used to input the sample prediction image features corresponding to the sample prediction image into the second initial diffusion model, and denoise the sample prediction image through the second initial diffusion model to obtain a sample material image corresponding to the sample object; the sample material image includes sub-sample material images of the sample object at each perspective;

[0196] The first training module 15 is used to perform light estimation on the sample prediction image to obtain the image light information of the sample prediction image, and perform image rendering on the image light information and the sample material image to obtain a sample rendering image corresponding to the sample object; the sample rendering image includes sub-sample rendering images of the sample object at each perspective;

[0197] The first training module 15 is used to adjust the parameters of the second initial diffusion model according to the sample rendering image and the sample prediction image to obtain a second diffusion model.

[0198] Optionally, the second training module 16 is used to obtain an initial denoising network, sample texture control data, a sample image, and a set of time steps; the set of time steps includes at least two time steps; the sample image has the texture indicated by the sample texture control data;

[0199] The second training module 16 is configured to perform random uniform sampling on the set of time steps to obtain noisy time steps in at least two time steps;

[0200] The second training module 16 is configured to determine, according to the sample image features corresponding to the sample image, the noisy Gaussian noise corresponding to the sample image features, and the noisy time steps, the noisy image features corresponding to the noisy time steps; the noisy Gaussian noise follows a normal distribution;

[0201] The second training module 16 is configured to input the noisy image features, the sample texture control features corresponding to the sample texture control data, and the noisy time step features corresponding to the noisy time steps into the initial denoising network, predict the noise in the noisy image features through the initial denoising network, and determine the noise in the noisy image features as the first noise;

[0202] The second training module 16 is configured to adjust the parameters of the initial denoising network according to the first noise and the noisy Gaussian noise to obtain a denoising network.

[0203] Optionally, the third training module 17 is configured to obtain the sample geometric images of the sample object at each of the S perspectives, and perform image stitching on the S sample geometric images to obtain a sample control image;

[0204] The third training module 17 is configured to determine an initial control network according to the denoising network, input the noisy image features, the sample texture control features, and the noisy time step features into the denoising network, input the noisy image features, the sample texture control features, the noisy time step features, and the sample control image features corresponding to the sample control image into the initial control network, predict the noise in the noisy image features through the denoising network and the initial control network, and determine the noise in the noisy image features as the second noise;

[0205] The third training module 17 is configured to adjust the parameters of the initial control network according to the second noise and the noisy Gaussian noise to obtain a control network.

[0206] Among them, for the specific implementation manners of the image acquisition module 11, the first denoising module 12, the second denoising module 13, and the multi-view baking module 14, reference may be made to the descriptions of steps S101 - S104 in the corresponding embodiments above, and details will not be elaborated here; for the specific implementation manner of the first training module 15, reference may be made to the descriptions of steps S201 - S205 in the corresponding embodiments above, and details will not be elaborated here; for the specific implementation manners of the second training module 16 and the third training module 17, reference may be made to the above Figure 3 corresponding embodiments' descriptions of steps S101 - S104, and will not be elaborated here; for the specific implementation manner of the first training module 15, reference may be made to the above Figure 7 corresponding embodiments' descriptions of steps S201 - S205, and will not be elaborated here; for the specific implementation manners of the second training module 16 and the third training module 17, reference may be made to the above Figure 8The descriptions of steps S301 - S303 in the corresponding embodiments will not be elaborated here. Additionally, the descriptions of the beneficial effects of using the same method will not be elaborated either.

[0207] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of that module or unit.

[0208] Furthermore, please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. This computer device can be a terminal device or a server. As Figure 10 shown, the computer device 1000 can include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 can further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, in some embodiments, the user interface 1003 can include a display screen (Display), a keyboard (Keyboard), and optionally the user interface 1003 can further include a standard wired interface and a wireless interface. Optionally, the network interface 1004 can include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 can be a high-speed RAM memory, or a non-volatile memory, for example, at least one disk memory. Optionally, the memory 1005 can also be at least one storage device located far from the aforementioned processor 1001. As Figure 10 shown, the memory 1005, as a computer-readable storage medium, can include an operating system, a network communication module, a user interface module, and a device control application program.

[0209] In the computer device 1000 as Figure 10 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the computer program stored in the memory 1005 to achieve:

[0210] Obtain the geometric images of the spatial object from each of the S perspectives, perform image stitching on the S geometric images to obtain a control image; S is an integer greater than 1;

[0211] Acquire texture control data, perform denoising on the noisy image according to the texture control data and the control image, and obtain a predicted image corresponding to the spatial object; the predicted image includes a sub-predicted image of the spatial object at each viewing angle;

[0212] De-noising the predicted image to obtain a material image corresponding to the spatial object; the material image includes a sub-material image of the spatial object at each viewing angle;

[0213] The spatial object is baked from multiple perspectives according to the material image to obtain the multi-perspective baked spatial object; the multi-perspective baked spatial object has the texture indicated by the texture control data.

[0214] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 3 , Figure 7 and Figure 8 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 9 The description of the data processing device 1 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.

[0215] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the data processing device 1 mentioned above. When the processor executes the computer program, it can execute the above-mentioned Figure 3 , Figure 7 and Figure 8 The description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0216] In addition, it should be noted that: the embodiment of the present application also provides a computer program product, which may include a computer program, and the computer program may be stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor may execute the computer program, so that the computer device performs the above Figure 3 , Figure 7 and Figure 8 The description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer program product embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0217] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0218] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A data processing method, characterized in that: include: Acquire a geometric image of the spatial object at each of the S viewing angles, perform image stitching on the S geometric images, and obtain a control image; Said S is an integer greater than 1; Acquire texture control data, perform denoising on the noisy image according to the texture control data and the control image, and obtain a predicted image corresponding to the spatial object; the predicted image includes a sub-predicted image of the spatial object at each viewing angle; De-noising the predicted image to obtain a material image corresponding to the spatial object; the material image includes a sub-material image of the spatial object at each viewing angle; The spatial object is baked from multiple perspectives according to the material image to obtain a multi-perspective baked spatial object; the multi-perspective baked spatial object has the texture indicated by the texture control data.

2. The method according to claim 1, characterized in that The denoising process is performed on the noise image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object, including: Acquire a texture control feature corresponding to the texture control data, a control image feature corresponding to the control image, and a noise image feature corresponding to the noise image; In a first diffusion model, the noise image feature is denoised according to the texture control feature and the control image feature to obtain a denoised image feature; The denoised image features are decoded to obtain a predicted image corresponding to the spatial object.

3. The method according to claim 2, characterized in that The step of performing denoising on the noise image feature according to the texture control feature and the control image feature to obtain the denoised image feature comprises: Obtaining a denoising time step feature corresponding to a denoising time step of the noise image feature; Input the noise image feature, the denoising time step feature and the texture control feature into a denoising network of the first diffusion model, input the noise image feature, the denoising time step feature, the texture control feature and the control image feature into a control network of the first diffusion model, and predict the noise in the noise image feature through the denoising network and the control network; A denoised image feature is determined based on the noise in the noisy image feature, the noisy image feature and a noise intensity coefficient of the denoising network.

4. The method according to claim 2, characterized in that: The denoising process is performed on the predicted image to obtain a material image corresponding to the spatial object, including: Acquire the denoised image features corresponding to the predicted image; In the second diffusion model, the denoised image features are denoised to obtain material denoised image features; The material denoising image features are decoded to obtain a material image corresponding to the spatial object.

5. The method according to claim 1, characterized in that The number of the material images is K, where K is a positive integer, and the K material images correspond to different physical materials respectively. The material image corresponding to each physical material includes a sub-material image of the spatial object at each viewing angle, and each material image is used to perform multi-view baking on the spatial object to obtain a multi-view baked spatial object for the physical material corresponding to each material image; the K material images include at least one of a color material image, a metal material image, or a rough material image.

6. The method according to claim 1, characterized in that The step of performing multi-perspective baking on the spatial object according to the material image to obtain the multi-perspective baked spatial object includes: Randomly initializing the texture map of the spatial object to obtain a rendering texture map corresponding to the spatial object; Sampling the rendering texture map according to the S viewing angles to obtain a rendering image at each viewing angle; Obtaining perspective weights corresponding to the S perspectives respectively, and determining perspective loss values ​​corresponding to the S perspectives respectively according to the S rendered images, the S sub-material images in the material image, and the S perspective weights; Sum the S viewing angle loss values ​​to obtain an image loss value corresponding to the rendered texture image; Performing image adjustment on the rendering texture map according to the image loss value to obtain a material texture map corresponding to the spatial object; the material texture map matches the physical material corresponding to the material image; Texture mapping is performed on the spatial object according to the material texture map to obtain a multi-view baked spatial object.

7. The method according to claim 6, characterized in that The S viewing angles include viewing angle G i , i is a positive integer less than or equal to S; The obtaining of the perspective weights respectively corresponding to the S perspectives includes: Obtain L object vertices of the spatial object; L is a positive integer greater than 1; According to the vertex coordinates corresponding to the L object vertices, the viewing angle G i The corresponding camera optical center position and the normal vectors corresponding to the L object vertices are determined to determine the L object vertices for the viewing angle G i The vertex weights of The L vertex weights are summed to obtain the viewing angle G i The corresponding viewing angle weight.

8. The method according to claim 6, characterized in that The S viewing angles include viewing angle G i , i is a positive integer less than or equal to S; The determining, according to the S rendered images, the S sub-material images in the material image, and the S viewing angle weights, the viewing angle loss values ​​respectively corresponding to the S viewing angles comprises: For the viewing angle G i The rendered image and the viewing angle G i The sub-material image under the difference operation is performed to obtain the viewing angle G i The corresponding difference image; For the viewing angle G i The corresponding difference image is subjected to absolute value calculation to obtain the viewing angle G i The corresponding absolute value image; For the viewing angle G i The corresponding viewing angle weight and the viewing angle G i The corresponding absolute value image is multiplied to obtain the viewing angle G i The corresponding viewing angle loss value.

9. The method according to claim 4, characterized in that The method further comprises: Acquire a second initial diffusion model, sample texture control data, and a sample geometric image of the sample object at each of the S viewing angles, and perform image stitching on the S sample geometric images to obtain a sample control image; De-noising the noise image according to the sample texture control data and the sample control image to obtain a sample prediction image corresponding to the sample object; the sample prediction image includes a sub-sample prediction image of the sample object at each viewing angle; Inputting the sample prediction image features corresponding to the sample prediction image into the second initial diffusion model, and performing denoising processing on the sample prediction image by using the second initial diffusion model to obtain a sample material image corresponding to the sample object; the sample material image includes a subsample material image of the sample object at each viewing angle; Performing illumination estimation on the sample prediction image to obtain image illumination information of the sample prediction image, performing image rendering on the image illumination information and the sample material image to obtain a sample rendered image corresponding to the sample object; the sample rendered image includes a subsample rendered image of the sample object at each viewing angle; According to the sample rendered image and the sample predicted image, parameters of the second initial diffusion model are adjusted to obtain the second diffusion model.

10. The method according to claim 3, characterized in that The method further comprises: Acquire an initial denoising network, sample texture control data, a sample image and a time step set; the time step set includes at least two time steps; the sample image has a texture indicated by the sample texture control data; Performing random uniform sampling on the time step set to obtain a noisy time step in the at least two time steps; Determine the noisy image feature corresponding to the noisy time step according to the sample image feature corresponding to the sample image, the noisy Gaussian noise corresponding to the sample image feature and the noisy time step; the noisy Gaussian noise obeys a normal distribution; Inputting the noisy image feature, the sample texture control feature corresponding to the sample texture control data, and the noisy time step feature corresponding to the noisy time step into the initial denoising network, predicting the noise in the noisy image feature through the initial denoising network, and determining the noise in the noisy image feature as the first noise; According to the first noise and the added Gaussian noise, parameters of the initial denoising network are adjusted to obtain the denoising network.

11. The method according to claim 10, characterized in that The method further comprises: Acquire a sample geometric image of the sample object at each of the S viewing angles, and perform image stitching on the S sample geometric images to obtain a sample control image; Determine an initial control network according to the denoising network, input the noisy image feature, the sample texture control feature and the noisy time step feature into the denoising network, input the noisy image feature, the sample texture control feature, the noisy time step feature and the sample control image feature corresponding to the sample control image into the initial control network, predict the noise in the noisy image feature through the denoising network and the initial control network, and determine the noise in the noisy image feature as the second noise; According to the second noise and the added Gaussian noise, parameters of the initial control network are adjusted to obtain the control network.

12. A data processing device, characterized in that: include: An image acquisition module, used to acquire a geometric image of the spatial object at each of S viewing angles, and to perform image stitching on the S geometric images to obtain a control image; S is an integer greater than 1; A first denoising module is used to obtain texture control data, and perform denoising on the noisy image according to the texture control data and the control image to obtain a predicted image corresponding to the spatial object; the predicted image includes a sub-predicted image of the spatial object at each viewing angle; A second denoising module is used to perform denoising processing on the predicted image to obtain a material image corresponding to the spatial object; the material image includes a sub-material image of the spatial object at each viewing angle; A multi-perspective baking module is used to perform multi-perspective baking on the spatial object according to the material image to obtain a multi-perspective baked spatial object; the multi-perspective baked spatial object has the texture indicated by the texture control data.

13. A computer device, characterized in that: include: Processor and memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer-readable storage medium and is suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 11.