Training method of image correction model, image correction method, device and equipment
By training the image correction model, using viewing angle diagrams and reference diagrams in multiple sets of training samples, the noise is automatically predicted and removed, and the problem of low image correction efficiency in the prior art is solved, and efficient and automatic image correction is achieved.
Patent Information
- Application Number
- CN202510098889.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-20
AI Technical Summary
In the prior art, image correction efficiency is low, and texture defects in images need to be manually corrected.
A training method for image correction model is provided. By obtaining multiple sets of training samples, including a first view, a second view, and a reference diagram of the three-dimensional model, the image correction model is used to process and predict these images, and the model is iteratively trained to predict and remove noise and correct the view diagram.
Improves image correction efficiency and quality, and enables automatic prediction and removal of noise in viewing maps without manual correction.
Smart Images

Figure CN120047343A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a training method for an image correction model, an image correction method, a device and equipment. Background Art
[0002] Image correction is an important task in image processing. If an image has defects such as blurred texture or texture error, the quality of the image will be reduced. Therefore, the image can be corrected by correcting the defects in the image to obtain a high-quality image.
[0003] In the related art, in order to obtain high-quality images, when correcting images, it is generally necessary to manually correct the images based on experience, which makes the image correction inefficient. Summary of the invention
[0004] The embodiment of the present application provides a training method for an image correction model, an image correction method, a device and an apparatus, which improves the efficiency and quality of image correction. The technical solution provided by the embodiment of the present application is as follows.
[0005] According to one aspect of an embodiment of the present application, a method for training an image correction model is provided, the method comprising:
[0006] Acquire multiple groups of first training samples, each group of first training samples includes a first perspective image, a second perspective image, and a reference image of a three-dimensional model, the texture of the three-dimensional model in the first perspective image has defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model;
[0007] Inputting the first viewing angle image, the second viewing angle image and the reference image into an image correction model, respectively processing the first viewing angle image, the second viewing angle image and the reference image through the image correction model to obtain a first viewing angle feature, a second viewing angle feature and a reference feature, adding a first noise feature to the second viewing angle feature to obtain a third viewing angle feature, performing prediction based on the first viewing angle feature, the third viewing angle feature and the reference feature to obtain a second noise feature used to represent a noise feature in the third viewing angle feature, and the first viewing angle feature, the second viewing angle feature and the reference feature are respectively used to describe the first viewing angle image, the second viewing angle image and the reference image;
[0008] The image correction model is iteratively trained based on the difference information between the first noise feature and the second noise feature of each of the multiple groups of first training samples. The image correction model is used to predict the noise in the viewing angle image and correct the viewing angle image by removing the noise.
[0009] According to another aspect of the embodiments of the present application, there is provided an image correction method, the method comprising:
[0010] Acquire a noise map, a viewing angle map of a three-dimensional model, and a reference map, wherein the texture of the three-dimensional model in the viewing angle map has defects, and the reference map is a viewing angle map that provides appearance features of the three-dimensional model;
[0011] Inputting the noise map, the viewing angle map and the reference map into an image correction model, respectively processing the noise map, the viewing angle map and the reference map through the image correction model to obtain noise features, viewing angle features and reference features, performing prediction based on the noise features, the viewing angle features and the reference features to obtain predicted noise features for representing noise features other than viewing angle features in the noise features, removing the predicted noise features from the noise features to obtain corrected viewing angle features, and obtaining a corrected viewing angle map based on the corrected viewing angle features;
[0012] The noise feature, the viewing angle feature and the reference feature are used to describe the noise image, the viewing angle image and the reference image respectively.
[0013] According to another aspect of an embodiment of the present application, a training device for an image correction model is provided, the device comprising:
[0014] an acquisition module, configured to acquire multiple groups of first training samples, each group of first training samples comprising a first perspective image, a second perspective image, and a reference image of a three-dimensional model, wherein the texture of the three-dimensional model in the first perspective image has defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model;
[0015] an input-output module, configured to input the first viewing angle image, the second viewing angle image, and the reference image into an image correction model, process the first viewing angle image, the second viewing angle image, and the reference image respectively through the image correction model to obtain a first viewing angle feature, a second viewing angle feature, and a reference feature, add a first noise feature to the second viewing angle feature to obtain a third viewing angle feature, perform prediction based on the first viewing angle feature, the third viewing angle feature, and the reference feature to obtain a second noise feature used to represent a noise feature in the third viewing angle feature, and the first viewing angle feature, the second viewing angle feature, and the reference feature are used to describe the first viewing angle image, the second viewing angle image, and the reference image respectively;
[0016] A training module is used to iteratively train the image correction model based on the difference information between the first noise characteristics and the second noise characteristics of each of the multiple groups of first training samples. The image correction model is used to predict the noise in the perspective image and correct the perspective image by removing the noise.
[0017] In some embodiments, the input-output module is used to:
[0018] splicing the first viewing angle feature and the third viewing angle feature to obtain a splicing feature;
[0019] Prediction is performed based on the splicing feature and the reference feature to obtain the second noise feature.
[0020] In some embodiments, the input-output module is used to:
[0021] The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features, wherein the convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range;
[0022] The processed splicing features and the reference features are processed by the attention module of the image correction model to obtain cross-attention features, wherein the query parameters in the attention module are the processed splicing features, and the key parameters and value parameters in the attention module are both the reference features;
[0023] The processed splicing features and the cross-attention features are fused through the denoising module of the image correction model, and prediction is performed based on the fused features to obtain the second noise feature.
[0024] In some embodiments, the acquisition module is further used to:
[0025] Based on a pre-trained model, the image correction model is obtained, the pre-trained model is trained based on multiple groups of second training samples, the model parameters of the convolution module corresponding to the second perspective feature, the model parameters of the attention module and the model parameters of the denoising module in the image correction model are respectively the model parameters of the convolution module, the attention module and the denoising module in the pre-trained model, and the second training sample includes a sample image and the sample image with noise added.
[0026] In some embodiments, the input-output module is used to:
[0027] The first viewing angle image and the second viewing angle image are processed respectively by the first encoding module of the image correction model to obtain the first viewing angle feature and the second viewing angle feature, wherein the first viewing angle feature is a representation of the first viewing angle image in a latent space, and the second viewing angle feature is a representation of the second viewing angle image in the latent space;
[0028] The reference image is processed by the second encoding module of the image correction model to obtain the reference feature, where the reference feature is a representation of the reference image in the embedding space.
[0029] According to another aspect of an embodiment of the present application, there is provided an image correction device, the device comprising:
[0030] An acquisition module, used to acquire a noise map, a viewing angle map of a three-dimensional model, and a reference map, wherein the texture of the three-dimensional model in the viewing angle map has defects, and the reference map is a viewing angle map that provides appearance features of the three-dimensional model;
[0031] An input-output module, used for inputting the noise map, the viewing angle map and the reference map into an image correction model, processing the noise map, the viewing angle map and the reference map respectively through the image correction model to obtain noise features, viewing angle features and reference features, making predictions based on the noise features, the viewing angle features and the reference features to obtain predicted noise features for representing noise features other than viewing angle features in the noise features, removing the predicted noise features from the noise features to obtain corrected viewing angle features, and obtaining a corrected viewing angle map based on the corrected viewing angle features;
[0032] The noise feature, the viewing angle feature and the reference feature are used to describe the noise image, the viewing angle image and the reference image respectively.
[0033] In some embodiments, the input-output module is used to:
[0034] Splicing the noise feature and the viewing angle feature to obtain a splicing feature;
[0035] Prediction is performed based on the splicing feature and the reference feature to obtain the predicted noise feature.
[0036] In some embodiments, the input-output module is used to:
[0037] The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features, wherein the convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range;
[0038] The processed splicing features and the reference features are processed by the attention module of the image correction model to obtain cross-attention features, wherein the query parameters in the attention module are the processed splicing features, and the key parameters and value parameters in the attention module are both the reference features;
[0039] The processed splicing features and the cross-attention features are fused through the denoising module of the image correction model, and prediction is performed based on the fused features to obtain the predicted noise features.
[0040] In some embodiments, the input-output module is used to:
[0041] The noise map and the viewing angle map are processed respectively by the first encoding module of the image correction model to obtain the noise feature and the viewing angle feature, wherein the noise feature is a representation of the noise map in a latent space, and the viewing angle feature is a representation of the viewing angle map in the latent space;
[0042] The reference image is processed by the second encoding module of the image correction model to obtain the reference feature, where the reference feature is a representation of the reference image in the embedding space.
[0043] In some embodiments, the input-output module is used to:
[0044] The corrected viewing angle feature is processed by the decoding module of the image correction model to obtain the corrected viewing angle map.
[0045] In some embodiments, the input-output module is used to:
[0046] Based on the noise feature, the viewing angle feature and the reference feature, multiple first iterative calculations are performed until a first iteration stop condition is reached to obtain the modified viewing angle feature; the multiple first iterative calculations include:
[0047] In a first iterative calculation, a prediction is performed based on the noise feature, the viewing angle feature and the reference feature to obtain a first predicted noise feature, wherein the first predicted noise feature is used to represent noise features other than the viewing angle feature in the noise feature, and the first predicted noise feature is removed from the noise feature to obtain a first modified viewing angle feature;
[0048] In the Mth first iterative calculation, prediction is performed based on the M-1th corrected viewing angle feature, the viewing angle feature and the reference feature to obtain the Mth predicted noise feature, wherein the Mth predicted noise feature is used to represent the noise feature other than the viewing angle feature in the M-1th corrected viewing angle feature, and the Mth predicted noise feature is removed from the M-1th corrected viewing angle feature to obtain the Mth corrected viewing angle feature, where M is an integer greater than 1.
[0049] In some embodiments, the apparatus further comprises:
[0050] A first correction module, configured to correct a plurality of viewing angles of the three-dimensional model by using the image correction model to obtain a plurality of corrected viewing angles, wherein the plurality of viewing angles are obtained by rendering based on a texture map of the three-dimensional model, wherein the texture map includes textures at a plurality of positions on the three-dimensional model;
[0051] The second correction module is used to correct the texture map based on the multiple viewing angle maps and the multiple corrected viewing angle maps.
[0052] In some embodiments, the second correction module is used to:
[0053] Based on the multiple viewing angle images and the multiple corrected viewing angle images, multiple second iterative calculations are performed until a second iteration stop condition is reached; the multiple second iterative calculations include:
[0054] In the first second iteration calculation, respectively determining the difference information between each viewing angle image and the corresponding corrected viewing angle image, and correcting the texture image based on the difference information between the multiple viewing angle images and the multiple corrected viewing angle images;
[0055] In the Nth second iteration calculation, the three-dimensional model is rendered based on the texture map corrected in the N-1th second iteration calculation and the multiple perspectives corresponding to the multiple perspective maps to obtain the Nth multiple perspective maps, and the difference information between each perspective map of the Nth and the corresponding corrected perspective map is determined respectively; based on the difference information between the Nth multiple perspective maps and the multiple corrected perspective maps, the texture map corrected in the N-1th second iteration calculation is corrected, where N is an integer greater than 1.
[0056] On the other hand, a computer device is provided, which includes a processor and a memory, the memory being used to store a computer program, the computer program being loaded and executed by the processor to implement the image correction model training method or the image correction method in the embodiment of the present application.
[0057] On the other hand, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the image correction model training method or image correction method in the embodiment of the present application.
[0058] On the other hand, a computer program product is provided, which includes a computer program, the computer program is stored in a computer-readable storage medium, a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the image correction model training method or the image correction method described in any of the above-mentioned implementation methods.
[0059] The embodiment of the present application provides a method for training an image correction model. During the training process, the features of each image are first extracted through the image correction model, and then noise features are added to the perspective features corresponding to the second perspective image after the texture defects are corrected. The added noise features are then predicted based on the perspective features and reference features corresponding to the first perspective image with texture defects, and the image correction model is trained based on the difference information between the predicted noise features and the actually added noise features, so that the image correction model has the ability to accurately predict the noise features. After the image correction model is trained by this method, for any perspective image with texture defects, the image correction model can predict the noise corresponding to the texture defects in the perspective image, and correct the perspective image by removing the noise without manual correction, thereby improving the correction efficiency. Moreover, when predicting noise features, the first viewing angle features and reference features are combined. Since the first viewing angle features provide prior information of the first viewing angle image, the predicted noise features are noise features that can make the viewing angle image close to the first viewing angle image but correct texture defects after being removed. Since the reference features provide appearance control information of the viewing angle image, the predicted noise features are noise features that can make the viewing angle image match the overall appearance of the three-dimensional model after being removed. That is, this method makes the noise features predicted by the image correction model accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;
[0061] Figure 2 It is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;
[0062] Figure 3 is a flowchart of a method for training an image correction model provided by an embodiment of the present application;
[0063] Figure 4 is a flowchart of a method for training an image correction model provided by an embodiment of the present application;
[0064] Figure 5 It is a schematic diagram of the principle of a pre-training model provided by an embodiment of the present application;
[0065] Figure 6 is a flowchart of a method for training an image correction model provided by an embodiment of the present application;
[0066] Figure 7 is a flow chart of an image correction method provided by an embodiment of the present application;
[0067] Figure 8is a flow chart of an image correction method provided by an embodiment of the present application;
[0068] Fig. 9 is a flow chart of an image correction method provided by an embodiment of the present application;
[0069] Fig.10 is a flowchart of texture map correction provided by an embodiment of the present application;
[0070] Fig.11 This is a comparison diagram of the effects before and after the texture image is corrected, provided by an embodiment of the present application;
[0071] Fig.12 is a flow chart of an image correction method provided by an embodiment of the present application;
[0072] Fig.13 is a flowchart of texture map correction provided by an embodiment of the present application;
[0073] Fig.14 is a block diagram of a training device for an image correction model provided by an embodiment of the present application;
[0074] Fig.15 is a block diagram of an image correction device provided by an embodiment of the present application;
[0075] Fig.16 It is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0076] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0077] Please refer to Figure 1 , which shows a schematic diagram of a training method for an image correction model provided by an embodiment of the present application. The method includes at least one of the following steps:
[0078] 1. Obtain training samples: Obtain multiple groups of training samples for training the image correction model, each group of training samples includes a first perspective image 101, a second perspective image 102 and a reference image 103 of the three-dimensional model, the first perspective image 101 is a perspective image with texture defects, the second perspective image 102 is the first perspective image after the defects are corrected, and the reference image 103 is a perspective image used to show the appearance of the three-dimensional model.
[0079] 2. Predicting noise features: Since there are texture defects in the first viewing angle image 101, the texture defects in the first viewing angle image 101 are corrected by the image correction model to obtain a high-quality viewing angle image. The first viewing angle image 101, the second viewing angle image 102 and the reference image 103 are processed respectively to obtain the first viewing angle feature 104, the second viewing angle feature 105 and the reference feature 106 corresponding to each other. Then, the first noise feature 107 is added to the second viewing angle feature 105 to obtain the third noise feature 108, and the first viewing angle feature 104 and the reference feature 106 are combined for prediction to obtain the second noise feature 109 used to represent the added noise feature.
[0080] 3. Training model: Based on the difference information between the predicted second noise feature 109 and the actually added first noise feature 107, iteratively train the image correction model so that the image correction model can predict accurate noise features.
[0081] In an embodiment of the present application, the noise characteristics are predicted by training an image correction model. Subsequently, when any viewing angle image with texture defects is given, the noise characteristics therein are predicted by the image correction model. The noise characteristics correspond to the part of the viewing angle image with texture defects. By removing the noise characteristics, the viewing angle image after the texture defects are corrected can be obtained.
[0082] Please refer to Figure 2 , which shows a schematic diagram of a computer system provided by an embodiment of the present application. The computer system includes at least one of the following: a terminal device 10 and a server 20.
[0083] The terminal device 10 may be an electronic device such as a mobile phone, a tablet computer, a multimedia playback device, a PC (Personal Computer), a wearable device, a vehicle-mounted terminal device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, an MR (Mixed Reality) device, etc. The terminal device 10 may have a client that runs a target application. The target application may be an application that needs to use a three-dimensional model, or an application for building a three-dimensional model, which is not limited in the embodiments of the present application. The embodiments of the present application do not limit the implementation form of the above-mentioned target application, for example, it may be an application that needs to be downloaded and installed, a small program that does not need to be installed, a web application, etc. For example, the target application may include at least one of the following: a game application, a video application, a game engine application, and a modeling application. The texture image is loaded into the target application to render the three-dimensional model.
[0084] In an embodiment of the present application, the server 20 is used to provide background services for the target application. For example, the server 20 is used to correct the perspective of the three-dimensional model. The server 20 is embedded with an image correction model, and the server 20 corrects the perspective of the three-dimensional model through the image correction model. In other embodiments, the server 20 provides the perspective of the three-dimensional model to the terminal 10, and the terminal 10 is embedded with an image correction model, and the terminal 10 corrects the perspective provided by the server 20 through the image correction model.
[0085] The server 20 may be a single server or a server cluster composed of multiple servers. The terminal device 10 may communicate with the server 20 via a network, such as a wireless or wired network.
[0086] Please refer to Figure 3 , Figure 3 This is a flowchart of a training method for an image correction model provided by an embodiment of the present application. The execution subject of each step of the method may be a computer device, such as the execution subject of each step of the method may be a terminal or a server. The method is described by taking the execution subject as a server as an example, and the method may include at least one of the following steps 301 to 303.
[0087] 301. Obtain multiple groups of first training samples, each group of first training samples includes a first perspective image, a second perspective image, and a reference image of a three-dimensional model, the first perspective image has texture defects of the three-dimensional model, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model.
[0088] In the embodiment of the present application, the three-dimensional model is composed of points, lines, and surfaces. The three-dimensional model can be used to represent entities in the real world, and can also be used to represent virtual objects. The three-dimensional model can be a model with a three-dimensional structure constructed in a virtual three-dimensional space.
[0089] In some embodiments, the first perspective image is an image of the three-dimensional model at a certain perspective, which may be a camera perspective. At this perspective, the three-dimensional model is rendered based on the texture map of the three-dimensional model to obtain the first perspective image. A rendering engine is used to render the texture on the texture map of the three-dimensional model onto the three-dimensional model to obtain the perspective image. The resolution of the perspective image can be set as needed, such as 512*512.
[0090] In an embodiment of the present application, the first perspective image can be a full-body perspective image of the three-dimensional model, or it can be a three-dimensional partial perspective image, such as the first perspective image is a face perspective image or a leg perspective image of the three-dimensional model.
[0091] The texture map of a three-dimensional model includes textures at multiple locations on the three-dimensional model. The surface of the three-dimensional model is composed of triangles or polygons, and the texture map is a two-dimensional plane map obtained by unfolding the surface of the three-dimensional model into a two-dimensional plane. The coordinates of the vertices on the texture map are called texture coordinates. Each vertex in the three-dimensional model has a corresponding texture coordinate on the texture map. The texture coordinates are used to indicate the coordinates of the vertex on the three-dimensional model on the texture map. When rendering the texture, the texture at the position indicated by the texture coordinate of the vertex on the texture map is mapped to the vertex on the three-dimensional model to render the texture onto the three-dimensional model.
[0092] The texture of a 3D model refers to a 2D image mapped on the surface of a 3D object, which is used to enhance the visual effect of the object and make it look more real and detailed. Texture usually includes properties such as color, smoothness, and reflectivity. Texture plays a key role in the rendering process of 3D models.
[0093] In the embodiment of the present application, the defects of the texture include at least one of texture blur, texture error, insufficient texture detail clarity, etc. Texture blur may refer to blunt texture edges, color distortion, etc., and texture error may refer to geometric distortion of texture lines, etc.
[0094] In the embodiments of the present application, the reference image is a perspective view for displaying the appearance of the three-dimensional model. Further, the reference image is a perspective view that provides as comprehensive texture features and structural features of the appearance of the three-dimensional model as possible, such as the reference image is a front view of the three-dimensional model.
[0095] In an embodiment of the present application, multiple three-dimensional models may be obtained first, and then multiple groups of first training samples may be obtained through the three-dimensional models. Among them, since the perspective images of the three-dimensional model can be obtained from different perspectives, then optionally, multiple groups of first training samples can be obtained based on multiple three-dimensional models, and one or more groups of first training samples can be obtained based on each three-dimensional model. Multiple groups of first training samples corresponding to the same three-dimensional model may include first perspective images obtained from different perspectives and the same reference image. For example, the multiple first perspective images included in at least part of the first training samples may be the front view, top view, side view or other perspective view images of the same three-dimensional model, etc., which are not specifically limited here.
[0096] 302. Input the first perspective image, the second perspective image and the reference image into the image correction model, process the first perspective image, the second perspective image and the reference image respectively through the image correction model to obtain the first perspective feature, the second perspective feature and the reference feature, add the first noise feature to the second perspective feature to obtain the third perspective feature, perform prediction based on the first perspective feature, the third perspective feature and the reference feature to obtain the second noise feature for representing the noise feature in the third perspective feature, and the first perspective feature, the second perspective feature and the reference feature are respectively used to describe the first perspective image, the second perspective image and the reference image.
[0097] In an embodiment of the present application, the first viewing angle feature, the second viewing angle feature, the third viewing angle feature, the reference feature, the first noise feature and the second noise feature may be vectors or matrices. In this embodiment, these features are described as matrices as an example.
[0098] Optionally, the first noise feature is a feature representation of random noise, such as the first noise feature is a feature representation of random Gaussian noise. The first noise feature has the same dimension as the second perspective feature, so it is convenient to add the first noise feature to the second perspective feature by summing. Accordingly, the first noise feature and the second noise feature have the same dimension, so it is convenient to determine the difference information between the two.
[0099] In an embodiment of the present application, a first noise feature is added to a second viewing angle feature, so that a second viewing angle image corresponding to the second viewing angle feature produces a texture defect. The first noise feature also corresponds to a feature of a texture defect portion in the second viewing angle image. Therefore, an image correction model is trained to predict the noise in the viewing angle image, that is, to predict the feature of the corresponding texture defect portion in the viewing angle image. Therefore, the texture defect in the viewing angle image can be corrected by removing the noise.
[0100] In an embodiment of the present application, the first perspective image is a perspective image with texture defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model. The noise characteristics in the first perspective image are predicted based on these three items. In this way, when predicting the noise characteristics, there is a priori knowledge of the first perspective image to generate a perspective image that is close to the first perspective image but of higher quality (i.e., the defects are corrected), and the reference image is used as a conditional constraint on the appearance of the three-dimensional model, so that the generated perspective image is not only of high quality in local appearance, but also matches the three-dimensional model in overall appearance, so that the corrected perspective image is of higher quality.
[0101] 303. Iteratively train an image correction model based on difference information between first noise features and second noise features of each of the multiple groups of first training samples, where the image correction model is used to predict noise in the viewing angle map and correct the viewing angle map by removing the noise.
[0102] In an embodiment of the present application, with the goal of minimizing the difference information between the first noise feature and the second noise feature, the parameters of the image correction model are iteratively adjusted so that the second noise feature predicted by the image correction model gradually approaches the actually added first noise feature.
[0103] Among them, by training the image correction model to predict the ability of noise features, after any viewing feature with texture defects is given, the image correction model is used to predict the noise features in the viewing feature, and then the corrected viewing feature can be obtained by removing the noise features in the viewing feature, thereby obtaining the corrected viewing feature, and then obtaining the corrected viewing map.
[0104] The embodiment of the present application provides a method for training an image correction model. During the training process, the features of each image are first extracted through the image correction model, and then noise features are added to the perspective features corresponding to the second perspective image after the texture defects are corrected. The added noise features are then predicted based on the perspective features and reference features corresponding to the first perspective image with texture defects, and the image correction model is trained based on the difference information between the predicted noise features and the actually added noise features, so that the image correction model has the ability to accurately predict the noise features. After the image correction model is trained in this way, for any perspective image with texture defects, the image correction model can predict the noise corresponding to the texture defects in the perspective image, and correct the perspective image by removing the noise without manual correction, thereby improving the correction efficiency. Moreover, when predicting noise features, the first viewing angle features and reference features are combined. Since the first viewing angle features provide prior information of the first viewing angle image, the predicted noise features are noise features that can make the viewing angle image close to the first viewing angle image but correct texture defects after being removed. Since the reference features provide appearance control information of the viewing angle image, the predicted noise features are noise features that can make the viewing angle image match the overall appearance of the three-dimensional model after being removed. That is, this method makes the noise features predicted by the image correction model accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0105] See also Figure 4 , Figure 4 This is a flowchart of a training method for an image correction model provided by an embodiment of the present application. The execution subject of each step of the method may be a computer device, such as the execution subject of each step of the method may be a terminal or a server. The method is described by taking the execution subject as a server as an example, and the method may include at least one of the following steps 401 to 407.
[0106] 401. Obtain multiple groups of first training samples, each group of first training samples includes a first perspective image, a second perspective image, and a reference image of a three-dimensional model, the first perspective image has texture defects of the three-dimensional model, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model.
[0107] In the embodiment of the present application, step 401 is the same as step 301 and will not be described in detail here.
[0108] 402. Input the first perspective image, the second perspective image and the reference image into the image correction model, and process the first perspective image and the second perspective image respectively through the first encoding module of the image correction model to obtain a first perspective feature and a second perspective feature, wherein the first perspective feature is a representation of the first perspective image in the latent space, and the second perspective feature is a representation of the second perspective image in the latent space.
[0109] In the embodiment of the present application, the first perspective feature is a matrix, the first perspective feature is a low-dimensional representation of the first perspective image in the latent space, and the first perspective feature is obtained by compressing and reducing the dimension of the first perspective image through the first encoding module. Similarly, the second perspective feature is a matrix, the second perspective feature is a low-dimensional representation of the second perspective image in the latent space, and the second perspective feature is obtained by compressing and reducing the dimension of the second perspective image through the first encoding module.
[0110] For example, the dimensions of the first viewing angle image and the second viewing angle image are 4*512*512 respectively. After being processed by the first encoding module, the first viewing angle feature and the second viewing angle feature both have dimensions of 4*64*64.
[0111] The first encoding module is an encoder, and the encoder may be an encoder in a variational autoencoder (VAE). Accordingly, the model parameters in the first encoding module are pre-trained, and the first encoding module can be directly used to process the first view image and the second view image without further training the first encoding module, so as to improve the training efficiency of the image correction model.
[0112] 403. Process the reference image through the second encoding module of the image correction model to obtain a reference feature, where the reference feature is a representation of the reference image in the embedding space.
[0113] In the embodiment of the present application, the reference feature is a matrix, and the reference image is converted into a matrix form so that the reference image can participate in the calculation in the image correction model.
[0114] The second encoding module is an encoder, and the encoder can be an image encoder in a contrastive language-image pre-training model (CLIP). Accordingly, the model parameters in the second encoding module are pre-trained, and the second encoding module can be directly used to process the reference image without further training the second encoding module, so as to improve the training efficiency of the image correction model.
[0115] It should be noted that the serial numbers of step 402 and step 403 are only for the convenience of explanation and are not used to limit the execution order of the two. For example, step 402 can be executed before step 403, or after step 403, or can be executed simultaneously. In this embodiment, the synchronous execution of the two is taken as an example to improve efficiency.
[0116] In an embodiment of the present application, the above steps 402 and 403 implement the process of processing the first perspective image, the second perspective image and the reference image respectively through the image correction model to obtain the first perspective feature, the second perspective feature and the reference feature. In this embodiment, the first perspective image and the second perspective image are mapped to the latent space by the first encoding module. Since the latent space usually has a lower dimension than the original image space, it helps to reduce the number of parameters of the model, thereby reducing the computational complexity and improving the training efficiency. In addition, the low-dimensional characteristics of the latent space help the model capture the essential characteristics of the data, thereby improving the generalization ability. By mapping the reference image to the embedding space through the second encoding module, feature extraction is realized for the reference image, which facilitates the image correction model to understand the reference image and improves the efficiency and accuracy of the image correction model in predicting noise features based on the reference image.
[0117] 404. Add the first noise feature to the second viewing angle feature through the image correction model to obtain a third viewing angle feature.
[0118] In some embodiments, the first noise feature and the second perspective feature have the same dimension, such as if the dimensions are both 4*64*64, the first noise feature can be added to the second perspective feature by summing up to obtain a third perspective feature with the same dimension as the second perspective feature. The third perspective feature is also the second perspective feature carrying the first noise feature.
[0119] In some embodiments, the image correction model includes a noise adding module, and the image correction model adds the first noise feature to the second viewing angle feature through the noise adding module.
[0120] 405. The first-view feature and the third-view feature are spliced together using the image correction model to obtain a spliced feature.
[0121] In the embodiment of the present application, if the dimensions of the first viewing angle feature and the third viewing angle feature are the same, the first viewing angle feature and the third viewing angle feature can be spliced in a certain sub-dimension to obtain a spliced feature. For example, if the dimensions of the two viewing angle features are 4*64*64 respectively, the two viewing angle features can be spliced in the channel sub-dimension to obtain a spliced feature with a dimension of 8*64*64.
[0122] 406. Predicting based on the splicing feature and the reference feature using the image correction model to obtain a second noise feature.
[0123] In some embodiments, the image correction model includes a convolution module, a cross-attention module and a denoising module. Accordingly, the above process of predicting the second noise feature based on the splicing feature and the reference feature through the image correction model includes the following steps: convolution processing is performed on the splicing feature through the convolution module of the image correction model to obtain the processed splicing feature, and the convolution processing is used to extract the splicing feature and limit the dimension of the splicing feature to a preset range; the processed splicing feature and the reference feature are processed through the attention module of the image correction model to obtain the cross-attention feature, the query parameter in the attention module is the processed splicing feature, and the key parameter and value parameter in the attention module are both reference features; the processed splicing feature and the cross-attention feature are fused through the denoising module of the image correction model, and prediction is performed based on the fused features to obtain the second noise feature.
[0124] The convolution module includes one or more convolution layers. The number of channels of the first convolution layer is the same as the channel dimension of the spliced features, so that the first convolution layer can process the spliced features. For example, if the dimension of the spliced features is 8*64*64, the number of channels of the first convolution layer is 8. The convolution module is used to extract preliminary features from the spliced features, which are subsequently used for deeper processing. Accordingly, the first convolution layer will increase the dimension of the spliced features, such as from 8 dimensions to 320 dimensions, so as to provide richer reference information for feature extraction in the subsequent downsampling stage.
[0125] Among them, since the cross-attention mechanism is used to extract the most important features in the input data, and by using the processed spliced features as query parameters and the reference features as key parameters and value parameters, the reference features and the view feature representation in the latent space are combined, so that the model considers additional conditional information in the process of predicting noise, thereby making the prediction result more accurate. Using the reference features as key parameters and value parameters can enhance the control of the reference image, that is, enhance the association between the denoised view image and the reference image, thereby improving the correction quality of the view image.
[0126] The fusion mode of the processed splicing features and the cross-attention features can be a superposition mode or a splicing mode, which is not specifically limited here. Optionally, the dimensions of the processed splicing features and the cross-attention features are the same, which facilitates the fusion of the two. The denoising module can be predicted by a convolutional network or other neural network, which is not specifically limited here.
[0127] It should be noted that the convolution module, the attention module and the denoising module are all modules in the prediction unit of the image correction model, and the prediction unit is, for example, a Unet (U-shaped network) network. The Unet network is a convolutional neural network for image segmentation, which can process the input image features at the pixel level and has good accuracy and robustness. The Unet network can process image features through the various processing modules it includes. For example, in this embodiment, the Unet network includes a convolution module, an attention module and a denoising module, and the Unet network predicts noise features through these modules.
[0128] In the embodiment of the present application, the splicing features are first processed by the convolution module to facilitate the subsequent processing of the splicing features by the attention module. The reference features and the processed splicing features are combined by the attention module, so that additional conditional information is considered in the process of predicting noise, and then the splicing features and the attention cross features are combined for prediction, so that the prediction results are more accurate, and the association between the denoised view image and the reference image can be enhanced, thereby improving the correction quality.
[0129] In some embodiments, before obtaining multiple groups of first training samples, an initial image correction model is first obtained. Optionally, the image correction model is an untrained model. Alternatively, the image correction model is obtained based on a pre-trained model, and the pre-trained model is trained based on multiple groups of second training samples. The model parameters of the convolution module corresponding to the second viewing angle feature, the model parameters of the attention module, and the model parameters of the denoising module in the image correction model are respectively the model parameters of the convolution module, the attention module, and the denoising module in the pre-trained model, and the second training sample includes a sample image and a sample image with noise added.
[0130] Among them, the pre-trained model is a generalized image generation model, which also generates images based on the denoising principle. For example, the pre-trained model is a diffusion model (Stable Diffusion Models). When using the pre-trained model to generate an image, a random noise is processed based on the text prompt information, a noise feature is predicted, and the predicted noise feature is removed from the random noise to obtain the predicted image. Accordingly, when the pre-trained model is trained, noise is added to the sample image, and the ability of the pre-trained model to predict noise features is trained based on the sample image and the sample image with added noise, so that the pre-trained model can generate an image by removing the noise from the random noise.
[0131] In this embodiment, an initial image correction model is obtained based on a pre-trained model, so that model training is performed based on the pre-trained model, which can reduce the number of training samples, reduce the number of iterations, and improve training efficiency. In addition, since the training samples of the pre-trained model are sample images and sample images with added noise, that is, the pre-trained model generates images based on the denoising principle, and the training samples of the pre-trained model are not limited to the perspective diagram of the three-dimensional model, the model parameters of the pre-trained model can also ensure the generalization of the image correction model, that is, the image correction module can achieve a controllable image correction function while retaining the generalization of big data.
[0132] It should be noted that the convolution module of the pre-trained model is generally only used to process a single perspective feature, and is not used to process splicing features with higher dimensions. Therefore, the number of channels of the convolution module of the pre-trained model does not correspond to the dimension of the splicing feature, and its number of channels only corresponds to the dimension of a single perspective feature. Therefore, the model parameters of the convolution module of the pre-trained model can only provide model parameters corresponding to the second perspective feature, and for the model parameters corresponding to the first perspective feature, it is necessary to increase the number of channels in the convolution module of the image correction model, and the model parameters corresponding to each newly added channel are trained from scratch, that is, the initial value of the model parameters corresponding to each newly added channel is zero. For example, the number of channels of the convolution module of the pre-trained model is 4, and the number of channels of the convolution module of the image correction model is 8, then the convolution module of the image correction model needs to add 4 channels on the basis of the 4 channels of the convolution module of the pre-trained model.
[0133] In some embodiments, the pre-trained model is a diffusion model. Figure 5 As shown in the figure. After multiple denoising, the random noise X3 is converted into the original image X0. Each denoising passes through the Unet network. The Unet network includes a convolution module, an attention module, and a denoising module. During training, the original image X0 can be input and different degrees of noise can be added to it to obtain a noisy image. Figure Xt, as shown in Formula 1 below. And a pre-trained model is used to predict the added noise, as shown in Formula 2 below. The difference information (error) is calculated based on the predicted noise and the actually added noise, and then the pre-trained model is trained based on the difference information to enable it to obtain the correct denoising ability. The pre-trained model can predict various noises during inference and obtain the expected original image X0 to be generated after multiple denoising operations, thereby realizing the image generation ability of the pre-trained model.
[0134] Formula 1:
[0135]
[0136] where x 0 represents the original image, x t represents the original image after adding noise, q(x t |x 0 ) represents the conditional probability distribution of the data x 0 at time step t given the data x t , ∫q(x 1:t |x 0 )dx 1:(t-1) represents the integration over all intermediate states x 1 to x t-1 , represents the overall normal distribution (Gaussian distribution), α t is a time-dependent parameter, usually a noise ratio, used to control the amount of noise added to the image at each step, represents the scaled part of the original image, (1 - α t )I represents the covariance matrix, and I is the identity matrix.
[0137] Formula 2:
[0138]
[0139] where L γ (∈ θ ) represents the loss function, represents the sum from time step 1 to T, where T is the total number of diffusion steps, and γ t represents the weight of a time step t, used to assign different weights to different time steps in the loss function. represents the expectation of the initial data x 0 , where x 0 is sampled from the data distribution q(x 0 ), represents the expectation of the noise ∈ t , where ∈ t is sampled from the standard normal distribution , represents the noise predicted at time step t, Represents the zoomed portion of the original image. represents the noise added at time step t, which is scaled and added to the scaled part of the original image. Represents the square of the Euclidean norm (L2 norm), which is used to calculate the difference information between the predicted noise and the actual noise.
[0140] In an embodiment of the present application, the above steps 405-406 implement a process of predicting based on the first viewing angle feature, the third viewing angle feature and the reference feature to obtain the second noise feature. In this embodiment, by splicing the first viewing angle feature and the second viewing angle feature, the prior of the viewing angle image with texture defects is added to the viewing angle feature, so that when the image correction model predicts the noise feature, the predicted noise feature is the noise feature that can make the viewing angle image close to the first viewing angle image but correct the texture defects after being removed, so that the prediction result is more accurate, thereby improving the correction quality.
[0141] 407. Iteratively train an image correction model based on difference information between the first noise features and the second noise features of each of the multiple groups of first training samples, where the image correction model is used to predict noise in the viewing angle image and correct the viewing angle image by removing the noise.
[0142] In an embodiment of the present application, the difference information between the first noise feature and the second noise feature may be a first loss function value, and the first loss function value is used to indicate the difference between the first noise feature and the second noise feature. Optionally, the first loss function value may be a mean square error loss function value or a cross entropy loss function value, etc., which is not specifically limited here.
[0143] In the embodiment of the present application, each iteration process uses one or at least two groups of first training samples. If each iteration process uses at least two groups of first training samples, the mean of the first loss function values of the at least two groups of first training samples is determined, and the model parameters of the image correction model are adjusted based on the mean. It should be noted that, since multiple groups of first training samples can be obtained through a three-dimensional model, each iteration process can optionally use multiple groups of first training samples corresponding to the same three-dimensional model.
[0144] In an embodiment of the present application, based on the first loss function value, the model parameters of the image correction model are iteratively adjusted until an iteration stop condition is reached. The iteration stop condition includes at least one of the following: the number of iterations reaches a first number; the first loss function value is less than or equal to a first threshold. The iteration stop condition is a criterion for determining when to stop updating the model parameters based on the first loss function value, and the first number is the maximum number of iterations. The first threshold represents the minimum value of the acceptable error of the first loss function value.
[0145] It should be noted that the image correction model includes a first encoding module, a second encoding module, a noise adding module, a decoding module and a Unet network, etc. The decoding module is used to decode the viewing angle feature into a viewing angle map. The Unet network includes a convolution module, an attention module and a denoising module, etc. Optionally, the first encoding module, the second encoding module, the noise adding module and the decoding module are pre-trained modules. If these modules are also included in the pre-training module, when adjusting the model parameters based on the first loss function value, it is only necessary to adjust the model parameters of the modules included in the Unet network, without adjusting the model parameters outside the Unet network, thereby improving the model training efficiency.
[0146] For example, see Figure 6 , Figure 6 It is a training flow chart of an image correction model provided by an embodiment of the present application. Among them, a first perspective image, a second perspective image and a reference image are input into the image correction model. The first perspective image is processed to obtain a first perspective feature. The second perspective image is processed to obtain a second perspective feature, and a first noise feature is added to the second perspective feature to obtain a third perspective feature. The first perspective feature and the third perspective feature are spliced to obtain a spliced feature. The reference image is processed to obtain a reference feature. The spliced feature and the reference feature are input into the Unet network, and predicted by the Unet network to obtain a second noise feature.
[0147] The embodiment of the present application provides a method for training an image correction model. During the training process, the features of each image are first extracted through the image correction model, and then noise features are added to the perspective features corresponding to the second perspective image after the texture defects are corrected. The added noise features are then predicted based on the perspective features and reference features corresponding to the first perspective image with texture defects, and the image correction model is trained based on the difference information between the predicted noise features and the actually added noise features, so that the image correction model has the ability to accurately predict the noise features. After the image correction model is trained in this way, for any perspective image with texture defects, the image correction model can predict the noise corresponding to the texture defects in the perspective image, and correct the perspective image by removing the noise without manual correction, thereby improving the correction efficiency. Moreover, when predicting noise features, the first viewing angle features and reference features are combined. Since the first viewing angle features provide prior information of the first viewing angle image, the predicted noise features are noise features that can make the viewing angle image close to the first viewing angle image but correct texture defects after being removed. Since the reference features provide appearance control information of the viewing angle image, the predicted noise features are noise features that can make the viewing angle image match the overall appearance of the three-dimensional model after being removed. That is, this method makes the noise features predicted by the image correction model accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0148] See also Figure 7 , Figure 7 : is a flowchart of an image correction method provided by an embodiment of the present application. The image correction model used in the method can be an image correction model trained by any of the above embodiments. The execution subject of each step of the method can be a terminal or a server. The method is described by taking the execution subject as a server as an example, and the method can include at least one of the following steps 701 to 702.
[0149] 701. Obtain a noise map, a perspective map of a three-dimensional model, and a reference map. The perspective map has defects in the texture of the three-dimensional model, and the reference map is a perspective map that provides appearance features of the three-dimensional model.
[0150] In the embodiment of the present application, the noise map may be a random noise map, such as a Gaussian noise map. The viewing angle map is the same as the first viewing angle map in step 301, and the reference map is the same as the reference map in step 301, which will not be described in detail.
[0151] 702. Input the noise map, the viewing angle map and the reference map into the image correction model, process the noise map, the viewing angle map and the reference map respectively through the image correction model to obtain noise features, viewing angle features and reference features, make predictions based on the noise features, viewing angle features and reference features to obtain predicted noise features for representing noise features other than viewing angle features in the noise features, remove the predicted noise features from the noise features to obtain corrected viewing angle features, and obtain a corrected viewing angle map based on the corrected viewing angle features; wherein the noise features, viewing angle features and reference features are respectively used to describe the noise map, the viewing angle map and the reference map.
[0152] In the embodiment of the present application, the dimensions of the noise feature and the view feature are the same, so that after removing the predicted noise feature from the noise feature, a modified view feature corresponding to the dimension of the view feature is obtained. The modified view feature is also the representation of the modified view map in the latent space.
[0153] The embodiment of the present application provides an image correction method. When the method uses an image correction model to correct a viewing angle map, the method first extracts the features of each image through the image correction model, and then predicts the noise features by combining the viewing angle features and reference features corresponding to the viewing angle map with texture defects. After removing the predicted noise features, the corrected viewing angle features are obtained. Based on the corrected viewing angle features, the corrected viewing angle map can be obtained without manual correction, thereby improving the correction efficiency. In addition, when predicting the noise features, the viewing angle features and the reference features are combined. Since the viewing angle features provide the prior information of the viewing angle map, the predicted noise features are noise features that can make the viewing angle map close to the original viewing angle map but correct the texture defects after being removed. Since the reference features provide the appearance control information of the viewing angle map, the predicted noise features are noise features that can make the viewing angle map match the overall appearance of the three-dimensional model after being removed. That is, through the viewing angle features and the reference features, the predicted noise features are accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0154] See also Figure 8 , Figure 8 1 is a flowchart of an image correction method provided by an embodiment of the present application. The image correction model used in the method may be an image correction model trained by any of the above embodiments. The execution subject of each step of the method may be a terminal or a server. The method is described by taking the execution subject as a server as an example, and the method may include at least one of the following steps 801-806.
[0155] 801. Obtain a noise map, a perspective map of a three-dimensional model, and a reference map. The perspective map has defects in the texture of the three-dimensional model, and the reference map is a perspective map that provides appearance features of the three-dimensional model.
[0156] In the embodiment of the present application, step 801 is the same as step 701 and will not be described again here.
[0157] 802. Input the noise map, the viewing angle map and the reference map into the image correction model, and process the noise map and the viewing angle map respectively through the first encoding module of the image correction model to obtain noise features and viewing angle features, wherein the noise feature is a representation of the noise map in the latent space, and the viewing angle feature is a representation of the viewing angle map in the latent space.
[0158] In the embodiment of the present application, the process of obtaining the noise feature and the viewing angle feature in step 802 is the same as the process of obtaining the first viewing angle feature and the second viewing angle feature in step 402, and will not be repeated here.
[0159] 803. Process the reference image through the second encoding module of the image correction model to obtain a reference feature, where the reference feature is a representation of the reference image in the embedding space.
[0160] In the embodiment of the present application, the process of obtaining the reference feature in step 803 is the same as the process of obtaining the reference feature in step 403, and will not be repeated here.
[0161] 804. The noise feature and the view feature are spliced together through the image correction model to obtain a spliced feature.
[0162] In the embodiment of the present application, the process of obtaining the splicing features in step 804 is the same as the process of obtaining the splicing features in step 405, and will not be repeated here.
[0163] 805. Predictions are made based on the splicing features and the reference features through the image correction model to obtain predicted noise features.
[0164] In some embodiments, the image correction model includes a convolution module, a cross-attention module and a denoising module. Accordingly, the above process of predicting based on the splicing features and the reference features through the image correction model to obtain the predicted noise features includes the following steps: the splicing features are convolved through the convolution module of the image correction model to obtain the processed splicing features, and the convolution processing is used to extract the splicing features and limit the dimensions of the splicing features to a preset range; the processed splicing features and the reference features are processed through the attention module of the image correction model to obtain the cross-attention features, the query parameters in the attention module are the processed splicing features, and the key parameters and value parameters in the attention module are all reference features; the processed splicing features and the cross-attention features are fused through the denoising module of the image correction model, and the prediction is performed based on the fused features to obtain the predicted noise features.
[0165] In some embodiments, in order to improve the correction quality of the viewing angle map, the above process of predicting the noise features other than the viewing angle features in the noise features based on the noise features, the viewing angle features and the reference features to obtain the predicted noise features, removing the predicted noise features from the noise features to obtain the corrected viewing angle features, also includes the following implementation method: performing multiple first iterative calculations based on the noise features, the viewing angle features and the reference features until the first iterative stop condition is reached to obtain the corrected viewing angle features; the multiple first iterative calculations include the following steps: in the first first iterative calculation, predicting based on the noise features, the viewing angle features and the reference features, The first predicted noise feature is obtained, and the first predicted noise feature is used to represent the noise features other than the viewing angle features in the noise features. The first predicted noise feature is removed from the noise features to obtain the first corrected viewing angle feature. In the Mth first iteration calculation, prediction is performed based on the M-1th corrected viewing angle feature, the viewing angle feature and the reference feature to obtain the Mth predicted noise feature, and the Mth predicted noise feature is used to represent the noise features other than the viewing angle features in the M-1th corrected viewing angle feature. The Mth predicted noise feature is removed from the M-1th corrected viewing angle feature to obtain the Mth corrected viewing angle feature, where M is an integer greater than 1.
[0166] The first iteration stop condition may be that the number of iterations reaches a second number. The second number is the maximum number of iterations, and the second number can be set as needed, such as any number between 20 and 50.
[0167] In this embodiment, the perspective map is subjected to multiple iterative cycles of denoising to improve the denoising effect, so that the noise in the perspective map is removed more thoroughly, that is, the texture defects in the perspective map are corrected more thoroughly, thereby improving the correction effect of the image correction model on the perspective map and improving the correction quality.
[0168] 806. Process the corrected viewing angle feature through a decoding module of the image correction model to obtain a corrected viewing angle map.
[0169] The modified view feature is a representation of the modified view map in the latent space. The decoding module is used to decode the modified view feature to obtain the modified view map. The dimension of the modified view map is the same as the dimension of the original view map, such as 3*512*512.
[0170] The decoding module is a decoder, and the decoder can be a decoder in a variational autoencoder. Accordingly, the model parameters in the decoding module are trained before training the image correction model, and the decoding module can be directly used to obtain the corrected view map without further training the decoding module, so as to improve the training efficiency of the image correction model.
[0171] For example, see Fig. 9 , Fig. 9 This is a flow chart of an image correction method provided by an embodiment of the present application. In which, the noise map, the viewing angle map and the reference map are input into the image correction model, the noise map is processed to obtain the noise feature, the viewing angle map is processed to obtain the viewing angle feature, and the reference map is processed to obtain the reference feature. Based on the noise feature, the viewing angle feature and the reference feature, a prediction is performed to obtain a predicted noise feature, and the predicted noise feature is removed from the noise feature to obtain a corrected noise feature, and a corrected viewing angle map is obtained based on the corrected noise feature.
[0172] In some embodiments, after the perspective map of the three-dimensional model is corrected by the image correction model, the texture map of the three-dimensional model can be further corrected by the corrected perspective map. The process includes the following implementation methods: based on the image correction model, multiple perspective maps of the three-dimensional model are corrected to obtain multiple corrected perspective maps, and the multiple perspective maps are rendered based on the texture map of the three-dimensional model, and the texture map includes textures at multiple positions on the three-dimensional model; based on the multiple perspective maps and the multiple corrected perspective maps, the texture map is corrected.
[0173] Among them, multiple perspective images cover the overall appearance of the three-dimensional model, that is, the textures on the multiple perspective images cover all the textures on the texture map, and then the texture map is fully corrected based on the multiple perspective images.
[0174] In this embodiment, since the viewing angle map is obtained by rendering the three-dimensional model based on the texture map of the three-dimensional model, the texture map can be corrected based on the difference information between the corrected viewing angle map and the viewing angle map before correction. The viewing angle map obtained based on the corrected texture map is close to the corrected viewing angle map, so the viewing angle map directly obtained based on the texture map is no longer including texture defects, and there is no need to correct the viewing angle map, thereby improving the efficiency of obtaining other subsequent viewing angle maps.
[0175] In some embodiments, the process of correcting the texture map based on multiple perspective images and multiple corrected perspective images to obtain the corrected texture map includes the following implementation methods: performing multiple second iterative calculations based on multiple perspective images and multiple corrected perspective images until the second iteration stop condition is reached; the multiple second iterative calculations include the following steps: in the first second iterative calculation, respectively determining the difference information between each perspective image and the corresponding corrected perspective image, and correcting the texture map based on the difference information between the multiple perspective images and the multiple corrected perspective images; in the Nth second iterative calculation, rendering the three-dimensional model based on the corrected texture map in the N-1th second iterative calculation and the multiple perspectives corresponding to the multiple perspective images to obtain the Nth multiple perspective images, respectively determining the difference information between each perspective image of the Nth time and the corresponding corrected perspective image, and correcting the corrected texture map in the N-1th second iterative calculation based on the difference information between the Nth multiple perspective images and the multiple corrected perspective images, where N is an integer greater than 1.
[0176] Optionally, differentiable rendering is performed on the three-dimensional model based on the texture map, and accordingly, back propagation is performed based on difference information between the multiple view maps and the multiple corrected view maps to correct the texture map.
[0177] In the embodiment of the present application, the texture map is corrected with the goal of minimizing the difference information between the perspective map and the corrected perspective map. Optionally, the difference information between each perspective map and the corresponding corrected perspective map is a second loss function value between the two, and the second loss function value can be a mean square error loss function value or a cross entropy loss function value, etc., which is not specifically limited here.
[0178] Optionally, based on the difference information between the multiple perspective images and the multiple corrected perspective images, the process of correcting the texture image includes the following steps: correcting the texture image based on the sum of the second loss function values between the multiple perspective images and the multiple corrected perspective images. The process of determining the sum of the multiple second loss function values is shown in the following formula 3.
[0179]
[0180] Wherein, n represents the number of multiple viewpoints. renderi represents the i-th view image, I gti represents the corrected view image corresponding to the i-th view image, X represents the texture image, T i Represents the i-th viewing angle. The viewing angle in this application refers to the camera viewing angle, that is, rendering is performed based on the i-th camera viewing angle and the texture map to obtain a viewing angle map. Each pixel in the viewing angle map is sampled from the texture map according to the camera viewing angle. The viewing angle maps of different camera viewing angles are different, but they all correspond to the same texture map.
[0181] The second iteration stopping condition may be that the second loss function value reaches convergence, such as the second loss function value is less than or equal to a second threshold value. The second threshold value represents the minimum value of the error that the second loss function value can accept.
[0182] In the embodiment of the present application, the original texture map of the three-dimensional model can be corrected based on the corrected view map; or a texture map can be randomly initialized and the texture map can be corrected based on the corrected view map to obtain a corrected texture map. Fig.10 , Fig.10 This is a flowchart of texture map correction provided by an embodiment of the present application, wherein a texture map is randomly initialized, and the texture map is iteratively corrected based on multiple view maps corrected by the image correction model until convergence is reached, thereby obtaining a texture map corrected by the three-dimensional model.
[0183] For example, see Fig.11 , Fig.11 It is a comparison diagram of the effects before and after the texture map is corrected, provided by an embodiment of the present application. Among them, when the initial texture map is all gray, multiple corrected perspective maps and multiple perspective maps are respectively shown in the first two rows of images. Since the rendering process is the process of sampling from the texture map, and the texture map is all gray at this time, the perspective map is also all gray. After the texture map is corrected, multiple corrected perspective maps and multiple perspective maps are respectively shown in the last two rows of images. Since the texture map is corrected, the perspective map obtained based on the texture map is almost the same as the corrected perspective map. When the value of the second loss function reaches the minimum value, that is, when convergence is reached, the correction of the texture map is completed, and the texture map is output.
[0184] The method provided in the embodiment of the present application can replace the super-resolution effect without changing the resolution of the image, and can handle texture defects that cannot be handled by super-resolution, such as texture errors and lack of texture details. In addition, by inputting a reference image as a control signal, the viewing angle map of the three-dimensional model is optimized to obtain a fine, beautiful, and textured three-dimensional model that is highly correlated with the input reference image. The whole process is short in time and highly efficient, such as only 10 seconds. In addition, the method provided in the embodiment of the present application can be applied to a post-processing module to correct a three-dimensional model with fuzzy textures generated by other algorithms, thereby improving the quality and usability of the generated three-dimensional model and reducing the production cost. At the same time, it can be decoupled from other algorithms and can be universally and flexibly applied after each generation algorithm.
[0185] For example, see Fig.12 , Fig.12This is a flow chart of an image correction method provided by an embodiment of the present application. In it, a reference image of a three-dimensional model and a perspective image with blurred texture are input into an image correction model, batch perspective images are obtained through multi-perspective rendering, and the texture image is corrected based on the corrected perspective image, and then a three-dimensional model corresponding to the corrected texture image is obtained, that is, a three-dimensional model with optimized texture is obtained. The whole process takes a short time, and the facial details and texture clarity of the corrected three-dimensional model can be improved.
[0186] For example, see Fig.13 , Fig.13 The present application provides a flowchart of texture map correction according to an embodiment. In which, the 3D model is firstly rendered from multiple perspectives using the texture map of the 3D model to obtain multiple perspective maps, which may include 6 full body perspective maps and 1 face perspective map, and then compared with the reference Figure 1 The image correction model is input in batches to obtain the corresponding 7 corrected view maps, and the texture map is corrected according to the 7 corrected view maps in combination with the texture coordinates of the texture map of the 3D model, thereby obtaining the corrected texture map.
[0187] The embodiment of the present application provides an image correction method. When the method uses an image correction model to correct a viewing angle map, the method first extracts the features of each image through the image correction model, and then predicts the noise features by combining the viewing angle features and reference features corresponding to the viewing angle map with texture defects. After removing the noise features, the corrected viewing angle features are obtained. Based on the corrected viewing angle features, the corrected viewing angle map can be obtained without manual correction, thereby improving the correction efficiency. In addition, when predicting the noise features, the viewing angle features and the reference features are combined. Since the viewing angle features provide the prior information of the first viewing angle map, the predicted noise features are noise features that can make the viewing angle map close to the first viewing angle map but remove the texture defects after being removed. Since the reference features provide the appearance control information of the viewing angle map, the predicted noise features are noise features that can make the viewing angle map match the overall appearance of the three-dimensional model after being removed. That is, through the first viewing angle features and the reference features, the predicted noise features are accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0188] The image correction method provided in the embodiment of the present application can be applied to a variety of scenarios. The following description will be given using a game production scenario as an example and applied to a three-dimensional model production process, which includes the following steps (1)-(7).
[0189] (1) The game production terminal obtains the original texture map of the three-dimensional model.
[0190] The game client includes a virtual scene and virtual objects that can be arranged in the virtual scene. Among them, the virtual scene can be a simulation environment of the real world, a semi-simulation and semi-fictitious environment, or a purely fictitious environment. The virtual scene can be any one of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, or a three-dimensional virtual scene. The embodiment of the present application does not limit the dimension of the virtual scene. For example, the virtual scene may include the sky, land, ocean, etc., and the land may include environmental elements such as deserts and cities. The user can control the virtual character to move in the virtual scene. Optionally, the virtual scene can provide a battle environment for virtual objects, and there are virtual resources available for virtual objects in the virtual scene, such as virtual resources including virtual props required for battle, virtual medicines required for treatment, virtual props required for upgrades, virtual gold coins required for transactions, etc.
[0191] In addition, virtual objects include virtual characters, virtual buildings, virtual props, etc. The virtual characters may be virtual people, virtual animals, virtual elves, cartoon characters, etc. The virtual character may be a virtual image in the virtual scene that is used to represent the user. A plurality of virtual characters may be included in the virtual scene, each of which has its own shape and volume in the virtual scene and occupies a portion of the space in the virtual scene. The three-dimensional model of the virtual character may be a three-dimensional character constructed based on three-dimensional human skeleton technology, and the same virtual character may show different external images by wearing different skins. In some embodiments, the virtual character may also be implemented using a 2.5-dimensional or 2-dimensional model, which is not limited in the embodiments of the present application.
[0192] There may be a variety of different virtual characters in the virtual scene. Exemplarily, the virtual characters may include player characters controlled by operations on the client, and may also include non-player characters (NPCs) set in the virtual scene interaction. Optionally, the virtual character may be a virtual person competing in the virtual scene. Optionally, the number of virtual characters participating in the interaction in the virtual scene may be pre-set, or may be dynamically determined according to the number of clients joining the interaction.
[0193] Therefore, in the process of making a game client, it is necessary to make a three-dimensional model of a virtual object. In the process of making the three-dimensional model, the three-dimensional model is rendered based on the texture map to obtain a rendered three-dimensional model.
[0194] First, the game production terminal obtains the original texture map of the three-dimensional model, which is a texture map with texture defects. The texture map can be uploaded to the game production terminal by the game production personnel, or downloaded from the game server by the game production terminal.
[0195] (2) The game production terminal sends a texture map correction request to the game server, and the texture map correction request carries the original texture map of the three-dimensional model.
[0196] (3) The game server renders the three-dimensional model from multiple perspectives based on the original texture map to obtain multiple perspective images of the three-dimensional model. The game server corrects the texture defects in the multiple perspective images through the image correction model to obtain multiple corrected perspective images. The texture map of the three-dimensional model is corrected based on the corrected multiple perspective images to obtain a corrected texture map.
[0197] (4) The game server sends the corrected texture map to the game production terminal.
[0198] (5) The game production terminal receives the modified texture map, renders the three-dimensional model based on the modified texture map, obtains the rendered three-dimensional model, and displays the rendered three-dimensional model.
[0199] (6) In response to the confirmation operation on the rendered three-dimensional model, the game production terminal uploads the rendered three-dimensional model to the game server, or uploads the corrected texture map.
[0200] The game production terminal displays the rendered 3D model, and the game production personnel can preview the rendering effect of the 3D model to determine whether the rendering effect of the 3D model meets the requirements. If the rendering effect of the 3D model is satisfactory, the confirmation operation can be triggered, indicating that the 3D model rendering is completed.
[0201] In another embodiment, if the game production staff is not satisfied with the rendering effect of the 3D model, a rejection operation can be triggered to correct the texture map for the 3D model and re-render it.
[0202] (7) The game server adds the rendered three-dimensional model or the modified texture map as a game resource of the three-dimensional model to the game resource package of the game client.
[0203] The game server will create a game resource package for the game client, which includes various game resources required for the game client to install or run. The game server will then add the rendered 3D model to the game resource package, or add the modified texture map to the game resource package. After the game server releases the game client, the terminal that downloads the game client can install and run the game client based on the game resource package, thereby displaying the rendered 3D model of the virtual object in the game client.
[0204] Of course, the method provided in the embodiment of the present application can also be applied to other scenarios of correcting texture images, and the embodiment of the present application does not limit this.
[0205] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0206] Please refer to Fig.14 , Fig.14 : is a block diagram of a training device for an image correction model provided by an embodiment of the present application. The device has the function of implementing the training method for the above-mentioned image correction model, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Fig.14 As shown, the device may include:
[0207] The acquisition module 1401 is used to acquire multiple groups of first training samples, each group of first training samples includes a first perspective image, a second perspective image and a reference image of the three-dimensional model, the texture of the three-dimensional model in the first perspective image has defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model;
[0208] The input-output module 1402 is used to input the first perspective image, the second perspective image and the reference image into the image correction model, process the first perspective image, the second perspective image and the reference image respectively through the image correction model to obtain the first perspective feature, the second perspective feature and the reference feature, add the first noise feature to the second perspective feature to obtain the third perspective feature, perform prediction based on the first perspective feature, the third perspective feature and the reference feature to obtain the second noise feature used to represent the noise feature in the third perspective feature, and the first perspective feature, the second perspective feature and the reference feature are used to describe the first perspective image, the second perspective image and the reference image respectively;
[0209] The training module 1403 is used to iteratively train the image correction model based on the difference information between the first noise characteristics and the second noise characteristics of each of the multiple groups of first training samples. The image correction model is used to predict the noise in the perspective image and correct the perspective image by removing the noise.
[0210] In some embodiments, the input / output module 1402 is used to:
[0211] Splicing the first-view feature and the third-view feature to obtain a spliced feature;
[0212] Prediction is performed based on the concatenated features and the reference features to obtain a second noise feature.
[0213] In some embodiments, the input / output module 1402 is used to:
[0214] The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features. The convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range.
[0215] The processed spliced features and reference features are processed through the attention module of the image correction model to obtain cross-attention features. The query parameter in the attention module is the processed spliced features, and the key parameter and value parameter in the attention module are both reference features.
[0216] Through the denoising module of the image correction model, the splicing features and the cross-attention features after fusion processing are fused, and prediction is performed based on the fused features to obtain the second noise feature.
[0217] In some embodiments, the acquisition module 1401 is further used to:
[0218] Based on the pre-trained model, an image correction model is obtained. The pre-trained model is trained based on multiple groups of second training samples. The model parameters of the convolution module corresponding to the second perspective features, the model parameters of the attention module, and the model parameters of the denoising module in the image correction model are respectively the model parameters of the convolution module, the attention module, and the denoising module in the pre-trained model. The second training samples include sample images and sample images with added noise.
[0219] In some embodiments, the input / output module 1402 is used to:
[0220] The first perspective image and the second perspective image are processed respectively by the first encoding module of the image correction model to obtain a first perspective feature and a second perspective feature, wherein the first perspective feature is a representation of the first perspective image in a latent space, and the second perspective feature is a representation of the second perspective image in a latent space;
[0221] The reference image is processed by the second encoding module of the image correction model to obtain a reference feature, which is a representation of the reference image in the embedding space.
[0222] The embodiment of the present application provides a training device for an image correction model. During the training process, the features of each image are first extracted through the image correction model, and then noise features are added to the perspective features corresponding to the second perspective image after the texture defects are corrected. Then, the perspective features corresponding to the first perspective image with texture defects and the reference features are combined to predict the added noise features, and the image correction model is trained based on the difference information between the predicted noise features and the actually added noise features, so that the image correction model has the ability to accurately predict the noise features. After the image correction model is trained in this way, for any perspective image with texture defects, the image correction model can predict the noise corresponding to the texture defects in the perspective image, and correct the perspective image by removing the noise without manual correction, thereby improving the correction efficiency. Moreover, when predicting the noise feature, the first viewing angle feature and the reference feature are combined. Since the first viewing angle feature provides prior information of the first viewing angle image, the predicted noise feature is a noise feature that can make the viewing angle image close to the first viewing angle image but correct texture defects after being removed. Since the reference feature provides appearance control information of the viewing angle image, the predicted noise feature is a noise feature that can make the viewing angle image match the overall appearance of the three-dimensional model after being removed. That is, the device makes the noise feature predicted by the image correction model accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0223] Please refer to Fig.15 , Fig.15 : is a block diagram of an image correction device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned image correction method, and the function can be implemented by hardware, or by hardware executing corresponding software. The device can be the computer device introduced above, or it can be set in a computer device. Fig.15 As shown, the device may include:
[0224] An acquisition module 1501 is used to acquire a noise map, a view map of a 3D model, and a reference map, wherein the texture of the 3D model in the view map has defects, and the reference map is a view map that provides appearance features of the 3D model;
[0225] The input-output module 1502 is used to input the noise map, the viewing angle map and the reference map into the image correction model, process the noise map, the viewing angle map and the reference map respectively through the image correction model to obtain noise features, viewing angle features and reference features, perform prediction based on the noise features, viewing angle features and reference features to obtain predicted noise features for representing noise features other than the viewing angle features in the noise features, remove the predicted noise features from the noise features to obtain corrected viewing angle features, and obtain a corrected viewing angle map based on the corrected viewing angle features;
[0226] Among them, the noise feature, the viewing angle feature and the reference feature are used to describe the noise image, the viewing angle image and the reference image respectively.
[0227] In some embodiments, the input / output module 1502 is used to:
[0228] Splicing the noise feature and the view feature to obtain a splicing feature;
[0229] Prediction is performed based on the concatenated features and the reference features to obtain the predicted noise features.
[0230] In some embodiments, the input / output module 1502 is used to:
[0231] The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features. The convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range.
[0232] The processed spliced features and reference features are processed through the attention module of the image correction model to obtain cross-attention features. The query parameter in the attention module is the processed spliced features, and the key parameter and value parameter in the attention module are both reference features.
[0233] Through the denoising module of the image correction model, the splicing features and cross-attention features after fusion processing are fused, and prediction is performed based on the fused features to obtain the predicted noise features.
[0234] In some embodiments, the input / output module 1502 is used to:
[0235] The noise map and the view map are processed respectively by the first encoding module of the image correction model to obtain noise features and view features, wherein the noise feature is the representation of the noise map in the latent space, and the view feature is the representation of the view map in the latent space;
[0236] The reference image is processed by the second encoding module of the image correction model to obtain a reference feature, which is a representation of the reference image in the embedding space.
[0237] In some embodiments, the input / output module 1502 is used to:
[0238] The corrected viewing angle features are processed by the decoding module of the image correction model to obtain a corrected viewing angle map.
[0239] In some embodiments, the input / output module 1502 is used to:
[0240] Based on the noise feature, the viewing angle feature and the reference feature, multiple first iteration calculations are performed until the first iteration stop condition is reached to obtain a modified viewing angle feature; the multiple first iteration calculations include:
[0241] In the first iteration calculation, prediction is performed based on the noise feature, the viewing angle feature and the reference feature to obtain the first predicted noise feature, the first predicted noise feature is used to represent the noise feature other than the viewing angle feature in the noise feature, and the first predicted noise feature is removed from the noise feature to obtain the first corrected viewing angle feature;
[0242] In the Mth first iteration calculation, prediction is performed based on the M-1th corrected viewing angle feature, the viewing angle feature and the reference feature to obtain the Mth predicted noise feature. The Mth predicted noise feature is used to represent the noise features other than the viewing angle features in the M-1th corrected viewing angle feature. The Mth predicted noise feature is removed from the M-1th corrected viewing angle feature to obtain the Mth corrected viewing angle feature, where M is an integer greater than 1.
[0243] In some embodiments, the apparatus further comprises:
[0244] A first correction module is used to correct multiple perspective images of the three-dimensional model based on the image correction model to obtain multiple corrected perspective images, where the multiple perspective images are rendered based on a texture image of the three-dimensional model, where the texture image includes textures at multiple positions on the three-dimensional model;
[0245] The second correction module is used to correct the texture map based on the multiple perspective maps and the multiple corrected perspective maps.
[0246] In some embodiments, the second correction module is used to:
[0247] Based on the multiple viewing angle graphs and the multiple corrected viewing angle graphs, multiple second iteration calculations are performed until the second iteration stop condition is reached; the multiple second iteration calculations include:
[0248] In the first second iteration calculation, the difference information between each viewing angle image and the corresponding corrected viewing angle image is determined respectively, and the texture image is corrected based on the difference information between the multiple viewing angle images and the multiple corrected viewing angle images;
[0249] In the Nth second iteration calculation, the three-dimensional model is rendered based on the texture map corrected in the N-1th second iteration calculation and the multiple perspectives corresponding to the multiple perspective maps to obtain the Nth multiple perspective maps, and the difference information between each perspective map of the Nth and the corresponding corrected perspective map is determined respectively; based on the difference information between the multiple perspective maps of the Nth and the multiple corrected perspective maps, the texture map corrected in the N-1th second iteration calculation is corrected, where N is an integer greater than 1.
[0250] The embodiment of the present application provides an image correction device. When the device uses an image correction model to correct a viewing angle map, the image correction model is first used to extract the features of each image. Then, the noise features are predicted by combining the viewing angle features and reference features corresponding to the viewing angle map with texture defects. After removing the predicted noise features, the corrected viewing angle features are obtained. Based on the corrected viewing angle features, the corrected viewing angle map can be obtained without manual correction, thereby improving the correction efficiency. In addition, when predicting the noise features, the viewing angle features and the reference features are combined. Since the viewing angle features provide prior information of the viewing angle map, the predicted noise features are noise features that can make the viewing angle map close to the original viewing angle map but remove texture defects after being removed. Since the reference features provide appearance control information of the viewing angle map, the predicted noise features are noise features that can make the viewing angle map match the overall appearance of the three-dimensional model after being removed. That is, through the viewing angle features and the reference features, the predicted noise features are accurate, which can improve the ability of the image correction model to correct texture defects, thereby improving the correction quality.
[0251] Please refer to Fig.16 , which shows a block diagram of a computer device 1600 provided in one embodiment of the present application. The computer device 1600 may be a terminal device 10 or a server 20 in an implementation environment, and the computer device 1600 is used to implement the training method of the image correction model or the image correction method provided in the above embodiment. Specifically:
[0252] Typically, the computer device 1600 includes a processor 1610 and a memory 1620 .
[0253] The processor 1610 may include one or more processing cores, such as a 4-core processor, a 16-core processor, etc. The processor 1610 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 1610 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit; the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1610 may include a GPU, which is responsible for executing the method steps provided in this application.
[0254] The memory 1620 may include one or more computer-readable storage media, which may be non-transitory. The memory 1620 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1620 is used to store a computer program, which is configured to be executed by one or more processors (such as a GPU) to implement the above-mentioned image correction model training method or image correction method.
[0255] Those skilled in the art will understand that Fig.16 The structure shown in the figure does not constitute a limitation on the computer device 1600, and the computer device 1600 may include more or less components than shown in the figure, or combine some components, or adopt a different arrangement of components.
[0256] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the image correction model training method or image correction method of any of the above-mentioned implementation methods.
[0257] An embodiment of the present application also provides a computer program product, which includes a computer program, the computer program is stored in a computer-readable storage medium, a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the image correction model training method or image correction method of any of the above-mentioned implementation methods.
[0258] In some embodiments, the computer program product involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected by a communication network.
[0259] It should be noted that the relevant data collection and processing in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained, and subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0260] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application are not limited to this.
[0261] All the above optional technical solutions can be combined in any way to form optional embodiments of the present application, which will not be described one by one here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an image correction model, characterized in that: The method comprises: Acquire multiple groups of first training samples, each group of first training samples includes a first perspective image, a second perspective image, and a reference image of a three-dimensional model, the texture of the three-dimensional model in the first perspective image has defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model; Inputting the first viewing angle image, the second viewing angle image and the reference image into an image correction model, respectively processing the first viewing angle image, the second viewing angle image and the reference image through the image correction model to obtain a first viewing angle feature, a second viewing angle feature and a reference feature, adding a first noise feature to the second viewing angle feature to obtain a third viewing angle feature, performing prediction based on the first viewing angle feature, the third viewing angle feature and the reference feature to obtain a second noise feature used to represent a noise feature in the third viewing angle feature, and the first viewing angle feature, the second viewing angle feature and the reference feature are respectively used to describe the first viewing angle image, the second viewing angle image and the reference image; The image correction model is iteratively trained based on the difference information between the first noise feature and the second noise feature of each of the multiple groups of first training samples. The image correction model is used to predict the noise in the viewing angle image and correct the viewing angle image by removing the noise.
2. The method according to claim 1, characterized in that The predicting based on the first viewing angle feature, the third viewing angle feature and the reference feature to obtain a second noise feature for representing a noise feature in the third viewing angle feature includes: splicing the first viewing angle feature and the third viewing angle feature to obtain a splicing feature; Prediction is performed based on the splicing feature and the reference feature to obtain the second noise feature.
3. The method according to claim 2, characterized in that The performing prediction based on the splicing feature and the reference feature to obtain the second noise feature includes: The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features, wherein the convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range; The processed splicing features and the reference features are processed by the attention module of the image correction model to obtain cross-attention features, wherein the query parameters in the attention module are the processed splicing features, and the key parameters and value parameters in the attention module are both the reference features; The processed splicing features and the cross-attention features are fused through the denoising module of the image correction model, and prediction is performed based on the fused features to obtain the second noise feature.
4. The method according to claim 3, characterized in that Before obtaining the plurality of groups of first training samples, the method further includes: Based on a pre-trained model, the image correction model is obtained, the pre-trained model is trained based on multiple groups of second training samples, the model parameters of the convolution module corresponding to the second perspective feature, the model parameters of the attention module and the model parameters of the denoising module in the image correction model are respectively the model parameters of the convolution module, the attention module and the denoising module in the pre-trained model, and the second training sample includes a sample image and the sample image with noise added.
5. The method according to claim 1, characterized in that: The first viewing angle image, the second viewing angle image and the reference image are processed respectively by the image correction model to obtain a first viewing angle feature, a second viewing angle feature and a reference feature, including: The first viewing angle image and the second viewing angle image are processed respectively by the first encoding module of the image correction model to obtain the first viewing angle feature and the second viewing angle feature, wherein the first viewing angle feature is a representation of the first viewing angle image in a latent space, and the second viewing angle feature is a representation of the second viewing angle image in the latent space; The reference image is processed by the second encoding module of the image correction model to obtain the reference feature, where the reference feature is a representation of the reference image in the embedding space.
6. An image correction method, characterized in that: The method comprises: Acquire a noise map, a viewing angle map of a three-dimensional model, and a reference map, wherein the texture of the three-dimensional model in the viewing angle map has defects, and the reference map is a viewing angle map that provides appearance features of the three-dimensional model; Inputting the noise map, the viewing angle map and the reference map into an image correction model, respectively processing the noise map, the viewing angle map and the reference map through the image correction model to obtain noise features, viewing angle features and reference features, performing prediction based on the noise features, the viewing angle features and the reference features to obtain predicted noise features for representing noise features other than viewing angle features in the noise features, removing the predicted noise features from the noise features to obtain corrected viewing angle features, and obtaining a corrected viewing angle map based on the corrected viewing angle features; The noise feature, the viewing angle feature and the reference feature are used to describe the noise image, the viewing angle image and the reference image respectively.
7. The method according to claim 6, characterized in that The performing prediction based on the noise feature, the viewing angle feature and the reference feature to obtain a predicted noise feature for representing a noise feature other than the viewing angle feature in the noise feature includes: Splicing the noise feature and the viewing angle feature to obtain a splicing feature; Prediction is performed based on the splicing feature and the reference feature to obtain the predicted noise feature.
8. The method according to claim 7, characterized in that The performing prediction based on the splicing feature and the reference feature to obtain the predicted noise feature includes: The convolution module of the image correction model is used to perform convolution processing on the splicing features to obtain processed splicing features, wherein the convolution processing is used to extract features from the splicing features and limit the dimensions of the splicing features to a preset range; The processed splicing features and the reference features are processed by the attention module of the image correction model to obtain cross-attention features, wherein the query parameters in the attention module are the processed splicing features, and the key parameters and value parameters in the attention module are both the reference features; The processed splicing features and the cross-attention features are fused through the denoising module of the image correction model, and prediction is performed based on the fused features to obtain the predicted noise features.
9. The method according to claim 6, characterized in that The step of processing the noise map, the viewing angle map and the reference map respectively by the image correction model to obtain noise features, viewing angle features and reference features includes: The noise map and the viewing angle map are processed respectively by the first encoding module of the image correction model to obtain the noise feature and the viewing angle feature, wherein the noise feature is a representation of the noise map in a latent space, and the viewing angle feature is a representation of the viewing angle map in the latent space; The reference image is processed by the second encoding module of the image correction model to obtain the reference feature, where the reference feature is a representation of the reference image in the embedding space.
10. The method according to claim 9, characterized in that The step of obtaining a corrected viewing angle map based on the corrected viewing angle feature comprises: The corrected viewing angle feature is processed by the decoding module of the image correction model to obtain the corrected viewing angle map.
11. The method according to claim 6, characterized in that The predicting based on the noise feature, the viewing angle feature and the reference feature to obtain a predicted noise feature for representing a noise feature other than the viewing angle feature in the noise feature, and removing the predicted noise feature from the noise feature to obtain a modified viewing angle feature, comprises: Based on the noise feature, the viewing angle feature and the reference feature, multiple first iterative calculations are performed until a first iteration stop condition is reached to obtain the modified viewing angle feature; the multiple first iterative calculations include: In a first iterative calculation, a prediction is performed based on the noise feature, the viewing angle feature and the reference feature to obtain a first predicted noise feature, wherein the first predicted noise feature is used to represent noise features other than the viewing angle feature in the noise feature, and the first predicted noise feature is removed from the noise feature to obtain a first modified viewing angle feature; In the Mth first iterative calculation, prediction is performed based on the M-1th corrected viewing angle feature, the viewing angle feature and the reference feature to obtain the Mth predicted noise feature, wherein the Mth predicted noise feature is used to represent the noise feature other than the viewing angle feature in the M-1th corrected viewing angle feature, and the Mth predicted noise feature is removed from the M-1th corrected viewing angle feature to obtain the Mth corrected viewing angle feature, where M is an integer greater than 1.
12. The method according to claim 6, characterized in that The method further comprises: Correcting multiple perspective images of the three-dimensional model by using the image correction model to obtain multiple corrected perspective images, wherein the multiple perspective images are rendered based on a texture image of the three-dimensional model, wherein the texture image includes textures at multiple positions on the three-dimensional model; The texture map is modified based on the multiple perspective maps and the multiple modified perspective maps.
13. The method according to claim 12, characterized in that The step of correcting the texture map based on the multiple perspective maps and the multiple corrected perspective maps comprises: Based on the multiple viewing angle images and the multiple corrected viewing angle images, multiple second iterative calculations are performed until a second iteration stop condition is reached; the multiple second iterative calculations include: In the first second iteration calculation, respectively determining the difference information between each viewing angle image and the corresponding corrected viewing angle image, and correcting the texture image based on the difference information between the multiple viewing angle images and the multiple corrected viewing angle images; In the Nth second iteration calculation, the three-dimensional model is rendered based on the texture map corrected in the N-1th second iteration calculation and the multiple perspectives corresponding to the multiple perspective maps to obtain the Nth multiple perspective maps, and the difference information between each perspective map of the Nth and the corresponding corrected perspective map is determined respectively; based on the difference information between the Nth multiple perspective maps and the multiple corrected perspective maps, the texture map corrected in the N-1th second iteration calculation is corrected, where N is an integer greater than 1.
14. A training device for an image correction model, characterized in that: The device comprises: an acquisition module, configured to acquire multiple groups of first training samples, each group of first training samples comprising a first perspective image, a second perspective image, and a reference image of a three-dimensional model, wherein the texture of the three-dimensional model in the first perspective image has defects, the second perspective image is the first perspective image after the defects are corrected, and the reference image is a perspective image showing the appearance of the three-dimensional model; an input-output module, configured to input the first viewing angle image, the second viewing angle image, and the reference image into an image correction model, process the first viewing angle image, the second viewing angle image, and the reference image respectively through the image correction model to obtain a first viewing angle feature, a second viewing angle feature, and a reference feature, add a first noise feature to the second viewing angle feature to obtain a third viewing angle feature, perform prediction based on the first viewing angle feature, the third viewing angle feature, and the reference feature to obtain a second noise feature used to represent a noise feature in the third viewing angle feature, and the first viewing angle feature, the second viewing angle feature, and the reference feature are used to describe the first viewing angle image, the second viewing angle image, and the reference image respectively; A training module is used to iteratively train the image correction model based on the difference information between the first noise characteristics and the second noise characteristics of each of the multiple groups of first training samples. The image correction model is used to predict the noise in the perspective image and correct the perspective image by removing the noise.
15. An image correction device, characterized in that: The device comprises: An acquisition module, used to acquire a noise map, a viewing angle map of a three-dimensional model, and a reference map, wherein the texture of the three-dimensional model in the viewing angle map has defects, and the reference map is a viewing angle map that provides appearance features of the three-dimensional model; An input-output module, used for inputting the noise map, the viewing angle map and the reference map into an image correction model, processing the noise map, the viewing angle map and the reference map respectively through the image correction model to obtain noise features, viewing angle features and reference features, making predictions based on the noise features, the viewing angle features and the reference features to obtain predicted noise features for representing noise features other than viewing angle features in the noise features, removing the predicted noise features from the noise features to obtain corrected viewing angle features, and obtaining a corrected viewing angle map based on the corrected viewing angle features; The noise feature, the viewing angle feature and the reference feature are used to describe the noise image, the viewing angle image and the reference image respectively.
16. A computer device, characterized in that: The computer device includes a processor and a memory, the memory is used to store a computer program, and the computer program is loaded by the processor and executes the training method of the image correction model described in any one of claims 1 to 5 or the image correction method described in any one of claims 6 to 13.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the training method of the image correction model described in any one of claims 1 to 5 or the image correction method described in any one of claims 6 to 13.
18. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the training method of the image correction model described in any one of claims 1 to 5 or the image correction method described in any one of claims 6 to 13.
Citation Information
Patent Citations
Text-based three-dimensional modeling method, image rendering method and device
CN116958423A
Model training method, video restoration method, device, equipment, medium and product
CN117670735A
Building texture image restoration method based on diffusion model
CN117788344A
Three-dimensional model generation method and device, electronic equipment and storage medium
CN118918257A