Method, apparatus, device, and medium for inverse rendering of an image

The method enhances inverse rendering by predicting geometric and material features and illuminance values, addressing limitations in representing complex materials and lighting, thereby improving the integration of virtual objects into real scenes.

JP2025519258AActive Publication Date: 2025-06-24REALSEE (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024572215
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-17
Filing Date
2023-02-07
Publication Date
2025-06-24
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

Existing inverse rendering methods in computer graphics and computer vision lack the ability to accurately represent complex materials and lighting environments, leading to limitations in the fusion of virtual objects with real scenes.

Method used

A method involving a feature prediction model to extract geometric and material features, followed by an illumination prediction model to determine illuminance values, enabling more detailed modeling of complex materials and lighting conditions.

Benefits of technology

Improves the physical accuracy of material and geometric representations, allowing for enhanced integration of virtual objects into real scenes by overcoming simplified material representations and enabling more precise lighting simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519258000001_ABST
    Figure 2025519258000001_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, an apparatus, an electronic device, and a storage medium for inverse rendering of an image. The method includes: inputting a processing target image into a feature prediction model, predicting geometric features and material features of the processing target image by the feature prediction model, and obtaining a geometric feature map and a material feature map of the processing target image, where the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map; inputting the processing target image, the geometric feature map, and the material feature map into an illumination prediction model, predicting an illuminance value of the processing target image for each pixel, and obtaining an illumination feature map of the processing target image; and performing a preset process on the processing target image based on the geometric feature map, the material feature map, and the illumination feature map. The limitation of a simplified material representation for appearance acquisition in the inverse rendering process is overcome, which helps to improve the physical accuracy of the material, geometric shape, and illumination predicted by the inverse rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This disclosure claims priority to a Chinese patent application filed with the China National Intellectual Property Administration on June 17, 2022, with application number CN202210689653.X and invention title "Method, apparatus, device, and medium for inverse rendering of images", the entire content of which is incorporated herein by reference.

[0002] This disclosure relates to the field of computer vision, and particularly to a method, apparatus, device, and medium for inverse rendering of images.

Background Art

[0003] Inverse rendering of images is an important application in the fields of computer graphics and computer vision, aiming to restore attributes such as the geometric shape, material, and lighting of an image from the image. In the fields of augmented reality and scene digitization, an image can be processed according to attributes such as the geometric shape, material, and lighting obtained by inverse rendering. For example, virtual objects can be generated within the image. Attributes such as the geometric shape, material, and lighting of the image obtained by inverse rendering are directly related to the fusion effect of virtual objects and the scene.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, device, and medium for inverse rendering of images, which are used to improve the effect of image processing that depends on the material representation obtained by inverse rendering.

Means for Solving the Problems

[0005] A method for inverse rendering of images according to one aspect of the embodiments of this disclosure is A step of inputting a processing target image into a feature prediction model, predicting geometric features and material features of the processing target image by the feature prediction model, and obtaining a geometric feature map and a material feature map of the processing target image, wherein the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map. A step of inputting the processing target image, the geometric feature map, and the material feature map into an illumination prediction model, predicting an illuminance value of the processing target image for each pixel by the illumination prediction model, and obtaining an illumination feature map of the processing target image. A step of performing a preset process on the processing target image based on the geometric feature map, the material feature map, and the illumination feature map.

[0006] An apparatus for inverse rendering of an image according to another aspect of an embodiment of the present disclosure includes A feature prediction unit configured to input a processing target image into a feature prediction model, predict geometric features and material features of the processing target image by the feature prediction model, and obtain a geometric feature map and a material feature map of the processing target image, wherein the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map. An illumination prediction unit configured to input the processing target image, the geometric feature map, and the material feature map into an illumination prediction model, predict an illuminance value of the processing target image for each pixel by the illumination prediction model, and obtain an illumination feature map of the processing target image. An image processing unit configured to perform a preset process on the processing target image based on the geometric feature map, the material feature map, and the illumination feature map.

[0007] An electronic device according to still another aspect of an embodiment of the present disclosure includes A memory for storing a computer program product. Execute the computer program product stored in the memory, and when the computer program product is executed, include a processor for realizing a method for inverse rendering of an image provided by any of the above embodiments of the present disclosure.

[0008] A computer-readable storage medium according to still another further aspect of the embodiments of the present disclosure stores program code, and the program code can be called by a processor to realize a method for inverse rendering of an image provided by any of the above embodiments of the present disclosure.

Advantages of the Invention

[0009] In the solution means provided by the embodiments of the present disclosure, a feature prediction model is used to predict the geometric features and material features of the image to be processed. The geometric features include normal features and depth features, and the material features include albedo, roughness, and metallicity. Then, an illumination prediction model is used to predict the illuminance value of the image to be processed, and the preset processing for the image can be performed according to the predicted geometric features, material features, and illuminance value. The depth features, albedo, roughness, and metallicity can more physically and accurately characterize the complex materials in the image to be processed. As a result, in subsequent processing processes, complex lighting environments such as specular reflections can be modeled in more detail, the limitations of simplified material representations for appearance acquisition in the inverse rendering process can be overcome, and it is useful for improving the physical accuracy of the materials, geometric shapes, and illuminations predicted depending on inverse rendering and for improving the effects of image processing depending on the material representations obtained by inverse rendering. For example, in the fields of augmented reality and scene digitization, the fusion effect between virtual objects and scenes can be improved in this way.

[0010] The technical solution of the present disclosure will be further described in detail below through drawings and examples.

Brief Description of the Drawings

[0011] The drawings that form a part of the specification illustrate embodiments of the present disclosure and are used to interpret the principles of the present disclosure together with the description.

[0012] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0013] Hereinafter, various exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. It should be noted that the relative arrangements, mathematical formulas, and numerical values of the members and steps described in these embodiments do not limit the scope of the present disclosure unless otherwise specified.

[0014] In the embodiments of the present disclosure, it should also be understood that "a plurality" can mean two or more, and "at least one" can mean one, two or more.

[0015] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only for distinguishing different steps, devices or modules, etc., and do not represent any specific technical meaning, nor do they represent an inevitable logical order between them.

[0016] It should also be understood that any member, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more unless there is an explicit limitation or contrary disclosure in the context.

[0017] In the description of each embodiment in the present disclosure, the differences between the embodiments are emphasized, but the same or similar points may be referred to each other, and for the sake of brevity, it should also be understood that they will not be described one by one.

[0018] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way intended to limit the present disclosure and its applications.

[0019] Although it may not be possible to discuss in detail the technologies, methods and devices known to those skilled in the relevant art, where appropriate, the said technologies, methods and devices should be regarded as part of the specification.

[0020] In addition, since similar symbols and characters represent similar items in the following drawings, once an item is defined in one drawing, there is no need for further discussion in subsequent drawings.

[0021] In addition, the term "and / or" in the present disclosure is only for describing the relationship of related objects, indicating that three types of relationships may exist. A and / or B can indicate three situations: A exists alone, A and B exist simultaneously, and B exists alone. Also, in the present disclosure, the character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0022] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers that can operate in many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use in electronic devices such as terminal devices, computer systems, and servers include personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable household appliances, network personal computers, mini-computer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, but are not limited thereto.

[0023] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules may be located in the storage media of a local computing system or a remote computing system including a storage device.

[0024] To make the technical solutions and advantages in the embodiments of the present disclosure clearer, exemplary embodiments of the present disclosure will be further described below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure and do not cover all embodiments. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0025] An exemplary description of a method for inverse rendering of an image of the present disclosure will be given below with reference to FIG. 1. FIG. 1 shows a flowchart of one embodiment of a method for inverse rendering of an image of the present disclosure. As shown in FIG. 1, this process includes the following steps: Step 110: Input the image to be processed into a feature prediction model, predict the geometric features and material features of the image to be processed by the feature prediction model, and obtain a geometric feature map and a material feature map of the image to be processed.

[0026] Here, the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map.

[0027] In this embodiment, the geometric features can characterize the geometric attributes of the image to be processed. For example, they can include normal features and depth features. The normal features can characterize the normal vector of a pixel point, and the depth features can characterize the depth of a pixel point. The material features can characterize the material attributes of the pixel points in the image to be processed. For example, they can include albedo (base color), roughness, and metallicity. The albedo can represent the ratio of the light flux scattered in at least one direction by all illuminated parts of the object's surface to the light flux incident on the object's surface. The roughness can represent the smoothness of the object's surface and is used to describe the behavior when light irradiates the object's surface. For example, the smaller the roughness of the object's surface, the closer the specular reflection is when light irradiates the object's surface. The metallicity is used to characterize the degree of metal of the object. The higher the metallicity, the closer the object is to metal and vice versa, closer to non-metal.

[0028] The feature prediction model can characterize the correspondence between the image to be processed and its geometric features and material features, predict the geometric features and material features of each pixel point in the image to be processed, and is used to form the corresponding feature map based on the predicted feature values. Accordingly, the normal map, depth map, albedo feature map, roughness feature map, and metallicity feature map can respectively represent the normal vector, depth, albedo, roughness, and metallicity of at least one pixel point in the image to be processed.

[0029] In one specific example, the feature prediction model may be any neural network model such as a convolutional neural network, a residual network, for example, a multi-branch encoder / decoder based on ResNet and Unet. The encoder may be ResNet-18, and the decoder may be composed of five convolutional layers with skip connections. After training the feature prediction model using sample data, the feature prediction model is used to perform processes such as feature extraction, downsampling, high-dimensional feature extraction, upsampling, decoding, layer hopping connection, and shallow feature fusion on the image to be processed. Finally, the normal feature, depth feature, albedo, roughness, and metallicity of each pixel point in the image to be processed are predicted, and a normal feature map, a depth feature map, an albedo feature map, a roughness feature map, and a metallicity feature map are respectively formed based on the predicted feature values, thereby obtaining the geometric features and material features of the image to be processed.

[0030] In one selectable example, this step 110 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a feature prediction unit operated by the processor.

[0031] In step 120, the image to be processed, the geometric feature map, and the material feature map are input into the illumination prediction model. The illumination prediction model predicts the illuminance value of the image to be processed for each pixel, and obtains the illumination feature map of the image to be processed.

[0032] In this embodiment, the illuminance value can characterize the lighting environment of a point in space. The illumination prediction model can characterize the correspondence between the image to be processed and its geometric features, material features, and the lighting environment.

[0033] In one specific example, for the lighting prediction model, any neural network model such as a convolutional neural network or a residual network can be used, for example, a multi-branch encoder / decoder based on ResNet and Unet. The execution entity (which may be, for example, a terminal device or a server) can, through preprocessing, superimpose in terms of the number of channels a target image to be processed, a geometric feature map (including a normal feature map and a depth feature map), and a material feature map (including an albedo feature map, a roughness feature map, and a metal amount feature map), input the superimposed image into the lighting prediction model, and predict the spatial lighting environment of each pixel point, that is, the illuminance value of each pixel point, through operations such as feature extraction, encoding, and decoding, and form a spatially continuous HDR lighting feature map based on the predicted illuminance values.

[0034] In one selectable example, this step 120 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a lighting prediction unit operated by the processor.

[0035] Step 130: Perform pre-set processing on the target image based on the geometric feature map, the material feature map, and the lighting feature map.

[0036] In this embodiment, through steps 110 and 120, inverse rendering of the image to be processed can be realized, and geometric features and material features of the image to be processed can be obtained. The pre-set processing represents subsequent processing performed on the image to be processed based on the geometric features and material features obtained by inverse rendering. For example, in the field of mixed reality, a real image collected by a camera is used as the image to be processed, and by inserting a virtual image into the real image, fusion of the real world and the virtual image can be realized. Also, for example, based on the geometric features and material features of the image to be processed, a virtual object can be generated in the image to be processed by dynamic virtual object synthesis. Also, for example, based on the geometric features and material features of the image to be processed, the material of an object in the image to be processed can be edited to present an object of a different material.

[0037] Hereinafter, a method for inverse rendering of an image in this embodiment will be exemplarily described with reference to the scene shown in FIG. 2. As shown in FIG. 2, the image 210 to be processed is an LDR panoramic image, and a geometric feature map 230 and a material feature map 240 of the image 210 to be processed can be predicted using a feature prediction model 220. The geometric feature map includes a normal feature map 231 and a depth feature map 232, and the material feature map includes an albedo feature map 241, a roughness feature map 242, and a metallicity feature map 243. Thereafter, the image 210 to be processed, the geometric feature map 230, and the material feature map 240 are input into a second prediction model 250 to obtain an illumination feature map 260. Thereafter, based on the geometric feature map 230 and the material feature map 240, virtual objects 271, 272, and 273 are generated in the image 210 to be processed to obtain a processed image 270.

[0038] In one selectable example, this step 130 may be executed by the processor calling corresponding instructions stored in the memory, or may be executed by an image processing unit operated by the processor.

[0039] In the method for inverse rendering of an image provided by this embodiment, a feature prediction model is used to predict the geometric features and material features of the image to be processed. The geometric features include normal features and depth features, and the material features include albedo, roughness, and metallicity. Then, an illumination prediction model is used to predict the illuminance value of the image to be processed, and a preset process can be performed on the image according to the predicted geometric features, material features, and illuminance value. The depth features, albedo, roughness, and metallicity can more physically and accurately characterize the complex materials in the image to be processed. As a result, in subsequent processing processes, complex illumination environments such as specular reflection can be modeled in more detail, the limitation of simplified material representation for appearance acquisition in the inverse rendering process can be overcome, and it is helpful for improving the physical accuracy of the predicted materials, geometric shapes, and illuminations depending on inverse rendering and for improving the effect of image processing depending on the material representation obtained by inverse rendering.

[0040] In some selectable embodiments of this embodiment, step 120 described above may further include: using an illumination prediction model to process the image to be processed, the geometric feature map, and the material feature map, predicting the illuminance value of a pixel point in the image to be processed, and generating a panoramic image corresponding to the pixel point based on the predicted illuminance value; and stitching the panoramic images corresponding to the pixel points in the image to be processed to obtain an illumination feature map.

[0041] In this embodiment, the illumination prediction model can predict the illumination environment of each pixel point in space by processing the image to be processed, the geometric feature map, and the material feature map. Since a point in space can receive light radiated from any angle in the space, the illumination environment of the point can be characterized using a 360° panoramic image. Then, according to the position of the pixel point in the image to be processed, the panoramic image corresponding to at least one pixel point is stitched into the illumination feature map.

[0042] In this embodiment, by predicting the illuminance value of a pixel point in the image to be processed using an illumination prediction model and characterizing the illuminance value of the pixel point using a panoramic image, the illumination characteristics of the image to be processed can be characterized more accurately.

[0043] Next, referring to FIG. 3, FIG. 3 shows a schematic diagram of a process for training a feature prediction model and an illumination prediction model in one embodiment of a method for inverse rendering of an image of the present disclosure. As shown in FIG. 3, this process includes the following steps: Step 310: Input a sample image into a pre-trained feature prediction model, predict the geometric features and material features of the sample image, and obtain a sample geometric feature map and a sample material feature map of the sample image.

[0044] In this example, the pre-trained feature prediction model represents a feature prediction model that can complete a prediction operation on an input image through training.

[0045] As an example, the pre-training of the feature prediction model can be realized using a virtual dataset. The virtual dataset can include virtual images obtained by forward rendering processing, virtual geometric feature maps and virtual material feature maps generated during the forward rendering process. Then, using the virtual image as the input of the initial feature prediction model and the virtual geometric feature map and the virtual material feature map as the desired output, and training the initial feature prediction model, a pre-trained feature prediction model can be obtained.

[0046] In one selectable example, this step 310 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a model training unit operated by the processor.

[0047] Step 320: Input the sample image, the sample geometric feature map, and the sample material feature map into a pre-trained illumination prediction model to predict the illuminance value of the pixel points in the sample image, and obtain the sample illumination feature map of the sample image.

[0048] In this embodiment, the pre-trained illumination prediction model represents an illumination prediction model that can complete the prediction operation on the sample image, the sample geometric feature map, and the sample material feature map through training.

[0049] As an example, the pre-training of the illumination prediction model can be realized by using a virtual dataset. The virtual dataset can include a virtual image obtained by forward rendering processing, a virtual geometric feature map generated in the forward rendering process, a virtual material feature map, and a virtual illumination feature map. By using the virtual image, the virtual geometric feature map, and the virtual material feature map as inputs and the virtual illumination feature map as the desired output to train the initial illumination feature prediction model, a pre-trained feature prediction model can be obtained.

[0050] In one selectable example, this step 320 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a model training unit operated by the processor.

[0051] Step 330: Use a differentiable rendering module to generate a rendering image based on the sample geometric feature map, the sample material feature map, and the sample illumination feature map.

[0052] In related technologies, when generating an image by rendering, in the ray tracing stage, since the relationship between the rays received by the camera and the entire scene cannot be determined, the rendering process becomes non-differentiable. Since the backpropagation of the neural network is realized by derivation, in a non-differentiable rendering process, constraints cannot be imposed on the neural network.

[0053] In this embodiment, the sample geometric feature map, the sample material feature map, and the sample illumination feature map obtained by inverse rendering are images obtained by mapping feature values to the camera space. The differentiable rendering module directly uses the sample geometric feature map, the sample material feature map, and the sample illumination feature map to calculate shading values without performing ray tracing, thereby generating a rendering image by a differentiable rendering process.

[0054] As an example, the differentiable rendering module determines the normal vector, albedo, roughness, and metal amount of each pixel point from the sample geometric feature map, the sample material feature map, and the sample illumination feature map, and then substitutes the normal vector, albedo, roughness, metal amount, and illuminance value into the rendering equation, and then solves the rendering equation by the Monte Carlo sampling method to determine the shading value of the pixel point. Here, in order to generate a more detailed specular reflection, the importance sampling method can be used to calculate the Monte Carlo integration.

[0055] The following equations (1) to (6) show the differentiable rendering process in this example, where equation (1) is the rendering equation.

Number

Number

Number

Number

Number

Number

Number

[0056] In one selectable example, this step 330 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a model training unit operated by the processor.

[0057] Step 340, based on the difference between the sample image and the rendered image, adjust the parameters of the pre-trained feature prediction model and the pre-trained lighting prediction model until the pre-set training completion conditions are met (that is, train the pre-trained feature prediction model and the pre-trained lighting prediction model), and obtain the feature prediction model and the lighting prediction model.

[0058] As an example, the pre-set training completion conditions may be that the loss function converges, or the number of repeated executions of steps 310 to 240 reaches the pre-set number of times.

[0059] For example, the execution body can use the L1 function or the L2 function as the rendering loss function, and then determine the value of the rendering loss function based on the difference between the sample image and the rendering image. Then, by utilizing the characteristics of the backpropagation of the neural network and deriving the rendering loss function, the parameters of the pre-trained feature prediction model and the pre-trained lighting prediction model are adjusted until the function value of the rendering loss function converges, and the feature prediction model and the lighting prediction model can be obtained.

[0060] Also, for example, when the number of repeated executions of steps 310 to 340 reaches the pre-set number of times, the training is terminated, and the feature prediction model and the lighting prediction model can be obtained.

[0061] In this embodiment, based on the geometric features, material features, and lighting features obtained by inverse rendering, a differentiable rendering process is used to generate a rendering image, and based on the difference between the rendering image and the sample image, the parameters of the pre-trained feature prediction model and the pre-trained lighting prediction model are adjusted, so that physical constraints can be imposed on the feature prediction model and the lighting prediction model. As a result, the accuracy of the feature prediction model and the lighting prediction model is improved, which is helpful for improving the accuracy of the attributes obtained by inverse rendering.

[0062] In one selectable example, this step 340 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a model training unit operated by the processor.

[0063] In some selectable embodiments of the above embodiment, the pre-training process of the lighting feature prediction model can adopt the process shown in FIG. 4. As shown in FIG. 4, this process includes the following steps: Step 410, obtain an initial lighting feature map obtained by processing sample data with an initial lighting feature prediction model.

[0064] As an example, the sample data can include a virtual image obtained by forward rendering processing, a virtual geometric feature map generated in the forward rendering process, a virtual material feature map, and a virtual illumination feature map. Here, the virtual image, the virtual geometric feature map, and the virtual material feature map may be used as inputs, and the virtual illumination feature map may be used as a sample label.

[0065] In one selectable example, this step 410 may be executed by the processor calling the corresponding instruction stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0066] Step 420, based on the difference between the initial illumination feature map and the sample label, determine the value of the prediction loss function.

[0067] In this embodiment, the prediction loss function characterizes the degree of difference between the output of the initial illumination prediction model and the sample label. For example, the L1 function or the L2 function can be used as the prediction loss function.

[0068] In one selectable example, this step 420 may be executed by the processor calling the corresponding instruction stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0069] Step 430, based on the difference in the illuminance values of adjacent pixel points and the difference in the depths of adjacent pixel points in the initial illumination feature map, determine the value of the spatial continuity loss function.

[0070] Normally, the lighting environments between two adjacent points in space are close, and accordingly, the lighting environments between two far-apart points are significantly different. After mapping these two points to an image, the distance between these two points in space can be represented by the depth between pixel points.

[0071] In this embodiment, the spatial continuity loss function can represent the difference in the illumination environment between adjacent pixel points. When the difference in depth between two adjacent pixel points is small, it indicates that the illumination environments of the two are close. At this time, the value of the spatial continuity loss function is also small. Conversely, when the difference in depth between two adjacent pixel points is large, it indicates that the illumination environments of the two can be significantly different. At this time, the value of the spatial continuity loss function is also large.

[0072] In one selectable example, this step 430 may be executed by the processor calling the corresponding instruction stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0073] Step 440: Based on the value of the prediction loss function and the value of the spatial continuity loss function, train the initial illumination feature prediction model to obtain a pre-trained illumination feature prediction model.

[0074] In one selectable example, this step 440 may be executed by the processor calling the corresponding instruction stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0075] In this embodiment, the execution body repeatedly executes the above steps 410 to 440, and based on the value of the prediction loss function and the value of the spatial continuity loss function, adjusts the parameters of the initial illumination feature prediction model until the prediction loss function and the spatial continuity loss function converge, or until the number of repeated executions of steps 410 to 440 reaches the preset number of times, terminates the training, and can obtain a pre-trained illumination prediction model.

[0076] The embodiment shown in FIG. 4 embodies the step of restricting the pre-training of the illumination prediction model by using the prediction loss function and the spatial continuity loss function, and the spatial continuity loss function can give an overall restriction on the local illumination of the image to be processed, preventing sudden changes in illumination. Thereby, the pre-training of the illumination prediction model is restricted, the accuracy of the illumination prediction model can be improved, and it helps to more accurately obtain the illumination characteristics of the image to be processed.

[0077] In some selectable embodiments of the embodiment shown in FIG. 4, the value of the spatial continuity loss function can be determined by the process shown in FIG. 5. As shown in FIG. 5, this process includes the following steps: Step 510, project the illuminance value of the pixel point in the initial illumination feature map onto the adjacent pixel points to obtain the projected illuminance value of the pixel point in the initial illumination feature map, and determine the difference between the illuminance value and the projected illuminance value of the pixel point in the initial illumination feature map.

[0078] In this embodiment, the difference between the illuminance value and the projected illuminance value of the pixel point in the initial illumination feature map can characterize the difference in the illumination environment between adjacent pixel points.

[0079] As an example, the execution entity can realize the projection of the illuminance value by the projection operator, obtain the projected illuminance value of each pixel point by projecting the illuminance value of each pixel point onto the adjacent pixel points in the preset direction, and then determine the difference between the illuminance value and the projected illuminance value of each pixel point.

[0080] In one selectable example, this step 510 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by the pre-training unit operated by the processor.

[0081] Step 520, determine the scaling factor based on the depth gradient of the pixel point in the initial illumination feature map and the preset continuity weight parameter.

[0082] Here, the scaling coefficient and the depth gradient have a positive correlation.

[0083] In this embodiment, the depth gradient of a pixel point can represent the spatial distance of the point corresponding to an adjacent pixel point. The value of the continuity weight parameter may usually be set according to experience.

[0084] For example, the executing entity can first predict the depth gradients of two adjacent pixel points, and then determine the scaling coefficient based on the depth gradient and the continuity weight parameter. The scaling coefficient can allow a certain deviation in the illumination environment between at least one pixel point.

[0085] In one selectable example, this step 520 may be executed by the processor calling the corresponding instruction stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0086] In step 530, based on the difference and the scaling coefficient, determine the value of the spatial continuity loss function.

[0087] As an example, the executing entity can multiply the difference corresponding to each pixel point by the corresponding scaling coefficient, and then use the average value of the sum of the products corresponding to all pixel points as the value of the spatial continuity loss function.

[0088] As an example, the spatial continuity loss function in this embodiment can adopt the following formula (7):

Number

Number

Number

[0089] In one selectable example, this step 530 may be executed by the processor calling the corresponding instructions stored in the memory, or may be executed by a pre-training unit operated by the processor.

[0090] In the process shown in FIG. 5, the difference between the illumination environments of adjacent pixel points is represented by the difference between the illuminance value of the pixel point and the projection illumination, the scaling coefficient is determined based on the depth gradient and the continuity weight parameter of the pixel point, and the value of the spatial continuity loss function is determined based on the difference between the illuminance value of the pixel point and the projection illumination and the scaling coefficient, so that the difference between the illumination environments of points at different positions in space can be represented more accurately. For example, the illumination environments of far-apart points can be very different, and the illumination environments of nearby points can also be close. Thereby, the pre-training process of the illumination prediction model is restricted, and the illumination prediction model can learn the potential correlation between the position of points in space and the illumination environment, thereby improving the prediction accuracy.

[0091] In some selectable embodiments of the above embodiments, after obtaining the geometric feature map and the material feature map of the image to be processed through step 110, the albedo feature map and the roughness feature map can also be processed as follows: Input the image to be processed, the geometric feature map and the material feature map into a guided filtering model, determine the filtering parameters, and smooth the albedo feature map and the roughness feature map based on the filtering parameters.

[0092] In this embodiment, an albedo feature map and a roughness feature map are smoothed using a guided filtering model, so that the image quality of the albedo feature map and the roughness feature map can be improved. By inputting the smoothed albedo feature map and roughness feature map into an illumination prediction model, it helps to improve the prediction accuracy of illumination features. At the same time, by using the smoothed albedo feature map and roughness feature map to perform pre-set processing on the image to be processed, the image quality of the processed image can be improved.

[0093] As an example, the guided filtering model may be a convolutional neural network embedded with a guided filtering layer.

[0094] Furthermore, the filtering parameters are obtained as follows: Based on the image to be processed, a geometric feature map, and a material feature map, an input image is generated. The resolution of the input image is lower than that of the image to be processed. The initial filtering parameters of the input image are predicted using a guided filtering model, and the initial filtering parameters are upsampled to obtain filtering parameters that match the resolution of the image to be processed.

[0095] As an example, the resolutions of the image to be processed, the geometric feature map, and the material feature map are reduced to half of the original resolution, then input into a guided filtering model to obtain the initial filtering parameters at half the resolution, and then the initial filtering parameters are upsampled to obtain filtering parameters that match the original resolution.

[0096] In this embodiment, by reducing the resolution of the input image to obtain the initial filtering parameters and then obtaining the filtering parameters that match the input image through upsampling, the filtering parameters can be obtained more quickly, which helps to improve the efficiency of the image smoothing process by the guided filtering model.

[0097] The method for inverse rendering of an image provided by an embodiment of the present disclosure may be executed by any suitable device with data processing capabilities, including but not limited to terminal devices and servers. Alternatively, the method for inverse rendering of an image provided by an embodiment of the present disclosure may be executed by a processor. For example, the processor executes the method for inverse rendering of an image described in the embodiments of the present disclosure by calling corresponding instructions stored in a memory. Repeated descriptions are omitted below.

[0098] Those skilled in the art can understand that all or part of the steps of the above method embodiments can be completed by a program instructing relevant hardware. The program may be stored in a computer-readable storage medium. When the program is executed, the steps of the above method embodiments are executed. The storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical disks.

[0099] Refer to FIG. 6 below. FIG. 6 is a schematic structural diagram of one embodiment of an apparatus for inverse rendering of an image of the present disclosure. The apparatus of this embodiment can be used to implement each of the above method embodiments of the present disclosure. As shown in FIG. 6, the apparatus is configured to input a processing target image into a feature prediction model, predict the geometric features and material features of the processing target image by the feature prediction model, and obtain a geometric feature map and a material feature map of the processing target image. The geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map. A feature prediction unit 610, configured to input the processing target image, the geometric feature map, and the material feature map into an illumination prediction model, predict the illuminance value of the processing target image for each pixel by the illumination prediction model, and obtain an illumination feature map of the processing target image. An illumination prediction unit 620, and an image processing unit 630 configured to perform a preset process on the processing target image based on the geometric feature map, the material feature map, and the illumination feature map.

[0100] In one of the embodiments, the lighting prediction unit 620 includes a prediction module configured to use a lighting prediction model to process the image to be processed, the geometric feature map, and the material feature map, predict the illuminance value of a pixel point in the image to be processed, and generate a panoramic image corresponding to the pixel point based on the predicted illuminance value, and a stitching module configured to stitch the panoramic images corresponding to the pixel points in the image to be processed to obtain a lighting feature map.

[0101] In one of the embodiments, the apparatus inputs a sample image into a pre-trained feature prediction model to predict the geometric features and material features of the sample image, obtains a sample geometric feature map and a sample material feature map of the sample image, inputs the sample image, the sample geometric feature map, and the sample material feature map into a pre-trained lighting prediction model to predict the illuminance value of a pixel point in the sample image, obtains a sample lighting feature map of the sample image, uses a differentiable rendering module to generate a rendering image based on the sample geometric feature map, the sample material feature map, and the sample lighting feature map, and adjusts the parameters of the pre-trained feature prediction model and the pre-trained lighting prediction model based on the difference between the sample image and the rendering image until the pre-set training completion conditions are met, and further includes a model training unit configured to obtain the feature prediction model and the lighting prediction model.

[0102] In one of the embodiments, the apparatus obtains an initial lighting feature map obtained by processing sample data with an initial lighting feature prediction model, determines the value of a prediction loss function based on the difference between the initial lighting feature map and the sample label, determines the value of a spatial continuity loss function based on the difference in the illuminance values of adjacent pixel points and the difference in the depths of adjacent pixel points in the initial lighting feature map, trains the initial lighting feature prediction model based on the values of the prediction loss function and the spatial continuity loss function, and further includes a pre-training unit configured to obtain a pre-trained lighting feature prediction model.

[0103] In one of the embodiments, the pre-training unit projects the illuminance value of a pixel point in the initial illumination feature map onto adjacent pixel points, obtains the projected illuminance value of the pixel point in the initial illumination feature map, determines the difference between the illuminance value and the projected illuminance value of the pixel point in the initial illumination feature map, determines a scaling coefficient based on the depth gradient of the pixel point in the initial illumination feature map and a preset continuity weight parameter, where the scaling coefficient has a positive correlation with the depth gradient, and further includes a loss function module configured to determine the value of the spatial continuity loss function based on the difference and the scaling coefficient.

[0104] In one of the embodiments, the device further includes a filtering unit configured to input the image to be processed, the geometric feature map, and the material feature map into a guided filtering model, determine filtering parameters, and smooth the albedo feature map and the roughness feature map based on the filtering parameters.

[0105] In one of the embodiments, the device further includes a parameter determination unit configured to generate an input image based on the image to be processed, the geometric feature map, and the material feature map, where the resolution of the input image is lower than that of the image to be processed, predict the initial filtering parameters of the input image using a guided filtering model, upsample the initial filtering parameters, and obtain filtering parameters that match the resolution of the image to be processed.

[0106] In addition, the embodiments of the present disclosure a memory for storing a computer program; a processor for executing the computer program stored in the memory, and when the computer program is executed, realizing the method for inverse rendering of an image according to any one of the above embodiments of the present disclosure. Further provided is an electronic device including the processor.

[0107] In addition, embodiments of the present disclosure further provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, can implement a method for inverse rendering of an image in any of the above embodiments.

[0108] Hereinafter, an electronic device according to an embodiment of the present disclosure will be described with reference to FIG. 7.

[0109] FIG. 7 shows a block diagram of an electronic device according to an embodiment of the present disclosure.

[0110] As shown in FIG. 7, the electronic device includes one or more processors and a memory.

[0111] The processor may be a central processing unit (CPU) or another form of processing unit having data processing capabilities and / or instruction execution capabilities, and can control other components within the electronic device to perform desired functions.

[0112] The memory can store one or more computer program products, and the memory can include various forms of computer-readable storage media such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory (cache). The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products may be stored in the computer-readable storage medium, and the processor can execute the computer program products to implement the method for inverse rendering of an image in each of the above embodiments of the present disclosure and / or other desired functions.

[0113] As an example, the electronic device can further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0114] Also, the input device can further include, for example, a keyboard and a mouse, etc.

[0115] The output device can output various information including the determined distance information, direction information, etc. to the outside. The output device can include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto, etc.

[0116] Of course, for the sake of simplicity, only some of the components related to the present disclosure in the electronic device are shown in FIG. 7, and components such as a bus and an input / output interface are omitted. In addition, depending on the specific application situation, the electronic device can further include any other appropriate components.

[0117] In addition to the above methods and devices, embodiments of the present disclosure may also be a computer program product including computer program instructions. When the computer program instructions are executed by a processor, the processor executes the steps of the method for inverse rendering of images according to various embodiments of the present disclosure described in the above part of this specification.

[0118] The computer program product can be written in any combination of one or more programming languages for executing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially executed on the user's device, executed as an independent software package, partially executed on the user's computing device and a remote computing device, or executed entirely on a remote computing device or a server.

[0119] Also, the embodiments of the present disclosure may be a computer-readable storage medium storing computer program instructions. When the computer program instructions are executed by a processor, the processor executes the steps of the method for inverse rendering of an image according to various embodiments of the present disclosure described in the above part of this specification.

[0120] The computer-readable storage medium may be any combination of one or more readable storage media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium includes, but is not limited to, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0121] As described above, the basic principles of the present disclosure have been explained in conjunction with specific embodiments. However, it should be noted that the advantages, superiority, and effects mentioned in the present disclosure are merely examples and not limiting. It is not considered that these advantages, superiority, and effects are necessary for each embodiment of the present disclosure. Also, the above specific details of the present disclosure are for illustrative purposes only, for ease of understanding, and not limiting. The above details do not limit that the present disclosure must be implemented using the above specific details.

[0122] Each example in this specification is described progressively. The main content described in each example is the difference from other examples. For the same or similar parts between each example, reference may be made to each other. The examples of the system basically correspond to the examples of the method, so they are briefly described. For the relevant details, reference may be made to the partial description of the examples of the method.

[0123] The block diagrams of the devices, apparatuses, equipment, and systems according to the present disclosure are merely exemplary examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. Those skilled in the art will recognize that these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. For example, words such as "comprising", "including", and "having" are open vocabularies, referring to "including but not limited to...", and are used interchangeably therewith. The vocabulary "or" and "and" used here refer to the vocabulary "and / or", and are used interchangeably therewith unless otherwise clearly indicated in the context. The vocabulary "for example" used here refers to the phrase "for example,... but not limited to these", and is used interchangeably therewith.

[0124] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented in software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustrative purposes only, but the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified in other ways. Also, in some embodiments, the present disclosure can also be implemented as a program recorded on a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also includes a recording medium storing a program for executing the method according to the present disclosure.

[0125] Also, in the apparatuses, devices, and methods of the present disclosure, it is necessary to point out that each component or each step can be disassembled and / or recombined again. These disassemblies and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0126] The above description of the disclosed embodiments is provided so that those skilled in the art can make or use the present disclosure. Various modifications to these embodiments will be very apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments shown herein, but rather to follow the broadest scope consistent with the principles and novel features disclosed herein.

[0127] The above description has been made for purposes of illustration and description. Also, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although several exemplary embodiments have been discussed above, some of their variations, modifications, changes, additions, and sub - combinations will be recognized by those skilled in the art.

Claims

1. A method for inverse rendering of an image, comprising: inputting a target image to be processed into a feature prediction model, predicting geometric features and material features of the target image to be processed by the feature prediction model, and obtaining a geometric feature map and a material feature map of the target image to be processed, wherein the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map; inputting the target image to be processed, the geometric feature map, and the material feature map into an illumination prediction model, predicting an illuminance value of each pixel of the target image to be processed by the illumination prediction model, and obtaining an illumination feature map of the target image to be processed; performing a preset process on the target image to be processed based on the geometric feature map, the material feature map, and the illumination feature map. A method for inverse rendering of an image, characterized by comprising the above steps.

2. The step of inputting the target image to be processed, the geometric feature map, and the material feature map into an illumination prediction model, predicting an illuminance value of each pixel of the target image to be processed by the illumination prediction model, and obtaining an illumination feature map of the target image to be processed comprises: processing the target image to be processed, the geometric feature map, and the material feature map using the illumination prediction model, predicting an illuminance value of a pixel point in the target image to be processed, and generating a panoramic image corresponding to the pixel point based on the predicted illuminance value; stitching panoramic images corresponding to pixel points in the target image to be processed to obtain the illumination feature map. The method according to claim 1, characterized by comprising the above steps.

3. The method further comprises the step of obtaining the feature prediction model and the illumination prediction model, and the step is as follows: inputting a sample image into a pre-trained feature prediction model, predicting geometric features and material features of the sample image, and obtaining a sample geometric feature map and a sample material feature map of the sample image; inputting the sample image, the sample geometric feature map, and the sample material feature map into a pre-trained illumination prediction model, predicting an illuminance value of a pixel point in the sample image, and obtaining a sample illumination feature map of the sample image. Using a differentiable rendering module, a rendering image is generated based on the sample geometric feature map, the sample material feature map, and the sample illumination feature map. Based on the difference between the sample image and the rendering image, the parameters of the pre-trained feature prediction model and the pre-trained illumination prediction model are adjusted until a pre-set training completion condition is satisfied, and the feature prediction model and the illumination prediction model are obtained. The method according to claim 2 is characterized by this.

4. Further including the step of obtaining the pre-trained illumination feature prediction model, and the step is as follows: Obtain an initial illumination feature map obtained by processing sample data with an initial illumination feature prediction model. Based on the difference between the initial illumination feature map and the sample label, determine the value of the prediction loss function. Based on the difference in the illumination values of adjacent pixel points and the difference in the depths of the adjacent pixel points in the initial illumination feature map, determine the value of the spatial continuity loss function. Based on the value of the prediction loss function and the value of the spatial continuity loss function, train the initial illumination feature prediction model to obtain the pre-trained illumination feature prediction model. The method according to claim 3 is characterized by this.

5. The step of determining the value of the spatial continuity loss function based on the difference in the illumination values of adjacent pixel points and the difference in the depths of the adjacent pixel points in the initial illumination feature map includes: Project the illumination value of a pixel point in the initial illumination feature map onto an adjacent pixel point to obtain the projected illumination value of the pixel point in the initial illumination feature map, and determine the difference between the illumination value and the projected illumination value of the pixel point in the initial illumination feature map. Determine a scaling coefficient based on the depth gradient of a pixel point in the initial illumination feature map and a pre-set continuity weight parameter, where the scaling coefficient has a positive correlation with the depth gradient. Determine the value of the spatial continuity loss function based on the difference and the scaling coefficient. The method according to claim 4 is characterized by including this.

6. After the geometric feature map and the material feature map of the image to be processed are obtained, the method is as follows: Inputting the image to be processed, the geometric feature map, and the material feature map into a guided filtering model to determine filtering parameters; The method according to any one of claims 1 to 5, further comprising smoothing the albedo feature map and the roughness feature map based on the filtering parameters.

7. The method further comprises obtaining the filtering parameters, and the step is as follows: Generating an input image based on the image to be processed, the geometric feature map, and the material feature map, wherein the resolution of the input image is lower than that of the image to be processed; Using the guided filtering model to predict the initial filtering parameters of the input image, upsampling the initial filtering parameters, and obtaining filtering parameters that match the resolution of the image to be processed. The method according to claim 6.

8. An apparatus for inverse rendering of an image, A feature prediction unit configured to input an image to be processed into a feature prediction model, predict the geometric features and material features of the image to be processed by the feature prediction model, and obtain a geometric feature map and a material feature map of the image to be processed, wherein the geometric feature map includes a normal map and a depth map, and the material feature map includes an albedo feature map, a roughness feature map, and a metallicity feature map; An illumination prediction unit configured to input the image to be processed, the geometric feature map, and the material feature map into an illumination prediction model, predict the illuminance value of each pixel of the image to be processed by the illumination prediction model, and obtain an illumination feature map of the image to be processed; An apparatus for inverse rendering of an image, comprising an image processing unit configured to perform a preset process on the image to be processed based on the geometric feature map, the material feature map, and the illumination feature map.

9. The illumination prediction unit is A prediction module configured to process the image to be processed, the geometric feature map, and the material feature map using the illumination prediction model, predict the illuminance value of a pixel point in the image to be processed, and generate a panoramic image corresponding to the pixel point based on the predicted illuminance value; A stitch module configured to stitch panoramic images corresponding to pixel points in the image to be processed to obtain the illumination feature map, the apparatus according to claim 8.

10. The apparatus inputs a sample image into a pre-trained feature prediction model to predict the geometric features and material features of the sample image, and obtains a sample geometric feature map and a sample material feature map of the sample image, inputs the sample image, the sample geometric feature map, and the sample material feature map into a pre-trained illumination prediction model to predict the illuminance value of a pixel point in the sample image, and obtains a sample illumination feature map of the sample image, uses a differentiable rendering module to generate a rendered image based on the sample geometric feature map, the sample material feature map, and the sample illumination feature map, further includes a model training unit configured to adjust the parameters of the pre-trained feature prediction model and the pre-trained illumination prediction model based on the difference between the sample image and the rendered image until a pre-set training completion condition is satisfied, to obtain the feature prediction model and the illumination prediction model, the apparatus according to claim 9.

11. The apparatus obtains an initial illumination feature map obtained by processing sample data with an initial illumination feature prediction model, determines the value of a prediction loss function based on the difference between the initial illumination feature map and a sample label, determines the value of a spatial continuity loss function based on the difference in illuminance values of adjacent pixel points and the difference in depths of the adjacent pixel points in the initial illumination feature map, further includes a pre-training unit configured to train the initial illumination feature prediction model based on the value of the prediction loss function and the value of the spatial continuity loss function to obtain the pre-trained illumination feature prediction model, the apparatus according to claim 10.

12. The pre-training unit is configured to determine the value of a spatial continuity loss function based on the difference in illuminance values of adjacent pixel points and the difference in depths of the adjacent pixel points in the initial illumination feature map, that is, Project the illuminance value of the pixel point in the initial illumination feature map onto adjacent pixel points to obtain the projected illuminance value of the pixel point in the initial illumination feature map, and determine the difference between the illuminance value and the projected illuminance value of the pixel point in the initial illumination feature map. A step of determining a scaling coefficient based on the depth gradient of the pixel point in the initial illumination feature map and a preset continuity weight parameter, wherein the scaling coefficient has a positive correlation with the depth gradient. The apparatus according to claim 11, further comprising a loss function module configured to determine a value of the spatial continuity loss function based on the difference and the scaling coefficient.

13. The apparatus Inputs the image to be processed, the geometric feature map, and the material feature map into a guided filtering model to determine filtering parameters. The apparatus according to any one of claims 8 to 12, further comprising a filtering unit configured to smooth the albedo feature map and the roughness feature map based on the filtering parameters.

14. The apparatus Generates an input image based on the image to be processed, the geometric feature map, and the material feature map, and the resolution of the input image is lower than that of the image to be processed. The apparatus according to claim 13, further comprising a parameter determination unit configured to use the guided filtering model to predict initial filtering parameters of the input image, upsample the initial filtering parameters, and obtain filtering parameters that match the resolution of the image to be processed.

15. An electronic device A memory for storing a computer program product A processor that executes the computer program product stored in the memory, and when the computer program product is executed, realizes the method according to any one of claims 1 to 7. The electronic device is characterized by comprising the above components.

16. A computer-readable storage medium storing computer program instructions When the computer program instructions are executed by a processor, the computer-readable storage medium is characterized by realizing the method according to any one of claims 1 to 7.