Systems and methods for image relighting and image compositing
By using neural networks to refine shading and estimate specularities within the intrinsic image domain, the system achieves highly realistic image compositing and relighting, addressing the challenges of illumination and geometry interactions in existing methods.
Patent Information
- Application Number
- PCT/CA2024/051565
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing image compositing and relighting methods struggle to produce realistic results, especially in real-world scenarios, due to inadequate handling of illumination and geometry interactions.
The system employs neural networks to refine shading and estimate specularities, using intrinsic image decomposition and a render engine to simulate illumination and geometry, enabling accurate image compositing and relighting.
This approach generates highly realistic composite images that accurately match the color and illumination of the background, overcoming previous limitations in image harmonization and relighting.
Smart Images

Figure CA2024051565_30052025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR IMAGE RELIGHTING ANDIMAGE COMPOSITINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to and benefit of United States provisional patent application no. 63 / 602,427, and entitled “INTRINSIC HARMONIZATION FOR ILLUMINATION-AWARE COMPOSITING”, the entirety of which is hereby incorporated by reference herein.TECHNICAL FIELD
[0002] The present disclosure relates to image relighting and image compositing and in particular to image relighting and compositing though illuminationbased harmonization using neural networks.BACKGROUND
[0003] Compositing, or inserting an object onto a novel background, is an important image editing task that requires the object to naturally blend into the new environment. This requires the appearance of the inserted object to be readjusted to fit the background in a process called image harmonization. For a realistic composite, the harmonized object should match the color content as well as the illumination present in the background. However, existing methods remain inadequate in producing realistic composites.
[0004] Accordingly, systems and methods that enable accurate image compositing and relighting, particularly for real world images, remains highly desirable.SUMMARY
[0005] In accordance with one aspect of the present disclosure, an image processing method for relighting an image is disclosed, comprising: generating a shading image representative of shading for the image using an illumination component defining one or more lighting parameters for relighting the image and a geometry representation of the image; generating an input image using the shading image and an albedo image representative of albedo in the image; and processingthe input image, the albedo image, and the shading image with a first neural network trained to refine the shading to generate a relighted image corresponding to the illumination component.
[0006] In some aspects, generating the relighted images comprises: generating a refined shading image representative of the shading according to the illumination component using the first neural network; and generating the relighted image using the refined shading image and the albedo image.
[0007] In some aspects, the geometry representation is: a normal map defining surface normals of the image, a depth map, a point cloud, a set of 3D points, Gaussian splats in 3D space, a 3D mesh representation, or a 3D representation of the image.
[0008] In some aspects, generating the shading image comprises rendering the shading image using the illumination component and the geometry representation of the image.
[0009] In some aspects, the illumination component is representative of lighting from a light source.
[0010] In some aspects, the one or more lighting parameters comprises: a direction of lighting, a directional intensity of lighting, an ambient intensity of lighting, a 3D location of a light source, a 2D representation of a light source, or combinations thereof.
[0011] In some aspects, the illumination component comprises a constant ambient illumination and / or illumination from the light source; the geometry representation of the image and the illumination from the light source are 3 dimensional vectors and the constant ambient illumination is a constant value; and the shading image is a sum of the constant ambient illumination and a dot product of the illumination from the light source and the geometry representation of the image.
[0012] In some aspects, where the illumination component comprises a plurality of subcomponents, each subcomponent representative of lighting from a respective light source in 3D space; and the shading depicted in the shading image corresponds to a plurality of light sources defined by the plurality of subcomponents.
[0013] In some aspects, the first neural network is trained using a dataset comprising photographs, shading images of the photographs, training shading images, and training input images, the shading images of the photographs are generated using image decomposition; the training shading images are generated by: determining illumination components of the photographs by rendering shading estimations using geometry representations of the photographs; and rendering the training shading images using the illumination components and the geometry representations of the photographs; the training input images are generated from the training shading images and albedo images of the photographs; and the first neural network is trained to refine the shading by training the first neural network to generate the shading images from the training shading images, the training input images, or both.
[0014] In some aspects, the first neural network is trained using a loss function comprising mean squared error of the refined shading image, mean squared error of the relighted image, multi-scale gradient loss of the refined shading image, multi-scale gradient loss of the relighted image, or combinations thereof.
[0015] In some aspects, the method further comprises: generating the shading image and the albedo image by performing image decomposition on the image.
[0016] In some aspects, the image decomposition generates: the shading image as a diffuse shading image, the albedo image as a diffuse albedo image, and a residual image representative of specularities in the image.
[0017] In some aspects, the method further comprises: processing the relighted image and the refined shading image with a second neural network trained to estimate the specularities to generate an adjusted residual image representative of the specularities according to the illumination component; and generating a second relighted image using the refined residual image and the relighted image.
[0018] In some aspects, the second neural network is trained using a dataset comprising photographs, residual images of the photographs, training diffuse images, and training diffuse shading images, the training diffuse shading images are generated by: determining illumination components of the photographs; renderingshading images using the illumination components and geometry representations of the photographs; and generating training diffuse shading images from the shading images; the training diffuse images are generated from the training diffuse shading images and diffuse albedo images of the photographs; and the second neural network is trained to estimate the specularities by training the second neural network to generate the residual images from the training diffuse shading images and the training diffuse images.
[0019] In some aspects, the method further comprises a use thereof in compositing an object depicted in the image onto a background image depicting a background scene, comprising: processing a background albedo image of the background image and an object albedo image of the object using a third neural network trained to match a foreground albedo to a background albedo to generate a composite albedo image as the albedo image, the composite albedo image comprising the object albedo image composited onto the background albedo image, wherein an albedo of the object albedo image is matched to an albedo of the background albedo image by the second neural network; processing the background shading image using a render engine to perform illumination estimations to determine an illumination component of the background scene as the illumination component; processing the shading image by compositing the shading image as an object shading image onto a background shading image of the background image prior to generating the input image; the relighted image being a composite image depicting the object in the background scene.
[0020] In some aspects, the method further comprises: generating the background albedo image and the background shading image by performing image decomposition on the background with a fourth neural network trained to perform intrinsic image decomposition; and generating the object albedo image by performing image decomposition on the object image with the fourth neural network trained to perform intrinsic image decomposition.
[0021] In some aspects, the method further comprises: generating a mask corresponding to a region occupied by the object in the background scene forcompositing the object onto the background image; and applying the mask to composite images comprising the object onto images of the background scene.
[0022] In some aspects, the third neural network is trained using a segmentation dataset comprising albedo images, foreground albedo images, and background albedo images, wherein the albedo images are segmented to generate the foreground albedo images and the background albedo images, and wherein colors of the foreground albedo images are transformed according to one or more image editing parameters; and the third neural network is trained to match the foreground albedo to the background albedo by training the third neural network to generate the albedo images from the foreground albedo images and background albedo images.
[0023] In some aspects, the first neural network is trained using a segmentation dataset comprising shading images of photographs, training shading images, and training input images; the training shading images are generated by: segmenting foreground shading images from the shading images; determining illumination components of the foreground shading images by rendering shading estimations of the foreground shading images or the photographs using geometry representations of the foreground shading images or the photographs; generating second foreground shading images from the illumination components and geometry representations of the photographs; generating the training shading images by compositing the second foreground shading images onto the shading images of the photographs; and the training input images are generated from the training shading images and albedo images of the photographs; and the first neural network is trained to refine the shading by training the first neural network to generate the shading images from the training shading images and the training input images.
[0024] In some aspects, the method further comprises: determining geometry representations of the background scene; and processing the geometry representations of the background scene using the render engine to determine the illumination component of the background shading image; the render engine being configured to determine the illumination component by rendering estimated shading images using estimated illumination components.
[0025] In some aspects, the illumination component of the background shading image comprises a constant ambient illumination and / or illumination from a light source; the geometry representations of the background scene and the illumination from the light source are 3 dimensional vectors and the constant ambient illumination is a constant value; and the background shading image is a sum of the constant ambient illumination and a dot product of the illumination from the light source and the geometry representations of the background scene.
[0026] In some aspects, the method further comprises: generating a mask for compositing the object onto the background image; generating a depth map of the background scene from the background image; determining geometry representations of the background scene; determining geometry representations of the object; and processing the mask, the depth map, the geometry representations of the background scene, and the geometry representations of the object with the first neural network to generate the refined shading image.
[0027] In some aspects, the render engine is configured to determine the illumination component using gradient-based least square optimization to minimize a difference between the background shading image and an estimated background shading based on the illumination component.
[0028] In accordance with another aspect of the present disclosure, a method of training neural networks for performing image processing is disclosed, comprising: obtaining shading images of photographs; generating training shading images by: determining illumination components of the photographs; and rendering the training shading images using the illumination components and geometry representations of the photographs; generating training input images from the training shading images and albedo images of the photographs; and training a first neural network to refine shading of an image by training the first neural network to generate the shading images from the training shading images and the training input images.
[0029] In some aspects, the shading images are diffuse shading images and wherein the albedo images are diffuse albedo images.
[0030] In some aspects, the method further comprises: obtaining residual images of the photographs; generating training diffuse shading images from the training shading images using the first neural network; generating the training diffuse images from the training diffuse shading images and the albedo images of the photographs; and training a second neural network to estimate specularities in the image by training the second neural network to generate the residual images from the training diffuse shading images and the training diffuse images.
[0031] In some aspects, the method further comprises: generating foreground albedo images and background albedo images by segmenting the albedo images; generating foreground shading images and background shading images by segmenting the shading images; transforming colors of the foreground albedo images according to one or more image editing parameters to generate transformed foreground albedo images; training a third neural network match a foreground albedo to a background albedo by training the third neural network to generate the albedo images using the transformed foreground albedo images and the background albedo images; the third neural network being trained to generate a composite albedo image; a render engine being configured to perform illumination estimations to determine an illumination component for generating an object shading image representative of shading of an object for compositing onto a background scene, the object shading image for compositing onto a background shading image representative of shading of the object under the illumination in the background scene for generating a composite shading image; the composite shading image for generating a composite image; the first neural network being trained to generate a refined shading image from the composite shading image and the composite image; and the refined shading image for generating a relighted image with the composite albedo image.
[0032] In accordance with another aspect of the present disclosure, an image processing method for compositing an object onto a background image depicting a background scene is disclosed, the method comprising: obtaining a background albedo image and a background shading image of the background image; obtaining an object albedo image of the object; processing the background albedo image and the object albedo image with a first neural network trained to match a foregroundalbedo to a background albedo to generate a composite albedo image, the composite albedo image comprising the object albedo image composited onto the background albedo image, wherein an albedo of the object albedo image is matched to an albedo of the background albedo image by the first neural network; processing the background shading image using a render engine configured to perform illumination estimations to determine an illumination component of the background shading image corresponding to illumination in the background scene; generating, according to the illumination component, a second object shading image representative of shading of the object under the illumination in the background scene; generating a composite shading image by compositing the second object shading image onto the background shading image; generating a composite image from the composite shading image and the composite albedo image; processing the composite image and the composite shading image with a second neural network trained to determine a shading for a composited image to generate a second composite shading image; and generating a second composite image from the second composite shading image and the composite albedo image.
[0033] In accordance with another aspect of the present disclosure, an image processing method for compositing a light emitting object onto a background image depicting a background scene is disclosed, the method comprising: obtaining a background albedo image and a background shading image of the background image; obtaining an object albedo image of the light emitting object; generating, according to the object albedo image and the background albedo image, a composite albedo image; processing the background shading image using a render engine configured to perform illumination estimations to determine an illumination component of the background shading image corresponding to illumination in the background scene; augmenting the illumination component with an illumination component of the light emitting object; generating, according to the illumination component, a composite shading image comprising shading of the light emitting object and shading of the background scene representative illumination from the augmented illumination component; generating a composite image from the composite shading image and the composite albedo image; processing the composite image and the composite shading image with a second neural network trained to determine a shading for acomposited image to generate a second composite shading image; and generating a second composite image from the second composite shading image and the composite albedo image.
[0034] In accordance with another aspect of the present disclosure, a system is disclosed, comprising one or more processing units configured to perform the method of any one of the above aspects.
[0035] In accordance with another aspect of the present disclosure, a non- transitory computer-readable medium having computer readable instructions stored thereon is disclosed, which, when executed by one or more processing units, causes the one or more processing units to perform the method of any one of any one of the above aspects.BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Further features and advantages of the present disclosure will become apparent from the following detailed description, taken in combination with the appended drawings, in which:FIGS. 1-2C depict a system for performing image compositing and / or image relighting, according to example embodiments.FIGS. 3A-3C depict images for use in and generated by the system of FIGS. 1-2C, according to example embodiments.FIG. 4 depicts a method for estimating illumination in an image for use in the system of FIGS. 1-2C, according to an example embodiment.FIG. 5 depicts a method for training a neural network to determine a shading for a composite image for use in the system of FIGS. 1-2C, according to an example embodiment.FIG. 6 depicts a method for training a neural network to determine an albedo for a composite image for use in the system of FIGS. 1-2C, according to an example embodiment.FIGS. 7 and 8 depict images generated by the system of FIGS. 1-2C in comparison to images generated by other methods, according to example embodiments.FIG. 9 depicts manual adjustments to a composite image generated by the system of FIGS. 1-2C, according to an example embodiment.FIG. 10A depicts a method for performing image compositing using the system of FIGS. 1-2C, according to an example embodiment.FIG. 10B depicts a method for performing image relighting using the system of FIGS. 1-2C, according to an example embodiment.
[0037] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0038] Image relighting has been studied in the field as an image to image conversion method
[0029] , where systems take the image under initial illumination as input and are expected to produce a new image under the a new illuminating environment. Due to the complexity of interactions between the illumination and the scene content, and the problem definition that requires the same scene illuminated under different conditions, it requires large volumes of ground truth data which is expensive to capture. This makes previous methods limited to specific class of objects such as portraits
[0029] or indoor environments limiting their real-world applicability.
[0039] The disclosed systems and methods can overcome these challenges by introducing self-supervised relighting systems and methods that uses data-driven techniques and computer graphics theory in combination to achieve in-the-wild generalization. In particular, the disclosed systems and methods for relighting may comprise four aspects:(1 ) generating an estimation of physical attributes of the input image including a geometric representation thereof as well as an albedo image thereof representing surface reflectance and a shading image thereof representing the effects of the illumination in the scene, where the albedo and shading image can be generated via intrinsic decomposition;(2) simulating shading representing the new illumination of the scene after relighting using a rendering formulation, such as normal-based Lambertian formulations or a complete rendering engine with ray tracing, using the geometric representation estimated for the input scene;(3) combining the simulated shading with the albedo estimated for the scene (e.g. albedo image) to generate an initial relit image; and(4) using a neural network to generate the final realistic relit image that takes the initial relit image and the simulated target shading as input.
[0040] By defining a target illuminating environment through simulating a target shading map, full physical control of lights in the scene can be achieved. As such, the processing performed by the neural network may be significantly simplified when compared to image-to-image relighting methods. The image-to-image relighting methods take the image under original illumination as input and may be required to learn to understand the existing light while also modeling how the relit image appearance will change. Instead, the neural network of the disclosed systems and methods may receive as input the illumination-invariant albedo layer, and a simulated shading which can model, up to a precision, how the relit image will appear after processing. As a result, the neural network can serve as a light refinement and enrichment modeling engine to generate higher quality results in addition to providing physical controllability.
[0041] One technical advantage of the disclosed systems and methods can be that, due to the physical modeling of the illumination using estimated scene geometry and intrinsics, the neural network can be trained using only a set of images without requiring relighting or any other ground truth. Accordingly, the disclosed systems and methods (e.g. the neural network) are self-supervised, and can generate its own training data from any given image. This makes can make required training dataset much cheaper, easier to obtain, and easier to scale, when compared to other relighting approaches. In particular, the disclosed systems and methods for training this network may comprise three aspects:(1) generating an estimation of physical attributes of an image in the dataset including a geometric representation thereof as well as an albedo image thereof and a shading image thereof via intrinsic decomposition;(2) determining the illuminating environment in the image by simulating the shading present in the scene using a rendering formulation, such as normal-based Lambertian formulations or a complete rendering engine with ray tracing, and using the geometric representation estimated for the input scene (e.g. optimizing the parameters of the illuminating environment such that the simulated shading best approximates the shading estimated for the input image);(3) combining the simulated shading with the albedo image for the scene to generate an approximation of the original image; and(4) training the neural network to estimate the original image and / or the estimated original shading using the simulated shading and the albedo image.
[0042] In particular, during training, the network may be trained to re-generate the original image only from a simulation of the existing illuminating environment and the illumination-invariant albedo. As such, the images in the dataset can be used as ground truth, representing the realistic image under this illuminating environment, where training images can be generated using the computer graphics techniques based on estimated geometric representation of the image.
[0043] Image harmonization has been studied in the field as an image editing problem where self-supervised methods may be used to estimate a set of color and tone adjustment operations to match the color contents of the object with the background. Relighting the object to match the illumination in the novel environment, however, has yet to be addressed in the field. This comes from the difficulty of realistically relighting an object in-the-wild, which requires an accurate estimation of the illuminating environment as well as a detailed and accurate geometry. As a result, most relighting methods in the field may be limited to a specific domain such as portraits [29, 37], While restricting the domain of relighting can make it possible togenerate training data in controlled environments for image-to-image relighting, such methods do not yield satisfactory results, particularly outside of the restricted domain, in part due to the deficiencies of the training data. As such, relighting methods in the field lacking in general and are inapplicable to the general problem of image compositing.
[0044] Despite significant advancements in network-based image harmonization techniques for use in image compositing by matching or harmonizing an inserted object orforeground to the background scene of an image, there still exists a domain disparity between typical training pairs and real-world composites encountered during inference. Most existing methods can be trained to reverse global edits made on segmented image regions, which fail to accurately capture the lighting inconsistencies between the foreground and background found in composite or composited images (used interchangeably herein). The present disclosure relates to a self-supervised illumination harmonization approach formulated in the intrinsic image domain. In particular, a global lighting model determined by a neural network from mid-level vision representations can be used to generate a rough shading for the foreground region or object. Another network can refine this inferred shading to generate a harmonious re-shading that aligns with the background scene. In order to match the color appearance of the foreground and background, harmonization approaches can be used to perform parameterized image edits in the albedo domain. The present disclosure can provide highly accurate composite images from challenging real-world composites.
[0045] According to a broad aspect, an illumination-aware image harmonization approach for image compositing is disclosed. The disclosed systems and methods may be formulated in the intrinsic image domain. First, at least one neural network can be used to generate albedo, shading and surface normals for the input composite or object and the background image. By using a first neural network, it is possible to harmonize the albedo of the background and foreground by predicting image editing parameters. Using the generated normals and shading, a render engine can be used to estimate a simple lighting model for the background illumination. With this lighting model, it is possible to render a shading, such as a Lambertian shading,for the foreground and refine the shading using a second neural network trained on segmentation datasets via self-supervision. Accordingly, composite images modeling realistic lighting effects can be generated.
[0046] In particular, the disclosed systems and methods may comprise three aspects:(1) perform harmonization on the estimated albedo of the scene using a first neural network;(2) estimate a geometry or geometric representation (e.g. surface normals) and shading of the background to optimize the parameters of a simple lighting model using a render engine; and(3) use the lighting model to render a predicted shading, such as a Lambertian shading, of the foreground. The Lambertian shading can be refined using a second neural network to generate a final realistic shading layer that can be combined with the harmonized albedo to create a composite image.
[0047] It should be noted that surface normal(s) are referenced herein as an example parameter for defining a geometry of an object or contents depicted in an image. However, other geometry representations and estimations are possible as well in the context of the present disclosure. Example geometry representations can include: a normal map, a point cloud, a set of 3D points, Gaussian splats in 3D space, a 3D mesh representation, or similar 3D representations. Further, Lambertian shading are referenced herein as an example of a predicted or estimated shading, for example rendered using lighting parameters and geometry representations. Predicted shadings are not limited to shading under the Lambertian assumption and the types of predicted or estimated shadings are possible as well in the context of the present disclosure.
[0048] In accordance with another aspect of the present disclosure, the disclosed systems and methods may also be applicable in image relighting. For example, instead of estimating the simple lighting model according to the background illumination, the simple lighting model can be defined or set according to a desiredlighting condition for relighting an image. As used herein, relighting or image relighting may refer to artificially altering the lighting conditions on contents, such as objects or a scene, depicted in an image, where a relighted image would depict how the content would appear under the altered lighting conditions. Specifically, the lighting model can be used to represent the altered lighting conditions instead of the lighting conditions of the background. The simple lighting model can be used to then render an estimated shading of the image to be relighted, for example, as Lambertian shading using surface normals. The second neural network can then refine the Lambertian (or other) shading of the image to be relighted to generate a shading representative of the relighted scene, which can be combined with the albedo image of the original image to be relighted to generate the relighted image. It should be noted that this process is based on the property that the albedo is illumination invariant and as such would be the same for the original and the relighted image. In other words, image compositing, for example as performed by the disclosed systems and methods, may be considered as a form of relighting. For example, in image compositing, in order to composite the object such that it appears naturally in the background scene, the object must be “relighted” to match the lighting conditions of the background scene. Accordingly, the generation and refinement of the predicted or estimated shading (e.g. by the second neural network) for the object in image compositing is for relighting the object under the background lighting conditions using the simple illumination model. As such, it is possible to relight the object or any other image by instead setting the simple illumination model to the desired lighting conditions for relighting. In the context of relighting, the object or foreground, the geometry (e.g. surface normals) thereof, shading thereof, and albedo thereof may be or may correspond to entire images, rather than segments or composites thereof.
[0049] In accordance with the present disclosure, it is possible to model the image harmonization problem of compositing an object or foreground onto a background scene of an image in the intrinsic domain through image decomposition. As used herein, the foreground may refer to an object or image to be composited onto another image. Specifically, the object may be a 2D image, such as one or more objects or scenes (e.g. a foreground scene) depicted in an image. Alternatively, the object may be a 3D object, such as a 3D rendering, representation, or model havinggeometry information; and the 3D object may be rendered, compressed, or otherwise represented as an image, for example after being rendered using a rendering engine or algorithm. As used herein, the background may refer to the image (i.e. a background image) such as a photograph on which the foreground is to be composited onto, or the scene which is depicted in the image. As used herein, harmonize, harmonized, or harmonizing may refer to matching of an element, parameter, depiction, or image to another. In particular, harmonizing may refer to matching the appearance of an image or portion thereof to another image or portion thereof. Specifically, harmonizing A to B can refer to transforming an image, A, such that the appearance of A once harmonized is as it would naturally appear in B. In the case of an object being harmonized to a background, harmonizing can refer to adjusting the depiction of the object such that it is reflective of how it would appear should the object be physically placed in the background. That is, the harmonized object should depict how the object would appear when subject to the same conditions as the background. As another example, harmonizing an illumination on A to B can refer to reproducing the illumination of B to alter the appearance of A to be representative of how A would appear under the illumination of B. As used herein, relighting may refer to altering an appearance of an object or image to be reflective of how the object or image would appear under different illumination or lighting conditions. As used herein, model, where appropriate (e.g. outside of the context of computer implemented artificial intelligence models / networks), such as illumination model and Lambertian model, may refer to a set of conditions, parameters, or governing relationships.
[0050] In accordance with the present disclosure, it is possible to model the combined problem of compositing an light emitting object onto a background scene of an image in the intrinsic domain through image decomposition. Compositing a light emitting can require, in addition to compositing the object, altering the existing illumination in the scene by also considering the newly composited light emitting object. The simulated shading, in this case, can be generated using the illuminating environment estimated for the background image together with the composited object added to the illumination representation as a light source. This process may beperformed using the described the neural networks without re-training, for example by using the compositing and relighting systems and methods.
[0051] Intrinsic image decomposition can be referred to as a mid-level vision problem that represents an image as the product of the reflectance of the materials and the effect of illumination in the scene following the relationship of: I = S-A, where I is an image and where S and A respectively represent the shading and albedo of the image. The shading and albedo may be considered as components of the image and can be depicted or represented using a shading and albedo image, respectively, where the shading image can be a single channel greyscale image and the albedo image can be a 3-channel RGB image. Specifically, albedo may refer to a reflectance of the contents, such as object(s) / material(s) / scene depicted in the image and shading may refer to an effect of illumination on the contents depicted in the image. By isolating the scene colors contained in the albedo from shading, the intrinsic representation of the image into the shading and albedo can divide the image harmonization problem into two: color harmonization and relighting.
[0052] As described herein, the disclosed systems and methods can harmonize the color content of the foreground to match that of the background scene in the albedo space, for example, using a first neural network. Turning to the relighting problem in the shading domain, it is possible to generate a new shading for the composited object that reflects the new illumination environment. For this purpose, it is possible to generate an initial estimation of the shading that can be derived using a shading model and object geometry, such as a Lambertian shading model and surface normals, estimated for the background and the inserted object. The illumination environment of the background can be estimated according to a parametric illumination model, for example, by using a render engine. The Lambertian shading for the inserted object can then be generated, with this shading map being composited onto the original shading of the background, for processing by a second neural network being a re-shading network. Together with the RGB composite, the initial shading can be used as input to train the second neural network to generate a realistic new shading for the object. As such, the re-shading network can be trained in a self-supervised manner using segmentation datasets.
[0053] Therefore, by dividing the harmonization problem into two and formulating a self-supervised relighting method, realistic composite images can be generated where the inserted object not only reflects the color content of the image but also matches the illumination present in the background. Accordingly, the systems and methods of the present disclosure can generate much more realistic composite images when compared to other image harmonization approaches, even in challenging scenarios.
[0054] As an example, previously proposed image harmonization methods have predominantly been trained in a self-supervised manner in order to undo various image edits performed on a specific region of a natural image [36, 6, 15, 5], These edits only represent image-level differences (brightness, saturation, hue, etc.) between the harmonized ground truth and the unharmonized input. This creates a gap between training and inference as real-world composites require more complex operations to be properly harmonized. Various approaches have been proposed to model the image harmonization problem with more realistic assumptions, including illumination harmonization. One approach
[0018] proposes a perceptually-inspired shading model to relight image segments. This approach first estimates a simple geometry model that can be reshaded and then utilize low-level algorithms to estimate shading effects that are not represented by the rough geometry.
[0055] The disclosed systems and methods utilize estimated geometry that can be reshaded, but can replace the approximate shading model with a data-driven shading refinement process that can be trained via self-supervision.
[0056] As another example, other methods [11 , 10] formulate an approach based on the intrinsic image domain to jointly separate and harmonize albedo and shading. This method often fails to generate novel shading due to a lack of proper training data and learning to harmonize and perform intrinsic decomposition with one network. Another approach
[0012] is to develop a generated dataset and method to relight humans in outdoor scenes. This approach focuses on a specific use case and requires ground-truth geometry as input, making it difficult to use in the wild. Similarly, another approach [1] proposes a dataset and method to perform illumination harmonization, but it mainly focuses on relighting objects on a flat ground plane inoutdoor scenes. Another approach [2] formulates an intrinsic approach that re-shades a masked object using Deep Image Prior. This approach fails to generate realistic shading estimations and requires multiple minutes of run time for a single image. Another approach
[0034] proposes a simpler method of modeling illumination changes in image harmonization, but this approach can only locally modulate color-based editing of the foreground region. Given this limitation and lack of explicit reasoning about albedo, this approach is not able to fully harmonize composites with major differences in lighting between foreground and background.
[0057] In contrast, the disclosed systems and methods can utilize image harmonization approaches while also fully modeling the illumination mismatch present in real-world composite images. A self-supervised training approach is disclosed herein that can employ large-scale segmentation datasets to perform realistic relighting of foreground regions that match the background lighting environment for in-the-wild composites.
[0058] As an example of object insertion, rather than attempting to composite image segments into novel scenes, the task of object insertion can also refer to inserting a 3D object into a 2D photograph of a scene. This can be accomplished by inverse rendering 2D scenes to infer characteristics of the 3D scene they depict. One approach
[0024] proposed an algorithm to recover lighting information from an image segment. The approach uses their lighting information to relight recovered 3D geometry of an object making it appear as if it belongs in the scene. Other approaches can recover a 3D representation of a scene complete with lighting that can then be used to re-render a 3D object using a typical rendering pipeline [13, 14], With the advent of deep learning, inverse rendering techniques may be data-driven, leveraging large-scale synthetic datasets of indoor scenes to train networks for estimating scene intrinsics [21 , 22],
[0059] In contrast, the disclosed systems and methods need not assume a 3D object as input and instead can work completely in the 2D image domain. Additionally, the disclosed systems and methods can work both indoors and outdoors while not requiring large-scale datasets for inferring scene characteristics and may make useof off-the-shelf networks for mid-level vision tasks such as generation of geometry representations and intrinsic decomposition.
[0060] As an example of image relighting, due to its complexity, some approaches focus on a specific use case such as portraits [29, 37] or outdoor structures [9], These approaches rely on large-scale, difficult-to-obtain datasets, or multi-view scenes [31 , 32, 28] in order to achieve realistic results . Additionally, if the target illumination is not provided, as can be the case in image harmonization, lighting must be predicted. Generating estimations of lighting configuration or illumination can be a difficult task and also may rely on hard-to-capture datasets [38, 8],
[0061] The disclosed systems and methods can relight the foreground in the wild, from a single image by leveraging mid-level estimations and without requiring ground-truth source and target illumination examples, which can be off-the-shelf. Additionally, the disclosed systems and methods can utilize a simple lighting estimation formulation to the second neural network (i.e. re-shading network) while still generating realistic shading estimations.
[0062] Embodiments are described below, by way of example only, with reference to Figs. 1-10B.
[0063] FIG. 1 depicts a system for performing image compositing and / or image relighting, according to an example embodiment, shown in FIG. 1 as one or more servers 108. The implementation of the servers 108 is not restrictive and servers 108 may be a physical server, cloud-based server, or a hybrid thereof, for example. A user 102 may interact with the servers 108 via a device 104 over a communications network 106 (e.g. the internet). The device 104 may be a computer, as depicted in FIG. 1 , but is not restricted to those expressly shown and may be any suitable device known in the art such as smart phones and tablets. The servers 108 may provide a graphical user interface (GUI) on the device 104 for ease of communication and operation control by the user. The implementation of the GUI is not restrictive and may be, for example, a mobile / computer application or a web page. The GUI can be used to provide input to and receive output from the servers 108. Additionally oralternatively, other user interfaces, such as an audio interface that allows receipt and processing of spoken commands, maybe used.
[0064] The user 102 may be interested in performing image compositing and / or image relighting via the servers 108. In particular, image compositing may be performed by compositing an object onto a background. The object may be referred to as the foreground, and may be an object image 120 depicting the object, in particular a 2D image, or a 3D object depicted, rendered, or represented as a 2D image. The background image 122 may be a 2D image depicting a background scene, such as a photograph. The object image 120 and the background image 122 may each be a 3-channal RGB image. Image compositing may comprise inserting the object image 120 onto the background image 122 in such a manner that the object appears naturally in the background image 122. That is, the composite image should depict the object of the object image 120 as it would naturally appear in the background scene of the background image 122. Image relighting may alter lighting, lighting conditions, or lighting parameters or otherwise simulate a new set of lighting, lighting conditions, or lighting parameters on object(s) or scene(s) depicted in an image 128 such that the object(s) or scene(s) appears to be under the altered or new lighting in the relighted image. The servers 108 may be configured to process the object image 120 and the background image 122 to perform image compositing in order to generate a composite image 124 for the object image 120 and the background image 122. In particular, the object image 120 and the background image 122 may be processed by at least one neural network 126 to generate the composite image 124, as described further herein. The servers 108 may also be configured to process the image 128 to with the at least one neural network 126 to perform image relighting in order to generate a relighted image 130, as described further herein. In some embodiments, image relighting may form a part of image compositing. For example, compositing the object image 120 onto the background image 122 may require the object to be relighted under the lighting of the background scene depicted in background image 122. That is, the object image 120 may be a specific embodiment of the image 128 and the composite image 124 may be a specific embodiment of the relighted image 130. The neural networks 126 may each be an artificial intelligence (Al) model or algorithm, a machine learning model or algorithm,and may each be, in particular, a convolutional neural network. The composite image 124 and the relighted image 130 may be returned to the user 102 from the server 108 to the device 104 for display, for example over the communications network 106.
[0065] According to the present disclosure, the image 128, the object image 120 and the background image 122 may be provided to or retrieved by the servers 108, for example, from the device 104. The image 128, the object image 120 and the background image 122 can also be retrieved from one or more external devices and / or one or more databases. The image 128, the object image 120 and the background image 122 may be requested and received using an application programing interface (API) via requests / calls and responses, for example over the communications network 106, although other forms of communication such as Bluetooth and near-field communication are possible as well. The servers 108 may also receive, retrieve, or determine surface normals of the image, the object and the background scene and / or depth maps thereof.
[0066] In a particular implementation, the servers 108 each comprise a CPU 110, a non-transitory computer-readable memory 112, a non-volatile storage 114, an input / output interface 116, and graphical processing units (“GPU”) 118. The non- transitory computer-readable memory 112 comprises computer-executable instructions stored thereon at runtime which, when executed by the CPU 110, configure the server to perform the above described processes of sustainability analysis. The non-volatile storage 114 has stored on it computer-executable instructions that are loaded into the non-transitory computer-readable memory 112 at runtime. The input / output interface 116 allows the server to communicate with one or more external devices such as the device 104 (e.g. via network 106). The non- transitory computer-readable memory 112 may also have stored thereon the at least one neural network 126. The GPU 118 may be used to control a display and may be used to process the object image 120 and the background image 122 to generate the composite image 124 and / or to process the image 128 to generate the relighted image 130, including for example by running the neural networks 126 to process the image 128, the object image 120 and the background image 122, as described further herein. In some embodiments, the neural networks 126 may be stored at one or moreseparate servers. The servers 108 and the device 104 may each provide a communications interface which allows software and data to be transferred, for example between the servers 108 and the device 104 over the communications network 106.
[0067] The CPU 110 and GPU 118 may be one or more processors or microprocessors, which are examples of suitable processing units, which may additional or alternatively comprise an artificial intelligence accelerator, programmable logic controller, a microcontroller (which comprises both a processing unit and a non-transitory computer readable medium), Al accelerator, neural processing unit (NPU), or system-on-a-chip (SoC). As an alternative to an implementation that relies on processor-executed computer program code, a hardware-based implementation may be used. For example, an application-specific integrated circuit (ASIC), field programmable gate array (FPGA), or other suitable type of hardware implementation may be used as an alternative to or to supplement an implementation that relies primarily on a processor executing computer program code stored on a computer medium.
[0068] It should be noted that while FIG. 1 depicts the device 104 and the servers 108 as separate entities coupled over the communication network 106, the device 104 and servers 108 may also be coupled directly / physically using cable(s) for data transfer. In some embodiments, the servers 108 may also be the device 104 or comprise the device 104 (e.g. the servers 108 being implemented as a part of a computer system). In such an embodiment, the servers 108 may directly retrieve the image 128, the object image 120 and the background image 122 (as well as any other required data) from fixed local storage or removable local storage. In some embodiments, the device 104 and / or the servers 108 may be coupled to an instrument / d evice (e.g. a camera) that is configured to capture the image 128, the object image 120 and the background image 122 such that the image 128, object image 120 and the background image 122 can be received directly from the instrument / d evice once captured.
[0069] Referring now to FIG. 2A, the object image 120 and the background image 122 can be processed to generate mid-level representations. For example,using the relationship of l=SA, the background image 122 depicting the background scene and the object image 120 depicting the foreground object to be composited can be represented as Iband lf, respectively. In which case, the object image 120 and the background image 122 can be decomposed into their shading and albedo components, each of which can be represented using an image, following the below relationship:If = Sf ■ Af, Ib= Sb• Ab, where S and A refer to the shading and albedo component, respectively, which can be in the form of images, and the superscripts f and b refer to the foreground (i.e. object) and the background (i.e. background scene), respectively.
[0070] That is, image decomposition, in particular intrinsic image decomposition 202 can be performed on the object image 120 and the background image 122 to respectively generate: 1) object albedo image 210 and object shading image 212, and 2) background albedo image 214 and background shading image 216. In particular, the albedo images referred to herein may be 3-channel RBG albedo images and the shading images referred to herein may be single channel greyscale shading images. Image decomposition 202 may be performed in any suitable manner. For example, one or more neural networks trained to determine a shading and an albedo component of an image may be used to perform image decomposition 202 by processing the object image 120 and the background image 122 to generate the albedo and shading images 210, 212, 214, 216. The at least one neural network comprise one or more publicly available or “off-the-shelf” networks, for example in use for specific tasks. The at least one neural networks may be implemented using suitable architecture, parameter tuning, and training parameters, as described further herein, in some embodiments, image decomposition 202 can be performed according to a plurality of known or “off-the-shelf” methods (e.g. as disclosed in [4]). In some embodiments, the albedo and shading images 210, 212, 214, 216 may be obtained using other methods, or are available for retrieval with the object image 120 and the background image 122.
[0071] The compositing of the images can be performed using a mask 220. The mask 220 may define a boundary of the object to be inserted, for example enclosing the object or otherwise identifying a location and a region occupied by the object. The mask 220 may be an overlay such as a transparent layer. The mask 220 may be coded as a layer of binary values of 0 and 1 for identifying pixels in an image that should be replaced with the object or the background scene, where the mask of the object image 120 may be an inverse of the mask of the background image 122. For example, the regions (i.e. selected pixels) corresponding to the object in the object image 120 may be coded as 0, with all other regions coded as 1 , with regions coded as 1 replaced with a transparency layer such that the background scene is shown in the region coded as 1 after compositing, for example by overlaying or inserting the object onto the background image 122. The mask 220 may also be used for inserting the object onto the background scene, for example by segmenting the object from an image (e.g. the object image 120), and / or by segmenting out from the background image 122, the region corresponding to the region occupied by the object when composited. Specifically, the mask 220 may be used to replace a portion of the background image 122 with the object at a desired insertion location and may be used to replace the non-object portions of the object image 120 with the background scene. In cases where the size of the object image 120 is different from the background image 122 or in cases where the object images 120 comprise only the object itself, the mask 220 may be expanded or scaled to the size of the background image, with the region corresponding to the object being positioned at a desired location for compositing onto the background scene. It should be noted that the mask 220 may be applied to the object image 120 or both the object image 120 and the background image 122. The mask 220 can be applied for composite images or for compositing images such as albedo and shading images. As used herein, the object and background images 120, 122, as well as albedo and shading images derived therefrom may be the unmodified images (e.g. unsegmented) corresponding to the original content of the images or a masked / segmented version thereof subject to the application of the mask 220. Accordingly, the shading and albedo of the object image 120 and the background image 122 can be defined and composited using the mask 220 following the below relationship:Ac= aAf+ (1 - a)Ab, Sc= aSf+ (1 - a)Sb, lc= Ac■ Sc, where a is the mask 220 and (1- a) is the inverse thereof, and where Ic, Ac, and Screpresent the input composite image 302 and its intrinsic components, specifically an input albedo image 304 and an input shading image 306, which respectively represent the naive or unmodified albedo and shading of the unmodified, original composite image 302. The input composite image 302 may be a naive or unmodified composite, Ic, can be derived by combining the above defined Ac, and Scdetermined from the masked albedo and shading of the albedo and shading images 210, 212, 214, 216 or by applying the mask 220 to the object image 120 and / or the background image 122 and compositing the images 120, 122.
[0072] Examples of the input composite image 302 are shown in FIGS. 3A-3C. FIGS. 3A and 3C also depict examples of the input shading image 306, as well as the mask 220 used to generate or composite the images 302, 304, 306. As depicted in FIGS. 3A-3C, the input composite image 302 comprises the object from the segmented or masked object image 120 composited onto the background scene of the background image 122, which may also be masked or segmented. Similarly, as depicted in FIGS. 3A and 3C, the input shading image 306 comprises the shading of the object from the segmented or masked object shading image 212 composited onto the shading of the background scene of the background shading image 216, which may also be masked or segmented; the input albedo image 304 comprises the albedo of the object from the segmented or masked object albedo image 210 composited onto the albedo of the background scene of the background albedo image 214, which may also be masked or segmented.
[0073] In particular, linear RGB may be used when performing any albedo and shading operations. For example, image decomposition 202 may assume linear RGB values. In some embodiments, when given a standard RGB image as input (e.g. for image decomposition 202), it is possible to reverse the gamma-correction process using a gamma value of 2.2. This separation of the compositing problem can allow coIor and illumination harmonization to be performed in two separate steps.
[0074] Referring back to FIG. 2A, a first neural network 204 can process the object albedo image 210 and the background albedo image 214, one or both of which may be segmented using the mask 220, to generate a harmonized albedo image 218. Specifically, the color content of a scene can be represented in its albedo, therefore a color-based harmonization method through the use of the first neural network 204 can be used to adjust the foreground albedo to harmonize with the colors in the background albedo. In particular, parameter-based harmonization approaches [34, 15] can be used. The first neural network 204 may be trained to harmonize or match the albedo of the object, as depicted in the object albedo image 210, to the albedo of the background scene, as depicted in the background albedo image 214. In order to perform albedo harmonization, the first neural network 204 may aim to find a set of editing parameters that control common image editing operations such as changing the exposure, saturation, color curve, and white balance to harmonize the albedo of the object to the background scene. The harmonized albedo image 218 may be a composite albedo image comprising the (harmonized) object albedo 308, which is harmonized to the albedo of the background scene, and composited or inserted onto the albedo of the background scene as represented by the background albedo image 214. The first neural network 204 may output the harmonized albedo image 218 directly, for example, instead of receiving the object albedo image 210 and the background albedo image 214 as input, the first neural network 204 may receive the input albedo image 304. The mask 220 may also be provided as input to identify the region in the input albedo image 304 corresponding to the object albedo to be harmonized. Alternatively, the first neural network may output the harmonized object albedo 308, for example as a harmonized object albedo image, or a segmented harmonized object albedo image, for example by applying the mask 220. The harmonized object albedo image can accordingly be composited onto the background albedo image 214, or a segmented version thereof, for example by applying the mask 220.
[0075] Examples of the harmonized albedo image 218 are depicted in FIGS. 3A-3C, each of which comprises the harmonized object albedo 308, which is composited onto the background albedo image 214, comprising the original / unmodified albedo of the background scene.
[0076] Referring now to FIG. 6, a method for training the first neural network 204 is depicted, according to an example embodiment. The training can be conducted through a self-supervised setup using segmentation datasets. In particular, the training data may comprise a plurality of original, natural images 602 such as photographs. Each image 602 may comprise or depict an object or foreground 604 in a background scene 618. The training data may also comprise a segmentation mask 606 for each image 602, which can be functionally identical and generally correspond to masks 202. The segmentation mask 606 may be configured to segment the object 604 of the image 602 from the background scene 618 of the image 602.
[0077] At 620, intrinsic image decomposition may be performed on the images 602 to generate original albedo images 608, for example, as described with respect to FIG. 2A. In some embodiments, the original albedo images 608 may be readily obtained or are available for retrieval. At 622, the original albedo images 608 may be segmented to generate the object or foreground albedo 616 corresponding to the unmodified albedo of the object 604 and background scene albedo 610 corresponding to the unmodified albedo of the background scene 618. For example, the segmentation mask 606 may be configured to segment the object albedo 616 of the object 604 from the background scene albedo 610 of the original albedo image 608. In some embodiments, the images 602 may be segmented prior to image decomposition, for example, to generate the object albedo 616 and the background scene albedo 610 separately. At 624, the object albedo 616 may be modified or transformed using image edits to generate modified object albedo 612. For example, one or more image editing operations, techniques or parameters may be applied to the object albedo 616, for example automatically using an image editing program or algorithm. The image edits may be changing the exposure, saturation, color curve, and white balance. In some embodiments, the image edits may be applied to the object 602 prior to image decomposition rather than the object albedo 616 such that image decomposition generates the modified object albedo 612. At 626, the modified object albedo 612 can be composited onto the background scene albedo 610 to generate the modified object albedo image 614, for example using the mask 606. That is, a set of image editing operations are applied to the object albedo 616 to create a mismatch between a simulated composited albedo being the modified object albedoimage 614 and the original albedo image 608. At 628, the first neural network 204 can be trained to generate the original albedo image 608 (e.g. as ground-truth) from the modified object albedo image 614, for example, by estimating the image edits that will re-create the original appearance of the original albedo image 608, Ac. That is, the first neural network 204 may be trained to generate the original albedo image 608 by reversing the applied image edits using a training dataset comprising a plurality of original albedo images 608 and modified object albedo images 614.
[0078] In one embodiment, datasets from MS-COCO
[0020] and Davis
[0030] can be used. These dataset can also comprise the segmentation mask 606 for segmenting each of the objects 604 present in the scene of the original image 602 from the background scene 618. In particular, each image 602 may have associated therewith a plurality of segmentation masks 606 for segmenting a plurality of different objects 604 in the scene. During training, one of the masks 606 can be selected, for example randomly, to segment the original image 602. At least one image edit or image editing operation / technique / parameter, for example a random number between 1 to 4, can be applied to the object albedo 616 corresponding to the albedo of the segmented object 604. An order for the at least one image edit, and values for each of the edits can be sampled, for example uniformly at random, from a pre-specified range for application to the region specified by the mask 606 (i.e. the object albedo 616). Example image edits and ranges thereof are shown below in Table 1.Table 1 : Example image edits for generating training data
[0079] In one embodiment, the neural network of Miangoleh et al.
[0025] for image editing may be used as a basis for the first neural network 204. This network can be trained to regress image edits that increase or decrease the saliency of a specific region, for example the object albedo 616. The target ranges for the affine transformsof the estimation head of parameters as outlined in Miangoleh et al.
[0025] can be updated to be [0.1 ,1] white balance, [0,2] saturation, [0,2] color curve values, and [0.5,2] exposure.
[0080] In one embodiment, the first neural network 204 may be trained with the ADAM optimizer, for example with a learning rate of 1 e-5for 100 epochs with a batch size of 64. A random ordering of the 4 image edits or image editing operations (e.g. out of 24 possible permutations) can be sampled for each training batch and provided to the first neural network 204 as input during training. The permutations may also be selected at random during inference.
[0081] Referring back to FIG. 2A, a render engine 206 and a second neural network 208 may be used to generate the shading of the composite image 124, for example, as an adjusted composite shading image 234. Accordingly, there can be a need to perform reshading for the object, for example, by matching the shading of the object to the shading of the background scene such that it is representative of the illumination in the background scene. This process may be considered as harmonizing of the illumination of the object (e.g. foreground region), in a composite image, and can be posed as a specific case of image relighting. Achieving physically- accurate relighting can require an accurate High Dynamic Range Imaging (HDRI) illumination environment together with a highly detailed geometry to physically rerender the object. However, the creation of a relighting dataset for data-driven relighting can be highly challenging and hard to scale due to the controlled setup such a dataset requires. Therefore, instead of training a network for image-to-image relighting, in order to create a novel foreground shading that matches the background environment, it is possible to determine Lambertian shading under a Lambertian shading model with a parametric illumination representation, for example using the render engine 206. The render engine may be a program, algorithm, platform, or other implementation configured to simulate how light interacts with objects in a 3D space and calculate how the objects should appear when viewed from a specific camera perspective. The render engine 206 may process or render geometry representations and light models to generate images such as shading images that depict the appearance of objects under various lightings. It should be noted that while thepresent disclosure makes references to Lambertian shading or Lambertian assumption, any suitable predicted or estimated shading may be utilized to estimate the illumination representation and as an intermediate shading for refinement by a second neural network 208. That is, Lambertian shading and the Lambertian assumption is provided as an example of shading modeling, prediction, and estimation. The second neural network 208 can be used to refine the predicted (e.g. Lambertian) shading in order to generate a realistic shading map, for example the adjusted composite shading image 234, to generate the final composite image 124. Specifically, by defining the relighting problem as the refinement of the predicted shading, for example Lambertian shading, the networks 206, 208 can be trained in a self-supervised manner using standard segmentation datasets.
[0082] To determine the illumination in the background scene of the background image 122, an illumination model may be determined or estimated for the background scene. The illumination model may define a set of lighting or illumination parameters that is indicative or representative of a lighting condition in a scene, in this case the background scene of the background image 122. The illumination model may be defined or determined using predicted shading, for example shading under the Lambertian shading assumption using the Lambertian shading model by assuming that the reflection of the light in a scene is perfectly diffuse. As used herein, the illumination component 232 can be used to refer to the illumination model. Specifically, the illumination component 232 can be defined as a combination of a directional light source represented as a combination of a directional light source represented by the vector I e IR3and a constant ambient illumination c. In some embodiments, the illumination component 232 can also comprises one or more lighting parameters such as a 3D location of a light source, a 2D representation of a light source. The Lambertian shading model can be used to represent the shading of a scene using the surface normals and the illumination environment. In relation to the illumination component 232, the predicted (e.g. Lambertian) shading for the background image 122, Sb, can be defined using the below relationship:Sb= nb• I + c,where nbcan represent a surface normal of the background scene, a background normal 224, at each pixel i, and where the surface normal of an object or scene depicted in an image can be a set of vectors perpendicular to a surface at each pixel i. It should be noted that surface normals, as referred to in the present disclosure, is provided as an example of estimated geometry representation of the contents depicted in an image and may also be a normal map, a point cloud, a set of 3D points, Gaussian splats in 3D space, a 3D mesh representation, or similar 3D representations.
[0083] That is, the shading of the background image 122, determined as the background shading image 216, Sb, as described above, may be decomposed or constructed as a function of the illumination component 232, representing the lighting condition in the background scene, and the background normal 224, representing the surface normals of the background scene, for example under the Lambertian shading assumption.
[0084] As depicted in FIG. 2A, the render engine 206 is configured to determine the illumination component 232 of the background scene and may take as inputs the background shading image 216 and the background normal 224. Specifically, the render engine 206 can, based on the above described relationship for the background shading image 216, Sb, determine the illumination component using the below defined least squares optimization:
[0085] That is, the render engine 206 may be configured to search for the settings or values of I and c, which forms the illumination component 232, such that a predicted shading such as Lambertian shading, represented using a shading image generated by rendering the background normal 224 with the determined or estimated values of the illumination component 232, best reconstructs the estimated shading of the background scene, which is the background shading image 216. To ensure that the values of the illumination component 232 (e.g. values of I and c) take on plausible values, the values of I and c to may be constrained to be positive. This keeps I inthe outward-facing hemisphere. The minimization problem as defined above may be solved by the render engine 206 using gradient-based optimization with the Adam optimizer
[0016] ,
[0086] The process for determining the illumination component 232 using the render engine 206 is depicted in FIG. 4, in accordance with an example embodiment. The background image 122 depicts a background scene under a specific set of lighting parameters, represented by the illumination component 232. As depicted in FIG. 4, the illumination component 232 may be visualized using a spherical lighting model, which shows the direction of the incoming light, the directional intensity thereof, and the ambient intensity thereof. As an example, the render engine 206 may estimate a plurality of illumination components 232, which can be combined with the background normal 224 to generate predicted shading images 404 corresponding to the estimated illumination components 232. As described above, the background normal 224 represents a set of vectors corresponding to the surface normals of the background scene depicted in the background image 122. The background normal 224 is depicted as a background normal image or map in FIG. 4. For each of the estimated illumination components 232, a corresponding predicted shading image 404 can be generated, for example by rendering (402) the estimated illumination component 232 onto / with background normal 224 with the render engine 206. Specifically, the predicted shading image 404 may be a shading image representative of the shading of the background scene subject to the lighting parameters defined by the respective predicted illumination component 232, for example under the Lambertian assumption. Each predicted shading image 404 may be compared to the background shading image 216, for example by determining a difference in the value of each pixel in greyscale. The render engine 206 can be configured to minimize the difference between the rendered predicted shading image 404 and background shading image 216 by generating or determining an illumination component 232 that can render a predicted shading image 404 that is most similar to the background shading image 216, for example, by ranking the plurality of generated illumination components 232.
[0087] Referring back to FIG. 2A, the illumination component 232 for the background scene determined by the render engine 206 can be combined with object normal 222, corresponding to the surface normals of the object as depicted in the object image 120. In particular, the illumination component may be determined by taking into account the object’s geometry, such as by rendering the illumination component 232 onto / with object normal 222 to generate an adjusted object shading image 226 corresponding to a predicted or estimated shading of the object when subject to the same lighting parameters (i.e. illumination component 232) as the background scene, for example under Lambertian assumption. The adjusted object shading image 22Q, Sf, can be represented using the following relationship, which corresponds to the relationship that defines the background shading image 216, Sb, for example under the Lambertian assumption:
[0088] In some embodiments, the gradient-based optimization using constraints for training the render engine 206 may be too slow. As such, the constraints may be removed such that the optimization may be solved with leastsquares. The least squares optimization may be performed using the geometry representations, for example normals 222, 224 and shading of the foreground region (i.e. the object) corresponding to, for example, the adjusted object shading image 226 and the object shading image 212, rather than the background corresponding to, for example background shading image 216 and the predicted shading image 404. This is possible since the original image 122 is already harmonized, and as such will yield reliable lighting parameters. Additionally, this can allow the second neural network 208 to learn to rely on the predicted (e.g. Lambertian) shading, for example the predicted shading image 404, provided as input which results in a controllable lighting refinement network.
[0089] As depicted in FIG. 2A, the adjusted object shading image 226 can be composited onto the background shading image 216, for example using the mask 220, to generate a composite shading image 228, Sc, which can be representative of the shading of the composite image comprising the object and the background scene, for example under the Lambertian assumption. It should also be noted that the mask220 can also be applied at other times to other object-based images, where desired / appropriate. For example, the mask 220 may be applied to the object normal 222, which can then generate a segmented adjusted object shading image 226 for compositing. Further, the surface normals 222, 224 for the object and the background scene (i.e. the object image 120 and the background image 122) may be available for retrieval, for example with the object image 120 and the background image 122. Geometry representations, for example surface normals 222, 224 may also be determined or generated using a number of different methods, such those readily available [7],
[0090] As depicted in FIG. 2A, the composite shading image 228 can be adjusted or refined by a second neural network 208 to generate an adjusted composite shading image 234, Sc, which more accurately represents the shading of the composite image, specifically the shading of the object if it is naturally present in the background scene. In particular, the adjusted composite shading image 234, Sc, can be a realistic shading map such that, when multiplied with the albedo, in particular the harmonized albedo image 218, reconstructs the illumination-harmonized composite image 124. This definition can enable the training of the second neural network 208 in a self-supervised fashion. The second neural network 208 can generate the adjusted composite shading image 234 using the composite shading image 228, a predicted composite image 230, the object normal 222, the background normal 224, the mask 220, as well as a depth map 234. The predicted composite image 230 can be a 3-channel RGB composite image generated from the adjusted composite shading image 234 and the harmonized albedo image 218, for example by multiplying the images 234, 218. The depth map 234 may provide depth information for the background scene of the background image 122. The depth map 234 may be available for retrieval, for example with the background image 122, or may be determined or generated using a number of different methods, such as described in
[0026] , The normals 222, 224 and the depth map 234 can be provided to the second neural network 208 for geometry context and the mask 220 can be provided to identify the location or position of the object as well as albedo and shading thereof in the background scene / background images. In some embodiments, the second neuralnetwork 208 may also directly generate the composite image 124 from the inputs thereto.
[0091] Referring now to FIG. 5, a method fortraining the second neural network 208 to determine an accurate shading for generating the composite image 124 by refining or reshading the composite shading image 228 is depicted, according to an example embodiment. The training data for the second neural network 208 may comprise real images 502 such as photographs captured with an image capturing device or obtained from available sources. The images 502 can be used to generate the input - ground-truth pairs with a segmentation mask 504, representing a composited region 544 corresponding to the object. For example, the mask 504 can be used to segment the object, represented as a region 544, corresponding to the object to be composited. The mask 504 can also be used to segment images such as albedo and shading images to generate images therefrom that correspond to the background scene and the object (i.e. region 544).
[0092] At 530, intrinsic image decomposition can be performed on each of the images 502 to generate original albedo image 506 and original shading image 508 for each image 502, for example, as described with respect to FIG. 2A. In some embodiments, the images 506, 508 may be readily obtained or are available for retrieval. At 532, the original albedo image 506 may be segmented to generate the object or foreground albedo 524 corresponding to the unmodified albedo of the object (i.e. region 544) and background scene albedo 522 corresponding to the unmodified albedo of the background scene. The original shading image 508 may also be segmented to separate the object or foreground shading corresponding to the unmodified shading of the object (i.e. region 544) and background scene shading 520 corresponding to the unmodified shading of the background scene. In some embodiments, the images 502 may be segmented prior to image decomposition, for example, to generate the object albedo 524 and the background scene albedo 522 separately as well as to separate the background scene shading 520. At 534, an illumination component 514 corresponding to the lighting conditions or parameters on the background scene can be determined. In particular, surface normal 510 of the scene depicted in the image 502 may be generated or obtained, as described above.The illumination component 514 may be determined by processing the surface normal 510 and the original shading image 508, for example using the render engine 206, as described with respect to FIGS. 2 and 4. Alternatively, the illumination component 514 may be determined by the neural network 206 from the background scene shading 520 and the corresponding object geometry such as its surface normals, for example by segmenting the surface normal 510 using the mask 504. At 536, the illumination component 514 may be used to generate an adjusted shading image 512 via rendering with the surface normal 510, as described above. The adjusted shading image 512 may be a predicted shading image, for example a Lambertian shading image representative of the shading of the scene depicted in the image 502 subject to the lighting parameters defined by the illumination component 514 under the Lambertian assumption. The illumination component 514 can be segmented using the mask 504 to generate an adjusted object shading 516, corresponding to the predicted shading of region 544 (i.e. the object). In some embodiments, the surface normal 510 can be segmented using the mask 504 instead to directly generate the adjusted object shading 516. At 538, the adjusted object shading 516 can be composited onto the original shading image 508 or the background scene shading 520, for example using the mask 504, to generate the composite shading image 526. At 540, a predicted composite image 528 can be generated from the composite shading image 526 and the original albedo image 506, for example by multiplying the images 506, 526. At 542, the second neural network 208 may be trained to generate the original shading image 508, for example by reversing the above process by learning to map the predicted (e.g. Lambertian) shading depicted in the composite shading image 526 to the accurate and realistic shading depicted in original shading image 508, defined as the ground-truth. The second neural network 208 can accept as input images 508, 526, 528 as well as the mask 504, the surface normal 510 (or segments thereof corresponding to the object / region 504 and the background scene) and a depth map of image 502, obtained as described above, where the training dataset can comprise a plurality of the inputs 508, 526, 528, 504, 510, and the depth maps. In some embodiments, the second neural network 208 can also be trained to directly generate the input images 508 from the inputs thereto.
[0093] In some embodiments, the training dataset for the second neural network 208 may comprise entire images that are not segmented images or composited images, for example generated by applying the mask 504. In particular, intrinsic decomposition may be performed on the original image 502 to generate the original albedo image 506 and the original shading image 508. The illumination component 514 and the surface normals 510 can be used to render the adjusted shading image 512, as described. The predicted composite image, for example corresponding to the Lambertian composite, can be generated from the adjusted shading image 512 and the original albedo image 506. That is, none of the images 506, 508, 512, 528 and the surface normals 510 are segmented or comprise composites. The second neural network 208 can generate the composite shading image 526 using images 506, 512, 528 and the surface normals 510, where the performance of the second neural network 208 may be optimized and tuned by evaluating the composite shading image 526 against the original shading image 508 as ground-truth as well as by evaluating a composite image generated from the composite shading image 526 and the albedo image 506 or generated directly against the original image 502 as ground-truth.
[0094] The second neural network 208 may be trained using losses defined on: (1) the shading (£s), for example between the original shading image 508 and the composite shading image 526; as well as (2) a composite image 124 (£ , generated using an adjusted composite shading image 234 and the original shading image 508, for example in comparison to the original image 502. In particular, the losses may be formulated using mean squared error as follows:where Scand Icare the original shading image 508 and the original image 502, respectively, and where Scis the estimated shading, corresponding to the adjusted composite shading image 234 generated by the second neural network 208, and Ic= Ac■ Scis the estimated composite image, corresponding to the composite image 124, generated using an adjusted composite shading image 234 and the original shading image 508. The second neural network 208 may also be trained using a multi-scalegradient loss
[0023] as an edge-aware smoothness loss for the shading (£sg) and the composite image 124as follows:where 7Sc>mcan represent the gradient of the original shading image 508, Sc, at scale m and , VIc>mcan represent the gradient of the original image 502, Ic, at scale m. In some embodiments, any of the above described losses may be combined during training to compute a final loss. For example, a overall loss, £, for the training of the second neural network 208 may be represented as follows:where the overall loss can be computed using unit weights.
[0095] In one embodiment, it is possible to channel-wise concatenate the inputs for the second neural network 208 as h x w x 9 input, where h and w represent the height and width of the image, respectively. In another embodiment, the training dataset for the second neural network 208 may comprise 3 datasets: for example the COCO Dataset
[0020] , a 50,000 image subset of the SA-1 B Dataset
[0017] , and the MultiIllumination Dataset (MID)
[0027] , For COCO and SA-1 B where the mask 504 is provided, the masks 504 are used to sample foreground segments corresponding to region 544 (i.e. the object) that are sufficiently large. For MID, the segments provided in the dataset are used to represent the region 544, where the provided multiple illuminations are for dataset augmentation. The second neural network 208 may comprise an encoder-decoder architecture, for example used by Ranftl et al.
[0033] which consists of a ResNext101
[0035] encoder and a RefineNet
[0019] decoder. The second neural network 208 may be trained using the Adam optimizer with a learning rate of 1 x 10“5for 2 million iterations.
[0096] In one embodiment, the second neural network 208 can be trained at a resolution of (384x384) with a batch size of 8. For each batch, training images (e.g. images 502) are non-uniformly sampled from each of the 3 datasets. In some embodiments, the image sampling may be biased towards the multi-illuminationdataset, as the intrinsic decomposition method used herein may be trained on this data and therefore can generate reliable shading and albedo estimates in original albedo images 506 and original shading images 508. This can be beneficial as there can be less shading information left in the albedo, and as such, the second neural network 208 can learn to generate novel outputs rather than recovering the source illumination conditions from cues in the albedo (i.e. albedo images 506).
[0097] Referring back to FIG. 2A, the adjusted composite shading image 234 can be used to generate the composite image 124, for example with the harmonized albedo image 218 by multiplying the images 218, 234. In some embodiments, the neural network 208 can also directly generate the composite image 124 rather than first generating the adjusted composite shading image 234.
[0098] Referring now to FIG. 2B, the disclosed system can also be applied for use in image relighting, according to an example embodiment. As depicted in FIG. 2B, image relighting may be performed on an image 128 to generate a relighted image 130. In particular, the image 128 may depict content such as object(s) or scene(s) under a set of illumination or lighting conditions or parameters; for relighting, the object or scene of the image 128 should be depicted as under an altered or new set of illumination or lighting conditions or parameters once relighted as the relighted image 130, where the set of illumination or lighting conditions or parameters may be chosen or set by the user 102. It should be noted that image compositing as described with respect to FIG. 2A may be a specific embodiment or implementation of image relighting. Specifically, in order to composite the object onto the background scene, the object is relighted according to the lighting conditions of the background scene, as represented by the illumination component 232. In general, the illumination component 232 can dictate the lighting conditions for an image. In the case relighting for image compositing, the illumination component 232 may control the lighting conditions for the object to be the same as the background scene. For a more general case of relighting, the illumination component 232 can be any new or altered lighting conditions to relight the contents of the image under. As such, the system of FIG. 2A is also applicable for image relighting.
[0099] To relight the image 128, an illumination component 232 of the lighting parameters for relighting can be determined, defined, selected, or generated. That is, the lighting parameters of the illumination component 232, such as the directional light source and the ambient light I and c can be set manually such that these parameters correspond to the desired relighting parameters being the new or altered lighting conditions. Surface normals 246, or another suitable geometry representation, for the image 128 or the contents depicted therein can be obtained or determined, as described above. The surface normals 246 and the illumination component 232 can be rendered using the rendering process 402 by the render engine 206 to generate a rendered shading image 248 representative of the shading of the contents depicted in the image 128 when subject to the relighting illumination parameters as defined in the illumination component 232 under e.g. the Lambertian assumption, as described above. The intrinsic image decomposition process 202 can also be applied to the image 128 to determine its albedo and shading as the albedo image 240 and shading image 242, respectively. As the albedo of the image 128 is illumination-invariant, an input image 250 can be generated using the albedo image 240 and the rendered shading image 248, for example by multiplying the images 248, 250. The input image 250 can be a RGB image that is representative of the contents depicted in the image 128 underthe relighted illumination parameters, for example based on the Lambertian assumption. The second neural network 208 can process one or more of the input image 250, the rendered shading image 248, the surface normals 246, as well as a depth map of the image 128 (e.g. obtained or generated as described above) to determine a realistic or refined shading of the image 128 under the new or altered lighting conditions as the adjusted shading image 234. The adjusted shading image 234 and the albedo image 240 can be used to generate the relighted image 130, for example by multiplying the images 234, 240. In some embodiments, the second neural network 208 can also directly generate the relighted image 130 without first generating the adjusted shading image 234.
[0100] It should be noted that the system and process described in reference to FIG. 2B generally corresponds those described in reference to FIG. 2A. Further, the system and process are also generally interchangeable for whole images, for example as described with respect to relighting, segments of image, for examplesegmented by the mask 220, or composite images, such as images comprising the object and the background scene and albedo / shading images thereof, as described with respect to image compositing. In particular, the illumination component 232 and utilization thereof is consistent, with the only difference being how it is determined. For image compositing, the illumination component 232 is determined based on the background scene in the background image 122; for image relighting, it is set, for example manually, to specific parameters representative of the new or altered lighting condition for relighting. The surface normals 246 may correspond to the object normal 222, which can be used with the illumination component 232, once determined, to generate the adjusted object shading image 226 in FIG. 2A. The adjusted object shading image 226 may be composited onto the background shading image 216 to generate the composite shading image 228, corresponding to the rendered shading image 248 in FIG. 2B. The images 228, 248 may be representative of an estimated shading of the final image, respectively the composite image 124 and the relighted image 130, where the shading estimate is based on the Lambertian assumption. The albedo image 240 may correspond to the harmonized albedo image 218. While the images 218, 240 may be determined differently, both of which are representative of the albedo for the final image, respectively the composite image 124 and the relighted image 130. The input image 250 may correspond to the predicted composite image 230, which represents an estimate of the final images 124, 130 under the Lambertian assumption. It should also be noted that while only portions of images 228, 230, for example the portion corresponding to the object, may reflect the predicted shading, for example under the Lambertian assumption, and thus requires refinement in shading by the second neural network 208, the second neural network 208 can process images where the entirety of the image reflects the predicted shading, for example images corresponding to the Lambertian assumption (e.g. images 248, 250) or where only a portion thereof reflects such an assumption (e.g. images 228, 230). The adjusted shading image 234 may correspond to the adjusted composite shading image 234, both of which representative of the shading in the final image, respectively the relighted image 130 and the composite image 124. It should be also noted that the composite image 124 may be considered as a relighted image 130, in that a portion thereof, for example corresponding to the object, is relighted.
[0101] In some embodiments, the illumination component 232 may be representative of lights from a light source, for example represented using the spherical lighting model or a set of lighting parameters described above. Alternatively, the illumination component 232 may comprise a plurality of light sources. For example, the illumination component 232 may comprise a plurality subcomponents, each of which corresponds to a respective light source. That is, each subcomponent may define a set of lighting parameters corresponding to the respective light source. Accordingly, it is possible to composite additional light source(s) by augmenting the illumination component 232 with additional subcomponent(s) defining lighting parameters of the additional light source(s). The compositing of light sources may also be an embodiment of relighting. The render engine 206 can render a predicted shading image 248 from the illumination component 232 corresponding to the plurality of light sources. The rendered shading image 248 can be used to generate the input image 250 with the albedo image 240 for shading refinement and relighting using the second neural network 208, as described above. In some embodiments, the albedo image of each light source may be composited onto the albedo image 240 to generate a modified albedo image, which can be used to generate the input image 250 with the rendered shading image 248 and with the adjusted shading image 234 to generate the relighted image 130.
[0102] In another embodiment, the disclosed systems and methods may also be used for compositing an object that is light emitting onto a background scene of the background image 122. In particular, illumination from the light emitting object may be defined considered as a subcomponent of the illumination component 232, as described above. That is, the light emitting object may be considered as a light source. The light emitting object may be rendered or represented as the object image 120. The object albedo image 210 can be determined using image decomposition 202, and used as input with the background albedo image 214 to the first neural network 204 to generate the harmonized albedo image 218, as described above. In some embodiments, the harmonized albedo image 218 may be generated by directly compositing the object albedo image 210 onto the background albedo image 214. The illumination component 232 of the background scene can be determine using the render engine 206, as described above. The determined illumination component maybe augmented with the illumination component of the light emitting object as a subcomponent to generate an adjusted illumination component 232 that comprises the light sources from the background scene and the light emitting object and can reflect the lighting conditions or parameters resulting from the original lighting in the background scene and the light emitted by the light emitting object. The render engine 206 can be used to render an estimated shading, corresponding to the composite shading image 228, using geometry representations and the adjusted illumination component 232. For example, the geometry representations may correspond to a combination or composite of geometry representations of the light emitting object and the background scene, such as a composite of object normals 222 and background normals 224. The composite shading image 228 can be used with the harmonized albedo image 218 to generate the predicted composite image 230 and can be used as input to the second neural network 208 for generating the adjusted composite shading image 234 and the composite image 124, as described with reference to FIG. 2A.
[0103] In brief, compositing a light emitting object onto a background image 122 may comprise obtaining the background albedo image 214 and the background shading image 216 of the background image 122; obtaining the object albedo image 210 of the light emitting object; generating a composite albedo image being the harmonized albedo 218 using the object albedo image 210 and the background albedo image 214 (e.g. by compositing the images 210, 214); processing the background shading image 216 using the render engine 206 to determine the illumination component 232 of the background shading image 216 corresponding to illumination in the background scene; augmenting the illumination component 232 with an illumination component of the light emitting object; generating, according to the augmented illumination component 232, the composite shading image 228 comprising shading of the light emitting object and shading of the background scene representative illumination from the augmented illumination component 232 (e.g. by rendering the geometric representations of the light emitting object and the background scene and the illumination component 232 using the render engine 206); generating the predicted composite image 230 from the composite shading image 228 and the composite albedo image (e.g. by multiplying the composite albedo imageand the composite shading image 228); processing the predicted composite image 230 and the composite shading image 228 with the second neural network 208 to generate the adjusted composite shading image 234; and generating the composite image 124 from the adjusted composite shading image 234 and the composite albedo image.
[0104] A further embodiment of the present disclosure is depicted in FIG. 2C, according to an example embodiment. In particular, intrinsic decomposition 202 may not be limited to the determination of albedo and shading for an image. Specifically, intrinsic decomposition on an image to generate only the albedo and shading image may be based on a greyscale intrinsic diffuse model where the shading is represented in a single, greyscale channel. Such a model may assume that all surfaces of contents depicted in an image are diffuse and may not consider specularities and specular surfaces. In some cases, the single channel shading may also cause color information to be embedded in the albedo, for example when there are multiple light sources and inter-reflections. Accordingly, intrinsic decomposition 202 may also be performed according to an intrinsic residual model represented using the relationship below: / = Ad* Sd+ R, where an image I can be decomposed into its diffuse albedo Adand colorful diffuse shading Sdcomponents corresponding to the diffuse illumination effects, with a residual layer R containing non-diffuse illumination effects, each of which are 3- channel RGB images.
[0105] That is, intrinsic decomposition 202 may be performed on the object image 120, the background image 122, and the image 128 to generate the diffuse albedo image, diffuse shading image, and residual image for each of the images 120, 122, 128. In particular, the object albedo image 210, background albedo image 214, and albedo image 242 may be diffuse albedo images and the object shading image 212, background shading image 216, and shading image 242 may be diffuse shading images. As such, the above described systems and methods, for example with reference to FIGS. 1-2B, can perform image compositing or image relighting, as the case may be, using the diffuse albedo images and diffuse shading images as theobject albedo image 210, object shading image 212, background albedo image 214, background shading image 216, and albedo image 242.
[0106] As depicted in FIG. 2C, the generated composite image 124 or the relighted image 130, as the case may be, can be images which only depicts the effects of diffuse illumination. A third neural network 260 can be utilized to process the composite image 124, optionally with the corresponding adjusted shading image 234, to determine a residual image 262 corresponding thereto representative of the effects of non-diffuse illumination on the contents depicted in the composite image 124. In particular, portions of the residual image 262 corresponding to the object may depict the effects of non-diffuse illumination on the object according to the illumination component 232, estimated for the background scene. Similarly, the third neural network 260 can be utilized to process the relighted image 130, optionally with the corresponding adjusted shading image 234, to determine the residual image 262 corresponding thereto representative of the effects of non-diffuse illumination on the contents depicted in the relighted image 128, for example corresponding to the altered lighting conditions defined by the illumination component 232. The residual image 262 and the composite image 124 can be used to generate an adjusted composite image 264 representative of both diffuse and non-diffuse illumination on the object and the background scene as depicted in the composite image according to the illumination component 232. Similarly, the residual image 262 and the relighted image 130 can be used to generate an adjusted relighted image 266 representative of both diffuse and non-diffuse illumination on content depicted in the image 128 according to the illumination component 232.
[0107] The third neural network 260 may be generally analogous to the second network 206 and may have the same general architecture and be trained using the same parameters and tuning settings. For training, a dataset comprising photographs may be used. Intrinsic decomposition may be performed on the photographs to generate diffuse albedo images thereof, diffuse shading images thereof, and residual images thereof. Diffuse training images can be generated from the diffuse albedo images and the diffuse shading images, for example through the multiplication of the two images. The diffuse albedo images, the diffuse shading images, and / or the diffusetraining images can be input to the third neural network 260 to train the network to estimate the specularities or effects of non-diffuse illuminations in an image. The third neural network 260 can be trained to output the residual image 262, which may be evaluated against the residual images of the photographs as ground-truth. Alternatively or additionally, the residual images 262 may be used (e.g. multiplied) with the diffuse training images to generate predicted images for evaluation against the photographs as ground-truth.Performance Evaluation
[0108] For evaluation, an example embodiment of the disclosed systems and methods was used to generate composites (i.e. composite or composited images 124), which were compared against various other image harmonization methods for generating composite images. The evaluation was focused on methods that also attempt to model illumination harmonization. A user study was also performed to compare multiple methods on difficult composite images.
[0109] To evaluate the effectiveness of the example embodiment, a qualitative user study was performed. The user study was modeled after that of Wang et. al.
[0034] and comprised a two alternatives forced choice survey. The example embodiment was compared to 3 other methods, including the method of Bhattad et al. [2], which also performs illumination harmonization in the intrinsic domain; Wang et al.
[0034] , which utilizes a gain map to modulate image edits to model non-global edits; and Ke et al.
[0015] , which proposes a parametric harmonization approach for high-definition images. Each of these methods can produce results at high-resolution making them suitable for qualitative comparison against the example embodiment.
[0110] Fifty composited images were generated using free-to-use images from Unsplash™ using each method with the aim of creating difficult examples with a mismatch in color and illumination between the foreground (i.e. object) and background regions. Examples of composited images generated by all methods are shown in FIG. 7. Specifically, FIG. 7 depicts the naive composite images 702 generated by compositing objects 706 onto the background scenes directly using masks 704 without modification. Composited images 710 were generated by theexample embodiment, which can be compared to composited images 708 generated using the other methods described above, as well as composited images (not shown) generated by an example embodiment that does not comprise the render engine 206 and as such did not perform any processing related thereto, such as the adjustment of the shading to account for the illumination component 232. For each composite 708, 710, the composited images 710 generated by the example embodiment was compared to the naive composite 702 and the composited images 708 generated by the other methods described above.
[0111] The images were generated as image pairs each comprising 702, 708, 710, resulting in 5x50=250 pairs. Each participant of the user study was given context about image compositing and told to “determine which image has the foreground object better matching the background environment’. They were then shown a random set of 50 pairs without duplicate photographs (i.e. the original background image). Responses from 70 subjects were collected, resulting in 3500 total comparisons. Following example harmonization works [34, 6], it is possible to analyze the responses using the Bradley-Terry model [3] to generate global ranking scores for each method. The results of the user study are shown in Table 2 below. The example embodiment appears to be preferred over all other methods. Additionally, the example embodiment is preferred significantly more when the illumination is harmonized, for example by determining the illumination component 232 using the render engine 206, showing that illumination is an important aspect of realism when it comes to in-the- wild composites.Table 2: Comparison of image compositing methods based on user responses[001 12] As shown in FIG. 7, the other methods fail to harmonize the illumination differences between the foreground (i.e. object) and background. Referring to the composited images 710, generated by the example embodiment, in comparison to the composited images 708, generated by the other methods, the composited images 710 are able to: attenuate the shadow on the hat from the foreground’s original outdoor environment (row 1 ); and estimate the bright outdoor light shining on the building facade in the background (row 2). Further, as seen in composited images 708, the other methods fail to soften the lighting on the box (row 3), leaving the original direct lighting that does not match the blue ambient lighting from the background. The example embodiment is also able to dim the overall brightness of the illumination on the boxes and also reflect the direction of the light coming from the window, as seen in composited images 710 (row 3).[001 13] FIG. 8 depicts images generated by the example embodiment in comparison to images generated by the other methods described above, according to example embodiments. In particular, FIG. 8 depicts the naive composited images 802, generated by compositing objects onto the background scenes directly using masks 804 without modification. Composited images 810 were generated by the example embodiment, which can be compared to composited images 816, generated using the other methods described above. FIG. 8 also depicts albedo images 806 and shading images 808 of the composited images 810, generated by the example embodiment, as well as albedo images 812 and shading images 814 of the composited images 810, generated by the other methods. In particular, FIG. 8 provides a comparison of the example embodiment to the other methods that also model intrinsics, or perform non-global edits to simulate altered lighting. The method of Guo et al.
[0011] is also included as it may perform image harmonization in the intrinsic domain. It should be noted that these results are not a part of the user study as the generated composited images are only at a resolution of (256x256). As seen in FIG. 8, the other methods struggle to estimate accurate intrinsic representations. The approach of Bhattad [2] cannot generate high-frequency details in its shading due to estimating at low resolution, therefore its results typically don’t exhibit novel illumination effects as shown in shading images 814. Guo et al.
[0011] also attempts to perform intrinsic decomposition and harmonization jointly, but fails to estimatemeaningful albedo (e.g. albedo images 812) and shading (e.g. shading images 814) due to a lack of ground-truth supervision for these quantities. The method of Wang et al.
[0034] predicts a gain map to modulate parametric edits. While their gain map does allow them to model local variations, this method cannot fully relight the foreground as they do not explicitly model albedo. This results in their harmonized foreground maintaining the illumination conditions from its source environment, as seen in composited images 816. The example embodiment can model the image compositing problem with physical accuracy and can therefore predict a detailed and accurate reshading of the foreground, as seen in composited images 810.
[0114] In accordance with the present disclosure, systems and methods for performing illumination-aware image harmonization are disclosed. For example, the second neural network 208 can be used to faithfully generate a realistic shading layer that obeys the conditions of the adjusted shading image 226. Further, the second neural network 208 can also be modified and utilized for interactive relighting, as shown in FIG. 9. The illumination component or parameters 232, for example generated by the render engine 206, can be modified to modify the adjusted object shading image 226 and thereby the composite image 230 for processing by the second neural network 208. This aspect of the disclosed systems and methods is also amenable to interactive image editing applications. FIG. 9 depicts an example naive composite image 902, generated by compositing an object (i.e. vase) onto the background scenes directly without modification. An example embodiment is able to generate an adjusted shading image 906 and a corresponding illumination component 904, a corresponding adjusted composite shading image 904, as well as a corresponding composite image 910, as described above. In some embodiments, the user 102 can manually alter the lighting or illumination parameters by modifying the illumination component 904, for example in case the estimated lighting or illumination parameters is inaccurate. The user 102 can alter lighting parameters 912 of the illumination component 904 by providing values for the lighting parameters 912, such as direction, directional intensity, and ambient intensity, which would generate a new illumination component 920. Accordingly, by using the new illumination component 920 the example embodiment can generate updated shading images 914, 916 corresponding to the new adjusted shading image and the new adjusted compositeshading image, to yield an updated composite image 918 that is more preferable to the user 102. As such, when manual adjustments are combined with the albedo harmonization formulation described above, the user 102 can have full control over the image harmonization process.
[0115] In some embodiments, the input and generated images of FIGS. 3A-9 are processed at 1024-pixel resolution, which may be the largest resolution where the other image compositing methods provide consistent results. The disclosed systems and methods can generate enough details while still being globally coherent for the object at this resolution.
[0116] FIG. 10A depicts a method for performing image compositing using the system of FIGS. 1-2C, according to an example embodiment. Image compositing may comprise compositing an object, region orforeground depicted in an object image 120 onto a background scene depicted in a background image 122 to generate a composite image 124 such that the object has a natural appearance in the background scene. The images 120, 122 may be captured using an image capture device such as a smartphone or camera. The albedo and shading for the object and the background scene may be determined or obtained, for example through image decomposition 202 at 1002 to generate the object albedo image 210, object shading image 212, background albedo image 214, and background shading image 216. At 1004, the mask 220 can be generated or obtained. The mask 220 can be used to identify the region (i.e. selected pixels) to be occupied by the object in the background scene and may be used to image segmentation and compositing. For example, the mask 220 may be used to composite the object, shading(s) thereof, or albedo(s) thereof onto the background scene, shading thereof, or albedo thereof, respectively, and may be used to segment the object, shading(s) thereof, or albedo(s) from the object image, shading(s) thereof, or albedo(s) thereof as well as the background image, shading(s) thereof, or albedo(s).
[0117] By using the mask 220, the albedo of the object as depicted in the object albedo image 210 can be composited onto the albedo of the background as depicted in the background albedo image 216 for processing by the first neural network 204 trained to harmonize the albedo of the object to the albedo of the background sceneat 1006. In some embodiments, the first neural network 204 may take as input the mask 220, the object albedo image 210, and the background albedo image 216 rather than a composited albedo image. That is, the mask 220 can provide context for compositing the object albedo onto the background albedo to the first neural network 204. The first neural network 204 can output a harmonized albedo image 218 comprising the object albedo composited onto the background albedo and where the object albedo is harmonized to the background albedo as to match the color context of the object albedo to the color context of the background scene.
[0118] At 1008, surface normals for the object and the background scene can be obtained or generated, for example as surface normal maps, images, or vectors comprising object normal 222 and background normal 224. At 1010, the background shading image 216 and the background normal 224 may be processed by the render engine 206 to determine lighting parameters or conditions on the background scene as the illumination component 232, for example by rendering predicted shading images of the background using predicted illumination components to determine a illumination component 232 that renders a predicted background shading image that is the closest to the background shading image 216. At 1012, the lighting parameters of the background scene as defined in the illumination component 232 can be applied to the object to determine a shading of the object under the lighting parameters of the background scene, for example, based on the Lambertian assumption by rendering the object normal 222 with the illumination component 232 to generate an adjusted object shading image 226 corresponding to a predicted Lambertian shading of the object under the lighting conditions of the background. At 1014, the adjusted object shading image 226 can be composited onto the background shading image 216 to generate a composite shading image 228, for example using the mask 220.
[0119] At 1016, a predicted composite image 230 can be generated from the composite shading image 228 and the harmonized albedo image 218, for example by multiplying the images 218, 228. At 1018, depth information for the background scene may be generated or determined, for example as the depth map 234 or as depth image, vectors, or values. At 1020, the composite shading image 228, the predicted composite image 230, the mask 220, depth information (e.g. the depth map 234) andthe normals 222, 224 can be processed by the second neural network 208 to refine the shading of the composite image 124 to generate the adjusted composite shading image 234. In particular, the second neural network 208 may be trained to map the Lambertian shading of the object to non-Lambertian, accurate shading as present for the background scene. At 1022, the composite image 124 can be generated from the adjusted composite shading image 234 and the harmonized albedo image 218, for example by multiplying the images 218, 234.
[0120] FIG. 10B depicts a method for performing image compositing using the system of FIGS. 1-2C, according to an example embodiment. In particular, an image 128 may be relighted to generate the relighted image 130. At 1030, the illumination component 232 may be determined, corresponding to the target or desired lighting conditions for the relighted image 130. For general relighting, the illumination component 232 may be set to one or more target parameters. For compositing, the illumination component 232 may be set to the determined background scene illumination. At 1032, image decomposition 202 may be performed on the image 128 to generate the albedo image 240. At 1034, the geometry representations, such as surface normals 246 of the contents in the image 128 may be obtained or determined, which can be used to with the illumination component 232 to render the rendered shading image 248 at 1036, which depicting a predicted shading representative of the shading of the image under the altered lighting conditions defined by the illumination component 232, which may be based on the Lambertian assumption. At 1038, the input image 250 can be generated from the albedo image 240 and the rendered shading image 248, for example by multiplying the images 240, 248. At 1040, the shading from the rendered shading image 248 may be refined to determine an accurate (e.g. non-Lambertian) shading of the image 128 under the altered lighting conditions. In particular, the second neural network 208 may process the surface normals 246, the rendered shading image 248, and the input image 250 to generate the adjusted shading image 234 corresponding to the refined shading. At 1042, the relighted image 130 can be generated from the adjusted shading image 234 and the albedo image 240, for example by multiplying the images 234, 240. It should be noted that other orders of performing the method of FIG. 10B are possible as well.
[0121] It would be appreciated by one of ordinary skill in the art that the system and components shown in the figures may include components not shown in the drawings. For simplicity and clarity of the illustration, elements in the figures are not necessarily to scale and are only schematic. It will be apparent to persons skilled in the art that a number of variations and modifications can be made without departing from the scope of the invention as described herein.
[0122] It is contemplated that any part of any aspect or embodiment discussed in this specification can be implemented or combined with any part of any other aspect or embodiment discussed in this specification, so long as such those parts are not mutually exclusive with each other.
[0123] It should be recognized that features and aspects of the various examples provided above can be combined into further examples that also fall within the scope of the present disclosure.
[0124] When used in this specification and claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps or integers are included. The terms are not to be interpreted to exclude the presence of other features, steps or components. Additionally, the term "connect" and variants of it such as "connected", "connects", and "connecting" as used in this description are intended to include indirect and direct connections unless otherwise indicated. For example, if a first device is connected to a second device, that coupling may be through a direct connection or through an indirect connection via other devices and connections. Similarly, if the first device is communicatively connected to the second device, communication may be through a direct connection or through an indirect connection via other devices and connections. Further, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0125] The embodiments have been described above with reference to flow, sequence, and block diagrams of methods, apparatuses, systems, and computer program products. In this regard, the depicted flow, sequence, and block diagrams illustrate the architecture, functionality, and operation of implementations of variousembodiments. For instance, each block of the flow and block diagrams and operation in the sequence diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified action(s). In some alternative embodiments, the action(s) noted in that block or operation may occur out of the order noted in those figures. For example, two blocks or operations shown in succession may, in some embodiments, be executed substantially concurrently, or the blocks or operations may sometimes be executed in the reverse order, depending upon the functionality involved. Some specific examples of the foregoing have been noted above butthose noted examples are not necessarily the only examples. Each block of the flow and block diagrams and operation of the sequence diagrams, and combinations of those blocks and operations, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0126] Use of language such as "at least one of X, Y, and Z," "at least one of X, Y, or Z," "at least one or more of X, Y, and Z," "at least one or more of X, Y, and / or Z," or "at least one of X, Y, and / or Z," is intended to be inclusive of both a single item (e.g., just X, or just Y, or just Z) and multiple items (e.g., {X and Y}, {X and Z}, {Y and Z}, or {X, Y, and Z}). The phrase "at least one of" and similar phrases are not intended to convey a requirement that each possible item must be present, although each possible item may be present.
[0127] The invention may also broadly consist in the parts, elements, steps, examples and / or features referred to or indicated in the specification individually or collectively in any and all combinations of two or more said parts, elements, steps, examples and / or features. In particular, one or more features in any of the embodiments described herein may be combined with one or more features from any other embodiment(s) described herein.References[1] Zhongyun Bao, Chengjiang Long, Gang Fu, Daquan Liu, Yuanzhen Li, Jiaming Wu, and Chunxia Xiao. Deep image-based illumination harmonization. In Proc. CVPR, 2022.[2] Anand Bhattad and David A. Forsyth. Cut-and-paste object insertion by enabling deep image prior for reshading. Proc. 3DV, 2022.[3] Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3 / 4): 324-345, 1952.[4] Chris Careaga and Yagiz Aksoy. Intrinsic image decomposition via ordinal shading. SFU Tech. Rep., 2023.[5] Wenyan Cong, Xinhao Tao, Li Niu, Jing Liang, Xuesong Gao, Qihao Sun, and Liqing Zhang. High-resolution image harmonization via collaborative dual transformations. In Proc. CVPR, 2022.[6] Wenyan Cong, Jianfu Zhang, Li Niu, Liu, Zhixin Ling, Weiyuan Li, and Liqing Zhang. Dovenet: Deep image harmonization via domain verification. In Proc. CVPR, 2020.[7] Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In Proc. ICCV, 2021.[8] Mathieu Garon, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, and Jean-Francois Lalonde. Fast spatially-varying indoor lighting estimation. In Proc. CVPR, June 2019.[9] David Griffiths, Tobias Ritschel, and Julien Philip. Outcast: Single image relighting with cast shadows. Comput. Graph. Forum, 2022.
[0010] Zonghui Guo, Zhaorui Gu, Bing Zheng, Junyu Dong, and Haiyong Zheng. Transformer for image harmonization and beyond. IEEE Trans. Pattern Anal. Mach. Intell., 2022.
[0011] Zonghui Guo, Haiyong Zheng, Yufeng Jiang, Zhaorui Gu, and Bing Zheng. Intrinsic image harmonization. In Proc. CVPR, 2021.
[0012] Zhongyun Hu, Ntumba Elie Nsampi, Xue Wang, and Qing Wang. Neursf: Neural shading field for image harmonization, 2021 .
[0013] Kevin Karsch, Varsha Hedau, David Forsyth, and Derek Hoiem. Rendering synthetic objects into legacy photographs. ACM T rans. Graph. , 30(6), 2011.
[0014] Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, Hailin Jin, Rafael Fonte, Michael Sittig, and David Forsyth. Automatic scene inference for 3d object compositing. ACM Trans. Graph., 33(3), 2014.
[0015] Zhanghan Ke, Chunyi Sun, Lei Zhu, Ke Xu, and Rynson W.H. Lau. Harmonizer: Learning to perform white-box image and video harmonization. In Proc. ECCV, 2022.
[0016] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. 2015.
[0017] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. 2023.
[0018] Zicheng Liao, Kevin Karsch, Hongyi Zhang, and David Forsyth. An approximate shading model with detail decomposition for object relighting. Int. J. Comput. Vision, 127, 2019.
[0019] G. Lin, A. Milan, C. Shen, and I. Reid. RefineNet: Multi-path refinement networks for high-resolution semantic segmentation. In Proc. CVPR, 2017.
[0020] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proc. ECCV, 2014.
[0021] Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Inverse rendering for complex indoor scenes: Shape, spatially-varying lighting and SVBRDF from a single image. In Proc. CVPR, 2020.
[0022] Zhengqin Li, Jia Shi, Sai Bi, Rui Zhu, Kalyan Sunkavalli, Milos Hasan, Zexiang Xu, Ravi Ramamoorthi, and Manmohan Chandraker. Physically-based editing of indoor scene lighting from a single image. In Proc. ECCV, 2022.
[0023] Zhengqi Li and Noah Snavely. MegaDepth: Learning single-view depth prediction from internet photos. In Proc. CVPR, 2018.
[0024] Jorge Lopez-Moreno, Sunil Hadap, Erik Reinhard, and Diego Gutierrez. Compositing images through light source detection. Computers & Graphics, 34(6):698-707, 2010.
[0025] S. Mahdi H. Miangoleh, Zoya Bylinskii, Eric Kee, Eli Shechtman, and Yagiz Aksoy. Realistic saliency guided image enhancement. 2023.
[0026] S. Mahdi H. Miangoleh, Sebastian Dille, Long Mai, Sylvain Paris, and Yagiz Aksoy. Boosting monocular depth estimation models to high-resolution via content- adaptive multi-resolution merging. In Proc. CVPR, 2021.
[0027] Lukas Murmann, Michael Gharbi, Miika Aittala, and Fredo Durand. A multiillumination dataset of indoor object appearance. In Proc. ICCV, Oct 2019.
[0028] Baptiste Nicolet, Julien Philip, and George Drettakis. Repurposing a relighting network for realistic compositions of captured scenes. In Proceedings of the ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games, 2020.
[0029] Rohit Pandey, Sergio Orts-Escolano, Chloe LeGendre, Christian Haene, Sofien Bouaziz, Christoph Rhemann, Paul Debevec, and Sean Fanello. Total relighting: Learning to relight portraits for background replacement. ACM Trans. Graph., 2021.
[0030] F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine- Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proc. CVPR, 2016.
[0031] Julien Philip, Michael Gharbi, Tinghui Zhou, Alexei A. Efros, and George Drettakis. Multi-view relighting using a geometry-aware network. ACM Trans. Graph., 38(4), 2019.
[0032] Julien Philip, Sebastien Morgenthaler, Michael Gharbi, and George Drettakis. Free-viewpoint indoor neural relighting from multi-view stereo. ACM Trans. Graph., 2021.
[0033] Rene Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Trans. Pattern Anal. Mach. Intell., 2020.
[0034] Ke Wang, Michael Gharbi, He Zhang, Zhihao Xia, and Eli Shechtman. Semisupervised parametric real-world image harmonization. In Proc. CVPR, 2023.
[0035] Saining Xie, Ross Girshick, Piotr DollAjr, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proc. CVPR, 2017.
[0036] Ben Xue, Shenghui Ran, Quan Chen, Rongfei Jia, Binqiang Zhao, and Binqiang Zhao. Dccf: Deep comprehensible color filter learning framework for high-resolution image harmonization. In Proc. ECCV, 2022.
[0037] Yu-Ying Yeh, Koki Nagano, Sameh Khamis, Jan Kautz, Ming-Yu Liu, and Ting- Chun Wang. Learning to relight portrait images via a virtual light stage and synthetic- to-real adaptation. ACM Trans. Graph., 2022.
[0038] Jinsong Zhang, Kalyan Sunkavalli, Yannick Hold-Geoffroy, Sunil Hadap, Jonathan Eisenmann, and Jean-Frangois Lalonde. All-weather deep outdoor lighting estimation. In Proc. CVPR, 2019.
[0039] Yuanming Hu, Hao He, Chenxi Xu, Baoyuan Wang, and Stephen Lin. Exposure: A white-box photo post-processing framework. ACM Trans. Graph., 37(2):26, 2018.
Claims
CLAIMS:1 . An image processing method for relighting an image, comprising: generating a shading image representative of shading for the image using an illumination component defining one or more lighting parameters for relighting the image and a geometry representation of the image; generating an input image using the shading image and an albedo image representative of albedo in the image; and processing the input image, the albedo image, and the shading image with a first neural network trained to refine the shading to generate a relighted image corresponding to the illumination component.
2. The method of claim 1 , wherein generating the relighted images comprises: generating a refined shading image representative of the shading according to the illumination component using the first neural network; and generating the relighted image using the refined shading image and the albedo image.
3. The method of claim 2, wherein the geometry representation is: a normal map defining surface normals of the image, a depth map, a point cloud, a set of 3D points, Gaussian splats in 3D space, a 3D mesh representation, or a 3D representation of the image; and wherein generating the shading image comprises rendering the shading image using the illumination component and the geometry representation of the image.
4. The method of claim 3,wherein the illumination component is representative of lighting from a light source; and wherein the one or more lighting parameters comprises: a direction of lighting, a directional intensity of lighting, an ambient intensity of lighting, a 3D location of a light source, a 2D representation of a light source, or combinations thereof.
5. The method of claim 4, wherein the illumination component comprises a constant ambient illumination and / or illumination from the light source; wherein the geometry representation of the image and the illumination from the light source are 3 dimensional vectors and the constant ambient illumination is a constant value; and wherein the shading image is a sum of the constant ambient illumination and a dot product of the illumination from the light source and the geometry representation of the image.
6. The method of claim 5, where the illumination component comprises a plurality of subcomponents, each subcomponent representative of lighting from a respective light source in 3D space; and wherein the shading depicted in the shading image corresponds to a plurality of light sources defined by the plurality of subcomponents.
7. The method of claim 2, wherein the first neural network is trained using a dataset comprising photographs, shading images of the photographs, training shading images, and training input images, wherein the shading images of the photographs are generated using image decomposition;wherein the training shading images are generated by: determining illumination components of the photographs by rendering shading estimations using geometry representations of the photographs; and rendering the training shading images using the illumination components and the geometry representations of the photographs; wherein the training input images are generated from the training shading images and albedo images of the photographs; and wherein the first neural network is trained to refine the shading by training the first neural network to generate the shading images of the photographs from the training shading images, the training input images, or both.
8. The method of claim 2, wherein the first neural network is trained using a loss function comprising mean squared error of the refined shading image, mean squared error of the relighted image, multi-scale gradient loss of the refined shading image, multi-scale gradient loss of the relighted image, or combinations thereof.
9. The method of claim 2, further comprising: generating the shading image and the albedo image by performing image decomposition on the image.
10. The method of claim 9, wherein the image decomposition generates: the shading image as a diffuse shading image, the albedo image as a diffuse albedo image, and a residual image representative of specularities in the image; the method further comprising: processing the relighted image and the refined shading image with a second neural network trained to estimate the specularities to generate an adjusted residual image representative of the specularities according to the illumination component; andgenerating a second relighted image using the refined residual image and the relighted image.
11. The method of claim 10, wherein the second neural network is trained using a dataset comprising photographs, residual images of the photographs, training diffuse images, and training diffuse shading images, wherein the training diffuse shading images are generated by: determining illumination components of the photographs; rendering predicted shading images using the illumination components and geometry representations of the photographs; and generating training diffuse shading images from the predicted shading images; wherein the training diffuse images are generated from the training diffuse shading images and diffuse albedo images of the photographs; and wherein the second neural network is trained to estimate the specularities by training the second neural network to generate the residual images of the photographs from the training diffuse shading images and the training diffuse images.
12. The method of claim 2 for use in compositing an object depicted in the image onto a background image depicting a background scene, the method further comprising: processing a background albedo image of the background image and an object albedo image of the object using a third neural network trained to match a foreground albedo to a background albedo to generate a composite albedo image as the albedo image, the composite albedo image comprising the object albedo image composited onto the background albedo image, wherein analbedo of the object albedo image is matched to an albedo of the background albedo image by the second neural network; processing the background shading image using a render engine to perform illumination estimations to determine an illumination component of the background scene as the illumination component; and processing the shading image by compositing the shading image as an object shading image onto a background shading image of the background image prior to generating the input image; wherein the relighted image is a composite image depicting the object in the background scene.
13. The method of claim 12, further comprising: generating the background albedo image and the background shading image by performing image decomposition on the background with a fourth neural network trained to perform intrinsic image decomposition; and generating the object albedo image by performing image decomposition on the object image with the fourth neural network trained to perform intrinsic image decomposition.
14. The method of claim 12, further comprising: generating a mask corresponding to a region occupied by the object in the background scene for compositing the object onto the background image; and applying the mask to composite images comprising the object onto images of the background scene.
15. The method of claim 12, wherein the third neural network is trained using a segmentation dataset comprising albedo images, foreground albedo images, and background albedo images, wherein the albedo images are segmented to generate the foregroundalbedo images and the background albedo images, and wherein colors of the foreground albedo images are transformed according to one or more image editing parameters; and wherein the third neural network is trained to match the foreground albedo to the background albedo by training the third neural network to generate the albedo images from the foreground albedo images and background albedo images.
16. The method of claim 12, wherein the first neural network is trained using a segmentation dataset comprising shading images of photographs, training shading images, and training input images; wherein the training shading images are generated by: segmenting foreground shading images from the shading images; determining illumination components of the foreground shading images by rendering shading estimations of the foreground shading images or the photographs using geometry representations of the foreground shading images or the photographs; generating second foreground shading images from the illumination components and geometry representations of the photographs; and generating the training shading images by compositing the second foreground shading images onto the shading images of the photographs; wherein the training input images are generated from the training shading images and albedo images of the photographs; and wherein the first neural network is trained to refine the shading by training the first neural network to generate the shading images of the photographs from the training shading images and the training input images.
17. The method of claim 12, further comprising: determining geometry representations of the background scene; and processing the geometry representations of the background scene using the render engine to determine the illumination component of the background shading image; wherein the render engine is configured to determine the illumination component by rendering estimated shading images using estimated illumination components.
18. The method of claim 12, wherein the illumination component of the background shading image comprises a constant ambient illumination and / or illumination from a light source wherein the geometry representations of the background scene and the illumination from the light source are 3 dimensional vectors and the constant ambient illumination is a constant value; and wherein the background shading image is a sum of the constant ambient illumination and a dot product of the illumination from the light source and the geometry representations of the background scene.
19. The method of claim 12, further comprising: generating a mask for compositing the object onto the background image; generating a depth map of the background scene from the background image; determining geometry representations of the background scene; determining geometry representations of the object; andprocessing the mask, the depth map, the geometry representations of the background scene, and the geometry representations of the object with the first neural network to generate the refined shading image.
20. The method of claim 12, wherein the render engine is configured to determine the illumination component using gradient-based least square optimization to minimize a difference between the background shading image and an estimated background shading based on the illumination component.
21. A method of training neural networks for performing image processing comprising: obtaining shading images of photographs; generating training shading images by: determining illumination components of the photographs; and rendering the training shading images using the illumination components and geometry representations of the photographs; generating training input images from the training shading images and albedo images of the photographs; and training a first neural network to refine shading of an image by training the first neural network to generate the shading images of the photographs from the training shading images and the training input images.
22. The method of claim 21 , wherein the shading images are diffuse shading images and wherein the albedo images are diffuse albedo images, the method further comprising: obtaining residual images of the photographs;generating training diffuse shading images from the training shading images using the first neural network; generating the training diffuse images from the training diffuse shading images and the albedo images of the photographs; and training a second neural network to estimate specularities in the image by training the second neural network to generate the residual images of the photographs from the training diffuse shading images and the training diffuse images.
23. The method of claim 22, further comprising: generating foreground albedo images and background albedo images by segmenting the albedo images; generating foreground shading images and background shading images by segmenting the shading images; transforming colors of the foreground albedo images according to one or more image editing parameters to generate transformed foreground albedo images; and training a third neural network match a foreground albedo to a background albedo by training the third neural network to generate the albedo images using the transformed foreground albedo images and the background albedo images; wherein the third neural network is trained to generate a composite albedo image; wherein a render engine is configured to perform illumination estimations to determine an illumination component for generating an object shading image representative of shading of an object for compositing onto a background scene, the object shading image for compositing onto a background shadingimage representative of shading of the object under the illumination in the background scene for generating a composite shading image; wherein the composite shading image is for generating a composite image; wherein the first neural network is trained to generate a refined shading image from the composite shading image and the composite image; and wherein the refined shading image is for generating a relighted image with the composite albedo image.
24. An image processing method for compositing an object onto a background image depicting a background scene, the method comprising: obtaining a background albedo image and a background shading image of the background image; obtaining an object albedo image of the object; processing the background albedo image and the object albedo image with a first neural network trained to match a foreground albedo to a background albedo to generate a composite albedo image, the composite albedo image comprising the object albedo image composited onto the background albedo image, wherein an albedo of the object albedo image is matched to an albedo of the background albedo image by the first neural network; processing the background shading image using a render engine configured to perform illumination estimations to determine an illumination component of the background shading image corresponding to illumination in the background scene; generating, according to the illumination component, an object shading image representative of shading of the object under the illumination in the background scene; generating a first composite shading image by compositing the object shading image onto the background shading image;generating a first composite image from the first composite shading image and the composite albedo image; processing the first composite image and the first composite shading image with a second neural network trained to determine a shading for a composited image to generate a second composite shading image; and generating a second composite image from the second composite shading image and the composite albedo image.
25. An image processing method for compositing a light emitting object onto a background image depicting a background scene, the method comprising: obtaining a background albedo image and a background shading image of the background image; obtaining an object albedo image of the light emitting object; generating, according to the object albedo image and the background albedo image, a composite albedo image; processing the background shading image using a render engine configured to perform illumination estimations to determine an illumination component of the background shading image corresponding to illumination in the background scene; determining an augmented illumination component by augmenting the illumination component of the background shading image with an illumination component of the light emitting object; generating, according to the augmented illumination component, a first composite shading image comprising shading of the light emitting object and shading of the background scene representative of illumination from the augmented illumination component; generating a first composite image from the first composite shading image and the composite albedo image;processing the first composite image and the first composite shading image with a second neural network trained to determine a shading for a composited image to generate a second composite shading image; and generating a second composite image from the second composite shading image and the composite albedo image.
26. A system comprising one or more processing units configured to perform the method of any one of claims 1 to 25.
27. A non-transitory computer-readable medium having computer readable instructions stored thereon, which, when executed by one or more processing units, causes the one or more processing units to perform the method of any one of claims1 to 25.
Citation Information
Patent Citations
Dynamically estimating lighting parameters for positions within augmented-reality scenes based on global and local features
US10665011B1
End-to-end relighting of a foreground object of an image
US20210295571A1
Inserting three-dimensional objects into digital images with consistent lighting via global and local lighting information
US20230037591A1
Marking-based portrait relighting
US20240404188A1
Photo relighting and background replacement based on machine learning models
WO2022231582A1