Methods, storage media, and systems for controlling color and lighting conditions of generative machine learning models

By inserting a calibration object to derive lighting conditions and controlling GML models, the technique addresses intra-image inconsistencies, ensuring accurate and realistic visual outputs in applications such as virtual reality and film production.

WO2025166130A1PCT designated stage Publication Date: 2025-08-07HOVER INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-30
Filing Date
2025-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Traditional generative machine learning (GML) techniques often result in inconsistent color and lighting conditions between original and new visual data, leading to diminished visual realism in applications like building image visualization, virtual reality, and film production.

Method used

Insert a simulated calibration object into an image scene, derive lighting conditions based on the object, and control the GML model using these conditions to generate new visual data with consistent color and lighting.

Benefits of technology

This approach ensures accurate representation of color and lighting conditions, enhancing realism and reducing biases, thus improving the quality of generated and rendered images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025013985_07082025_PF_FP_ABST
    Figure US2025013985_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, storage media, and systems for constraining color and lighting conditions of generative machine learning (GML) models are disclosed. An image is provided, a simulated calibration object is inserted into a portion of the scene in the image, and color and lighting conditions of the scene in the image are derived based on the at least one simulated calibration object. The image is input into a GML model, which is constrained based on the derived color and lighting conditions. A visualization including new visual data is generated using the constrained GML model.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS, STORAGE MEDIA, AND SYSTEMS FOR CONTROLLING COLOR AND LIGHTING CONDITIONS OF GENERATIVE MACHINE LEARNING MODELSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to provisional patent application number 63 / 627,735 entitled ‘‘METHODS, STORAGE MEDIA, AND SYSTEMS FOR DERIVING ARTIFICIAL COLOR AND LIGHTING CONDITIONS” filed January' 31, 2024, provisional patent application number 63 / 712.004 entitled “METHODS, STORAGE MEDIA, AND SYSTEMS FOR CONSTRAINING COLOR AND LIGHTING CONDITIONS OF GENERATIVE MACHINE LEARNING MODELS” filed October 25, 2024, and provisional patent application number 63 / 751,745 entitled “METHODS, STORAGE MEDIA, AND SYSTEMS FOR CONTROLLING COLOR AND LIGHTING CONDITIONS OF GENERATIVE MACHINE LEARNING MODELS” filed January 30, 2025, which are hereby incorporated by reference in their entirety and made part of the present application for all purposes.BACKGROUNDFIELD OF DISCLOSURE

[0002] The present disclosure generally relates to the technical field of image processing and, more specifically, to controlling color and lighting conditions in generative machine learning models for rendering visual data.BRIEF SUMMARY OF RELATED ART

[0003] The technical challenges in image processing, computer graphics, generative machine learning (GML), or a combination thereof, related to intra-image color and lighting inconsistencies are significant and multifaceted. In the realm of image processing and computer graphics, visualization of scenes often involves the rendering of visual data that can be affected by various color and lighting conditions. Traditional GML techniques have been employed to predict and generate new visual data (sometimes referred to as “output inference visual data”). These predictions are typically informed by original visual data (sometimes referred to as "previously provided visual data" or "input inference visual data") and training data (the data used to train the GML techniques), which may include certain inherent color and lightingconditions.

[0004] Challenges arise when the color and lighting conditions of the original visual data are not taken into account when predicting and generating the new visual data. This can lead to the new visual data having color and lighting conditions different than that of the original visual data. As a result, the resultant color and lighting conditions of a rendered image including both the original visual data and the new visual data may be inconsistent, which can significantly impact the fi delity and authenticity of the visual output.

[0005] The ability7to accurately predict and generate visual data is crucial in various applications, including but not limited to, building image visualization, virtual reality, augmented reality, film production, video games, and the like. In these applications, the realism and accuracy of the generated scenes are paramount for creating an immersive and believable experience. Therefore, there is a need for improved techniques that can address the challenges associated with intra-image color and lighting inconsistencies in GML outputs including original visual data and new visual data.SUMMARY

[0006] In some embodiments, the present disclosure pertains to the field of image processing, particularly to techniques for controlling color and lighting conditions in generative machine learning (GML) models for rendering visual data. This field is instrumental in enhancing the realism of visual outputs in various applications such as building image visualization, virtual reality, augmented reality, film production, video games, and the like.

[0007] In some embodiments, a problem in image processing is a rendering of a GML technique having inconsistent color and lighting conditions. Traditional GML techniques have been utilized to predict and generate new visual data (sometimes referred to as “output inference visual data7’), but these techniques often modify color and lighting conditions of the original visual data (sometimes referred to as “previously provided visual data” or “input inference visual data”) when predicting and generating the new visual data and / or when rendering an image including both the original visual data and the new visual data. Consequently, this oversight can result in the new visual data, the rendered image, or both, having intra-image color and lighting condition inconsistencies, specifically between the color and lighting conditions of the original visual data and the new visual data, which maymisrepresent the underlying objects in the image and in turn diminish the visual realism.

[0008] In some embodiments, this problem may be addressed through manual adjustments of color and lighting in post-processing or use of generic lighting models that attempt to characterize the color and lighting conditions in a generic manner. However, these methods are time-consuming, often require expert intervention, and may not accurately capture the complex interplay of color and lighting of objects within a scene, especially when dealing with a dynamic or intricate color and lighting environment including a plurality of objects.

[0009] In some embodiments, this problem may be addressed by using machine learning models to predict and apply color and lighting conditions based on training data. However, this method may struggle with scenes not found in the training data and may introduce biases from the training data.

[0010] One aspect of the disclosure relates to methods, storage media, and systems controlling color and lighting conditions of a generative machine learning model. They may include providing an image, inserting at least one simulated calibration object into a portion of a scene in the image, deriving a lighting condition of the scene in the image based on the at least one simulated calibration object, providing the image as input to a generative machine learning model, controlling the generative machine learning model based on the derived lighting condition, and generating, using the controlled generative machine learning model, a first visualization comprising new visual data.

[0011] One aspect of the disclosure relates to methods, storage media, and systems controlling color and lighting conditions of a generative machine learning model. They may include providing an image, inserting at least one simulated calibration object into a portion of a scene in the image, deriving a lighting condition of the scene in the image based on the at least one simulated calibration object, generating a view of a 3D object within a 3D environment comprising a representation of the derived lighting condition, providing the image and the view of the 3D object as inputs to a generative machine learning model, and generating, using the generative machine learning model, a first visualization comprising new visual data.

[0012] In some embodiments, the present disclosure pertains to the field of image processing, particularly to techniques for deriving lighting conditions in rendered images. This field is instrumental in enhancing the realism of visual outputs in various applications such as building image visualization, virtual reality, augmented reality, film production, video games, and the like.

[0013] In some embodiments, a problem in image processing is an output of an imagerendering technique having inconsistent color and lighting conditions. Traditional generative machine learning (GML) techniques have been utilized to predict and generate images that are then rendered, but these methods often overlook unique color and lighting conditions of the original data. Consequently, this oversight can result in generated and / or rendered images with inconsistent color and lighting conditions, which may diminish the visual realism.

[0014] Previous approaches to solving this problem have included manual adjustments of color and lighting in post-processing or use of generic lighting models that attempt characterize the color and lighting conditions of the output in a generic manner. However, these methods are time-consuming, often require expert intervention, and may not accurately capture the complex interplay of color and light within a scene, especially when dealing with dynamic or intricate color and lighting environments.

[0015] The problem of deriving underlying scene information subject to an inconsistent resultant lighting condition of an output of an image rendering technique is solved by detecting a change between a calibration object in an input of the image rendering technique and a modified version of the calibration object in the output of the image rendering technique. The calibration object is inserted into the input. The calibration object is not subject to the original lighting conditions of the input. The image rendering technique renders a new visualization of the input including the inserted calibration object. The output includes a modified version of the calibration object. A change can be detected between the calibration object in the input and the modified version of the calibration object in the output. Actual color and lighting conditions of the output can be derived based on the detected change. This process allows for a more accurate representation of the color and lighting of the output. In some embodiments, by implementing this technique, image processing techniques such as GLM can generate and / or render new visual data and use the calibration object's change to derive color and lighting conditions of the new visual data and even additional visual data that may be added to the new visual data.

[0016] Potential applications of this technique are vast, ranging from improving the visual quality of generated and / or rendered images in entertainment and media.

[0017] One aspect of the disclosure relates to methods, storage media, and systems for deriving color and lighting conditions. They may include providing an image, inserting at least one simulated calibration object into a portion of the image, rendering a new visualization of the image including a modified version of the at least one simulated calibration object, detecting a change between the at least one simulated calibration object in the image and themodified version of the at least one simulated calibration object in the new visualization, and deriving a color and lighting condition of the new visualization based on the detected change.

[0018] These and other features, and characteristics of the present technology, as well as the methods of operation and functions of the related elements of structure and the combination of parts and economies of manufacture, will become more apparent upon consideration of the following description and the appended claims with reference to the accompanying drawings, all of which form a part of this specification, wherein like reference numerals designate corresponding parts in the various figures. It is to be expressly understood, however, that the drawings are for the purpose of illustration and description only and are not intended as a definition of the limits of the invention. As used in the specification and in the claims, the singular form of 'a', 'an', and 'the' include plural referents unless the context clearly dictates otherwise.DRAWINGS

[0019] Figure (FIG.) 1 illustrates a system configured for controlling color and lighting conditions of a generative machine learning (GML) model, according to some embodiments.

[0020] FIG. 2 illustrates a method for controlling color and lighting conditions of a GML model, according to some embodiments.

[0021] FIG. 3A illustrates an image, according to some embodiments.

[0022] FIG. 3B illustrates a simulated mirror sphere inserted into a portion of an image, according to some embodiments.

[0023] FIG. 3C illustrates a visualization including a motorcycle, according to some embodiments.

[0024] FIG. 4 illustrates a system configured for deriving color and lighting conditions, according to some embodiments.

[0025] FIG. 5 illustrates a method for deriving color and light conditions, according to some embodiments.

[0026] FIG. 6A illustrates an image, according to some embodiments.

[0027] FIG. 6B illustrates an image including an empty crop portion, according to some embodiments.

[0028] FIG. 6C illustrates an image including a mirror sphere in a portion of an empty cropportion, according to some embodiments.

[0029] FIG. 6D illustrates a new visualization of an image including a modified mirror sphere, according to some embodiments.

[0030] FIG. 6E illustrates an adjusted new visualization including a car, according to some embodiments.DETAILED DESCRIPTION

[0031] The problem of intra-image color and lighting condition inconsistencies in a generative machine learning (GML) output including original visual data and new visual data is solved by controlling a GML model based on derived color and lighting conditions of the original visual data. The disclosure includes details related to the technical problems in the field of image processing, computer graphics, GML, or a combination thereof, specifically, the problem of intra-image color and lighting condition inconsistencies in a GML output including original visual data and new visual data. The disclosure includes details related to the technical solutions of controlling the GML model based on derived color and lighting conditions of the original visual data. The disclosure improves the output of the GML models resulting in more consistent visualizations. The disclosure includes details of controlling the GML models in an unconventional way, specifically by controlling the GML models based on color and lighting conditions of the original visual data. This is not well-understood, routine, or conventional activity in the field of image processing, computer graphics, GML, or a combination thereof.

[0032] The present disclosure addresses the aforementioned and related challenges by introducing methods, storage media, and systems for controlling color and lighting conditions of a GML model. An image is provided. At least one simulated calibration object is inserted into a portion of the scene in the image. Color and lighting conditions are derived based on the at least one simulated calibration object inserted into the portion of the scene in the image. The image is provided as input to a GML model. The GML model is controlled based on the derived color and lighting conditions. A visualization is generated using the controlled GML model, where the visualization includes new visual data.

[0033] In some embodiments, an image of a building structure may be provided as input to a GML model and a new visualization may be generated using the GML model. In some embodiments, the new visualization may include a new object and the image.

[0034] Absent controlling the GML model based on the derived color and lighting conditions of the scene in the image, the new object and the scene in the image may haveinconsistent color and lighting conditions, and the new object may be inaccurately depicted with respect to the scene in the image. In other words, the color and lighting values of the new object may not match the color and lighting values of the scene in the image, resulting in visual inconsistency.

[0035] Controlling the GML model based on the derived color and lighting conditions improve the outputs of the GML model. When controlling the GML model based on the derived color and lighting conditions, the new object and the scene in the image have consistent color and lighting conditions, and the new object may be accurately depicted with respect to the scene in the image. In other words, the color and lighting values of the new object may match the color and lighting values of the scene in the image, resulting in visual consistency. This allows for and results in a more accurate representation of the color and lighting of the new object with respect to the scene in the image.

[0036] The controlled GML model produces more consistent visualizations by using derived color and lighting conditions from the original image based on the simulated calibration object. This approach ensures that new visual data maintains consistent color and lighting conditions with respect to the original image, addressing the problem of intra-image color and lighting inconsistencies.

[0037] The controlled GML model can adapt to scenes that may not be present in the training data, overcoming limitations of existing machine learning models that struggle with such scenarios. This adaptability is achieved through the use of control networks that modulate the GML models based on the derived color and lighting conditions.

[0038] Unlike some existing machine learning solutions that may introduce color and lighting biases from training data, the disclosed approaches of deriving color and lighting conditions of the original image helps reduce such biases, resulting in more consistent representations.

[0039] By automating the processes of deriving color and lighting conditions and controlling the GML model based on the derived color and lighting conditions, expert intervention may be minimized or eliminated, which is often required in traditional manual adjustment methods. This automation leads to more efficient and consistent results.

[0040] The ability to accurately predict and generate visual data is crucial in various applications, including but not limited to, building image visualization, virtual reality, augmented reality, mixed reality , film production, video games, and the like.

[0041] In building image visualization, such as architectural design and real estate, thedisclosure can significantly improve the accuracy of virtual staging and property visualization. When adding new furniture or decor to an existing room image, the controlled GML model ensures that the color and lighting of the added elements match the existing space. This results in more realistic and convincing visualizations, helping users better envision the final product.

[0042] For virtual reality (VR) applications, such as gaming or training simulations, the disclosure can enhance the realism of dynamically generated content. As users move through a VR space, new elements can be generated on-the-fly with color and lighting conditions that match the existing virtual world. This seamless integration of new visual data eliminates jarring transitions or inconsistencies that could break immersion, leading to a more believable and effective VR experience.

[0043] For augmented reality (AR) applications, such as in retail or education, the disclosure can improve the integration of virtual objects into real -world environments. When placing a virtual product in a real space, the disclosure ensures that the color and lighting of the virtual object match the real-world color and lighting conditions. This results in more convincing AR experiences, leading to a more believable and effective AR experience.

[0044] In game development, particularly for open-world games with dynamic time-of-day and weather systems, the disclosure can enhance the realism of procedurally generated content. As the game generates new areas or objects based on player actions, the disclosure ensures that these elements have consistent color and lighting with the surrounding environment. This leads to more immersive and visually coherent gaming experiences.

[0045] FIG. 1 illustrates a system 100 configured for controlling color and lighting conditions of a GML model, in accordance with one or more implementations. In some implementations, system 100 may include one or more computing platforms 102. Computing platform(s) 102 may be configured to communicate with one or more remote platforms 104 according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. Remote platform(s) 104 may be configured to communicate with other remote platforms via computing platform(s) 102 and / or according to a client / server architecture, a peer-to-peer architecture, and / or other architectures.

[0046] Users may access system 100 via remote platform(s) 104. Computing platform(s) 102 may be configured by machine-readable instructions 106. Machine-readable instructions 106 may include one or more instruction modules. The instruction modules may include computer program modules. The instruction modules may include one or more of image module 108, object insertion module 1 10, color and lighting conditions derivation module 112,generative machine learning module 114, control module 116, visualization generation module 118, and / or other instruction modules.

[0047] Image module 108 may be configured to provide, receive, capture, or otherwise obtain an image. In some embodiments, the image is a real-world image captured by a real- world camera. In some embodiments, the image is a synthetic image generated by a model, such as a GML model. In some embodiments, the image may include a building structure, which may be an interior or exterior view of the building.

[0048] Object insertion module 110 may be configured to insert at least one simulated calibration object into a portion of the scene in the image. In some embodiments, object insertion module 110 replaces pixels of the image with the at least one simulated calibration object. In some embodiments, the at least one simulated calibration object includes a chromatic object, such as a mirror sphere, or a matte object, such as a gray sphere. In some embodiments, the at least one simulated calibration object is a color chart. The simulated calibration object may be a three-dimensional (3D) object inserted in two dimensions of the image. The flexibility in the type of simulated calibration object and the techniques for inserting the simulated calibration object allows this technique to be adaptable and surpass capabilities of more rigid methods where the type of calibration object and the insertion technique may be fixed.

[0049] In some embodiments, object insertion module 110 may be configured to insert the at least one simulated calibration object with respect to 2D space. In these embodiments, the image from image module 108 and an inpaint mask may be provided as input to GML module 114. In some embodiments, GML module 114 may comprise one or more GML models. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control netw orks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. Stable diffusion models are a type of deep learning model known for their ability to generate high-quality images while maintaining coherence and stability in the output. The one or more control networks may use the input mask to adjust parameters or outputs of the one or more GML models, ensuring that the generated visual data inserts the at least one simulated calibration object according to the inpaint mask. In these embodiments, the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image.

[0050] In some embodiments, object insertion module 110 may be configured to insert theat least one simulated calibration object with respect to 2D and 3D space. In these embodiments, the image from image module 108, an inpaint mask, and a depth cue may be provided as input to GML module 114. In some embodiments, GML module 114 may comprise one or more GML models. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors. In some embodiments, the inpaint mask may be superimposed with the depth cue to generate a depth cue with an inpaint mask. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the depth cue with the inpaint mask to adjust parameters or outputs of the one or more GML models, ensuring that the generated visual data inserts the at least one simulated calibration object according to the depth cue with the inpaint mask. In some embodiments, the one or more control networks may be one or more depth conditioned control networks.

[0051] In some embodiments, the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image and for portions of the scene not observed by the camera associated with the image. For the portions of the scene that are observed by the camera associated with the image, the at least one simulated calibration object may be generated based on the depth cue and portions of the image that are in front of the at least one simulated calibration object in 3D space. For the portions of the scene not observed by the camera associated with the image, the at least one simulated calibration object may be generated based on the depth cue and portions of the scene in the image that are behind the at least one simulated calibration object in 3D space.

[0052] In some embodiments, object insertion module 110 may be configured to extend a boundary portion of the image to create an empty crop portion. In these embodiments, the object insertion module 110 may be configured to insert the at least one simulated calibration object into the empty crop portion, or a portion thereof.

[0053] In some embodiments, object insertion module 110 may be configured to insert the at least one simulated calibration object into the portion of the scene in the image using one or more GML techniques such as, for example, inpainting.

[0054] In some embodiments, object insertion module 110 may be configured to insert a simulated calibration object into a portion of scene in the image. In these embodiments, thesimulated calibration object may be used to derive global color conditions, lighting conditions, or both, of the scene in the image as a whole, for example by color and light conditions derivation module 112. In these embodiments, while the derived global color conditions, lighting conditions, or both, may be representative of the scene in the image as a whole, they may be biased to the portion of the scene in the image the simulated calibration object is inserted into. In some embodiments, the derived global color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in the scene in the image according to the derived global color conditions, lighting conditions, or both, for the scene in the image, for example by visualization generation module 118.

[0055] In some embodiments, object insertion module 110 may be configured to insert a plurality7of simulated calibration objects into various portions of the scene in the image. For example, the image may be divided into a grid and a simulated calibration object may be inserted into each cell of the grid. At the most granular level, a simulated calibration object may be inserted at each pixel location of the image. In these embodiments, the plurality7of simulated calibration object may be used to derive local color, conditions, lighting conditions, or both, of various portions of the scene in the image, for example by color and light conditions derivation module 112. In some embodiments, the derived local color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in one or more portions of the scene in the image according to the derived local color conditions, lighting conditions, or both, for the one or more portions of the scene in the image, for example by visualization generation module 118.

[0056] In some embodiments, object insertion module 110 may be configured to insert at least one simulated calibration object into a portion of the scene in the image according to the image. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and Y of the image. In some embodiments, obj ect insertion module 110 may be configured to insert at least one simulated calibration object into a portion of the scene in the image according to depth cues. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and the Y of the image and the Z of the depth cues. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors.

[0057] In some embodiments, object insertion module 110 may be configured to perform operations a plurality of times and a representation (e.g., average) may be used.

[0058] Color and lighting conditions derivation module 112 may be configured to derivecolor conditions, lighting conditions, or both, of the scene in the image based on the at least one simulated calibration object. The derived color and lighting conditions quantify and represent the color and lighting conditions of the scene in the image. The derived color and lighting conditions may estimate the 3D lighting of the scene depicted in the image. Derived color and lighting conditions may be an image-based and may be encoded into derived color and lighting conditions in feature-space. The derived color and lighting conditions in featurespace may be represented as a vector.

[0059] In some embodiments, color and lighting conditions derivation module 112 may be configured to unwrap the at least one simulated calibration object, such as a mirror sphere or a gray sphere, to generate an equirectangular map (sometimes referred to as an environment map), and derive the color conditions, the lighting conditions, or both, of the scene in the image based on the equirectangular map. In some embodiments, color and lighting conditions derivation module 112 may be configured to apply a high dynamic range image map onto the at least one simulated calibration object, such as a mirror sphere or a gray sphere, and derive the color conditions, the lighting conditions, or both based on the high dynamic range image map. In some embodiments, color and lighting conditions derivation module 112 may be configured to map the at least one simulated calibration obj ect. such as a color chart, to a known calibration object, such as a known color chart, and derive the color conditions, the lighting conditions, or both, based on the mapping. While specific techniques for deriving the color conditions, the lighting conditions, or both, are disclosed herein, one of ordinary' skill in the art may appreciate other techniques, such as those used in visual effects, may be used.

[0060] The derived color conditions may include various parameters such as camera colors. These parameters provide a comprehensive representation of the color environment in the image. The derived lighting conditions may include various parameters such as light directionality, light intensify, and light temperature. These parameters provide a comprehensive representation of the lighting environment in the image.

[0061] In some embodiments, color and light derivation module 112 may be configured to derive color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image. In other words, color and light derivation module 1 12 may be configured to derive color conditions, lighting conditions, or both, of a front half of the scene in the image where the front half of the scene is the portion visible by the camera associated with the image. In some embodiments, color and light derivation module 112 may be configured to derive color conditions, lighting conditions, or both, of the front half of the scene in the image based on a 3D position of the at least onesimulated calibration object in the scene in the image and scene data that is substantially in front of the at least one simulated calibration object. In these embodiments, portions of the equirectangular map that correspond to portions of the scene not observed by the camera associated with the image may be filled in, for example using interpolation, based on portions of the equirectangular map that correspond to portions of the scene observed by the camera associated with the image.

[0062] In some embodiments, operation 206 may include deriving color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by the camera associated with the image and for portions of the scene not observed by the camera associated with the image. In other words, operation 206 may include deriving color conditions, lighting conditions, or both, of a front half and a back half of the scene in the image where the front half of the scene is the portion visible by the camera associated with the image and where the back half of the scene is the portion not visible by the camera. In some embodiments, color and light derivation module 112 may be configured to derive color conditions, lighting conditions, or both, of the front half and the back half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of and substantially behind the at least one simulated calibration object. In these embodiments, the equirectangular map may correspond to the entire scene, or substantially the entire scene.

[0063] Color and lighting conditions derivation module 112 may use the at least one simulated calibration object as a reference point to represent the color and lighting conditions of the scene in the image, for example rather than relying on pre-existing lighting models, which may be generic, and rather than manual adjustments, which may require expert intervention and may be time consuming.

[0064] Object insertion module 110 inserts the at least one simulated calibration object into the portion of the scene in the image and color and lighting conditions derivation module 112 derives a color and lighting conditions of the scene in the image based on the at least one simulated calibration object. Object insertion module 110 introduces an element (i.e., the at least one simulated calibration object) into the image and color and lighting condition derivation module 112 relies on the at least one simulated calibration object rather than solely the image.

[0065] Generative machine learning (GML) module 114 may be configured to implement one or more GML models. The image obtained by image module 108 may be provided as inputinto the one or more GML models. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by one or more control networks. Control networks may use the input derived color and lighting conditions to adjust parameters or outputs of the GML model, ensuring that the generated visual data aligns with the image’s color and lighting environment.

[0066] Control module 116 may be configured to control the one or more GML models based on the derived color and lighting conditions. The derived color and lighting conditions from color and lighting conditions derivation module 112 may be provided as input to one or more control networks. The one or more GML modules may be modulated by the one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. Controlling the one or more GML modules based on the derived color and lighting conditions may ensure that changes made by the one or more controlled GML models are informed by the derived color and lighting conditions derived by the color and lighting conditions derivation module 112 and of the scene in the image obtained by image module 108, for example such that the changes maintain consistency with respect to the derived color and lighting conditions from color and lighting conditions derivation module 112 and of the scene in the image obtained by image module 108.

[0067] The one or more control networks may be one or more neural network architectures that may be used to control the one or more stable diffusion models. The one or more neural network architectures may include one or more neural networks. The derived color and lighting conditions may be input into blocks of a neural network, where a network block refers to a set of neural layers that are put together to form a single unit of a neural network. The one or more control networks may be applied to one or more encoder / decoder levels of the one or more stable diffusion models.

[0068] The one or more control networks may add localized and task specific input conditions, such as the derived color and lighting conditions, to the one or more stable diffusion models. The one or more control networks may be used to control the image generation process of the one or more stable diffusion models. The one or more stable diffusion models may be image diffusion models that are trained to progressively add noise to images and denoise images to generate samples / outputs. The one or more control networks may control the noising and / or denoising processes, and thereby controlling the generated samples / outputs.

[0069] The one or more control networks may add a conditioning image to the one or morestable diffusion models. A control may be added to a conditional output of the one or more stable diffusion models and may be multiplied by a weight. The control may be an image-space constraint that is encoded into a feature-space constraint. For example, the derived color and lighting conditions in image-space may be encoded into the derived color and lighting conditions in feature-space.

[0070] Visualization generation module 118 may be configured to generate, using the one or more controlled GML models, a first visualization comprising new visual data. The new visual data may be integrated with the scene in the image obtained by image module 108 and maintain consistent color and lighting conditions with respect to the scene in the image. Visualization generation module 118 may use the image and the control to generate the first visualization comprising the new visual data based on the scene in the image obtained by image module 108. The first visualization comprising the new visual data may be generated based on the scene in the image obtained by image module 108 may have the same color and lighting conditions as the scene in the image obtained by image module 108. In this way, the the image and the control may be used to guide encoding / decoding latent data to replicate the scene in the image obtained by image module 108. Feature data of the control may therefore decrease the number of plausible inferences made by the network, resulting in more accurate and lower computationally expensive results.

[0071] In some embodiments, visualization generation module 118 may change the scene in the image obtained by image module 108 using the one or more controlled GML models via the GML module 114 and control module 116. In some embodiments, visualization generation module 118 may selectively change the scene in the image such that not all pixels of the scene in the image are modified. In these embodiments, use of the one or more controlled GML models will result in the color and lighting conditions of the selectively changed portions of the scene in the image matching the color and lighting conditions of the non-changed portions of the scene in the image.

[0072] In some embodiments, one or more modules of system 100 may be configured to derive one or more parameters of one or more objects in the first visualization based on the derived color and lighting conditions. In some embodiments, the one or more parameters of the one or more objects may include color information. In some embodiments, the one or more parameters of the one or more objects may include material information.

[0073] In some embodiments, system 100 may be used in a variety of real -world applications, such as architecture and real-estate visualizations. In some examples, such asinterior applications, one or more controlled GML models may be used to generate modifications (e.g., updates, replacements, new placements, etc.) of floors, walls, ceilings, doors, windows, and other elements, while maintaining consistent color and lighting conditions with existing or unchanged portions. In some examples, such as exterior applications, one or more controlled GML models may be used to generate modifications (e.g., updates, replacements, new placements, etc.) of landscaping, walls, roofing elements, doors, windows, and other elements, while maintaining consistent color and lighting conditions with existing or unchanged portions. Integration of interior and / or exterior capabilities provides comprehensive tools for architects, designers, real-estate professionals, craftspersons, and homeowners, to create realistic and accurate representations including both existing elements and modifications.

[0074] In some implementations, computing platform(s) 102, remote platform(s) 104, and / or external resources 120 may be operatively linked via one or more electronic communication links. For example, such electronic communication links may be established, at least in part, via a network such as the Internet and / or other networks. It will be appreciated that this is not intended to be limiting, and that the scope of this disclosure includes implementations in which computing platform(s) 102. remote platform(s) 104, and / or external resources 120 may be operatively linked via some other communication media.

[0075] A given remote platform 104 may include one or more processors configured to execute computer program modules. The computer program modules may be configured to enable an expert or user associated with the given remote platform 104 to interface with system 100 and / or external resources 120. and / or provide other functionality attributed herein to remote platform(s) 104. By way of non-limiting example, a given remote platform 104 and / or a given computing platform 102 may include one or more of a server, a desktop computer, a laptop computer, a handheld computer, a tablet computing platform, aNetBook, a Smartphone, a gaming console, and / or other computing platforms.

[0076] External resources 120 may include sources of information outside of system 100, external entities participating with system 100, and / or other resources. In some implementations, some or all of the functionality attributed herein to external resources 120 may be provided by resources included in system 100.

[0077] Computing platform(s) 102 may include electronic storage 122, one or more processors 124, and / or other components. Computing platform(s) 102 may include communication lines, or ports to enable the exchange of information with a netw ork and / orother computing platforms. Illustration of computing platform(s) 102 in FIG. 1 is not intended to be limiting. Computing platform(s) 102 may include a plurality of hardware, software, and / or firmware components operating together to provide the functionality attributed herein to computing platform(s) 102. For example, computing platform(s) 102 may be implemented by a cloud of computing platforms operating together as computing platform(s) 102.

[0078] Electronic storage 122 may comprise non-transitory storage media that electronically stores information. The electronic storage media of electronic storage 122 may include one or both of system storage that is provided integrally (i.e., substantially nonremovable) with computing platform(s) 102 and / or removable storage that is removably connectable to computing platform(s) 102 via, for example, a port (e.g., a USB port, a firewire port, etc.) or a drive (e.g., a disk drive, etc.). Electronic storage 122 may include one or more of optically readable storage media (e.g.. optical disks, etc.), magnetically readable storage media (e.g., magnetic tape, magnetic hard drive, floppy dnve, etc.), electrical charge-based storage media (e.g., EEPROM, RAM, etc.), solid-state storage media (e.g., flash drive, etc.), and / or other electronically readable storage media. Electronic storage 122 may include one or more virtual storage resources (e.g., cloud storage, a virtual private network, and / or other virtual storage resources). Electronic storage 122 may store software algorithms, information determined by processor(s) 124, information received from computing platform(s) 102, information received from remote platform(s) 104, and / or other information that enables computing platform(s) 102 to function as described herein.

[0079] Processor(s) 124 may be configured to provide information processing capabilities in computing platform(s) 102. As such, processor(s) 124 may include one or more of a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information. Although processor(s) 124 is show n in FIG. 1 as a single entity, this is for illustrative purposes only. In some implementations, processor(s) 124 may include a plurality of processing units. These processing units may be physically located within the same device, or processor(s) 124 may represent processing functionality of a plurality of devices operating in coordination.

[0080] FIG. 2 illustrates a method 200 for controlling color and lighting conditions in GML models, in accordance with one or more implementations. The operations of method 200 presented below are intended to be illustrative. In some implementations, method 200 may be accomplished with one or more additional operations not described, and / or without one or more of the operations discussed. Additionally, the order in which the operations of method 200 areillustrated in FIG. 2 and described below is not intended to be limiting.

[0081] In some implementations, method 200 may be implemented in one or more processing devices (e.g., a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information). The one or more processing devices may include one or more devices executing some or all of the operations of method 200 in response to instructions stored electronically on an electronic storage medium. The one or more processing devices may include one or more devices configured through hardware, firmware, and / or software to be specifically designed for execution of one or more of the operations of method 200.

[0082] An operation 202 may include providing an image. This operation may be performed by an image module similar to or the same as image module 108, in accordance with one or more implementations. In some embodiments, the image is a real-world image captured by a real-world camera. In some embodiments, the image is a synthetic image generated by a model, such as a GML model. In some embodiments, the image may include a building structure, which may be an interior or exterior view of the building. Referring briefly to FIG. 3 A, it illustrates image 300 include building structure 302, according to some embodiments.

[0083] An operation 204 may include inserting at least one simulated calibration object into a portion of a scene in the image. This operation may be performed by an object insertion module similar to or the same as object insertion module 110, in accordance with one or more implementations. In some embodiments, operation 204 includes replacing pixels of the image with the at least one simulated calibration object. In some embodiments, the at least one simulated calibration object includes a chromatic object, such as a mirror sphere, or a matte object, such as a gray sphere. In some embodiments, the at least one simulated calibration object is a color chart. The simulated calibration object may be a three-dimensional (3D) object inserted in two dimensions of the image. Referring briefly to FIG. 3B, it illustrates simulated mirror sphere 304 inserted into a portion of image 300, according to some embodiments.

[0084] In some embodiments, the at least one simulated calibration object may be inserted with respect to 2D space. In these embodiments, the image from operation 202 and an inpaint mask may be provided as input to one or more GML models. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. Stablediffusion models are a pe of deep learning model known for their ability to generate high- quality images while maintaining coherence and stability- in the output. The one or more control networks may use the input mask to adjust parameters or outputs of the one or more GML models, ensuring that the generated visual data inserts the at least one simulated calibration object according to the inpaint mask. In these embodiments, the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image.

[0085] In some embodiments, the at least one simulated calibration object may be inserted with respect to 2D and 3D space. In these embodiments, the image from operation 202, an inpaint mask, and a depth cue may be provided as input to one or more GML models. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors. In some embodiments, the inpaint mask may be superimposed with the depth cue to generate a depth cue with an inpaint mask. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control netw orks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the depth cue with the inpaint mask to adjust parameters or outputs of the one or more GML models, ensuring that the generated visual data inserts the at least one simulated calibration object according to the depth cue with the inpaint mask. In some embodiments, the one or more control networks may be one or more depth conditioned control networks.

[0086] In some embodiments, the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image and for portions of the scene not observed by the camera associated with the image. For the portions of the scene that are observed by the camera associated with the image, the at least one simulated calibration object may be generated based on the depth cue and portions of the image that are in front of the at least one simulated calibration object in 3D space. For the portions of the scene not observed by the camera associated with the image, the at least one simulated calibration object may be generated based on the depth cue and portions of the scene in the image that are behind the at least one simulated calibration object in 3D space.

[0087] In some embodiments, operation 204 may include extending a boundary- portion of the image to create an empty crop portion. In these embodiments, operation 204 may includeinserting the at least one simulated calibration object into the empty crop portion, or a portion thereof.

[0088] In some embodiments, operation 204 may include inserting the at least one simulated calibration object into the portion of the scene in the image using one or more GML techniques such as, for example, inpainting.

[0089] In some embodiments, operation 204 may include inserting a simulated calibration object into a portion of scene in the image. In these embodiments, the simulated calibration object may be used to derive global color conditions, lighting conditions, or both, of the scene in the image as a whole, for example in operation 206. In these embodiments, while the derived global color conditions, lighting conditions, or both, may be representative of the scene in the image as a whole, they may be biased to the portion of the scene in the image the simulated calibration object is inserted into. In some embodiments, the derived global color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in the scene in the image according to the derived global color conditions, lighting conditions, or both, for the scene in the image, for example in operation 212.

[0090] In some embodiments, operation 204 may include inserting a plurality of simulated calibration objects into various portions of the scene in the image. For example, the image may be divided into a grid and a simulated calibration obj ect may be inserted into each cell of the grid. At the most granular level, a simulated calibration object may be inserted at each pixel location of the image. In these embodiments, the plurality of simulated calibration object may be used to derive local color, conditions, lighting conditions, or both, of various portions of the scene in the image, for example in operation 206. In some embodiments, the derived local color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in one or more portions of the scene in the image according to the derived local color conditions, lighting conditions, or both, for the one or more portions of the scene in the image, for example in operation 212.

[0091] In some embodiments, operation 204 may include inserting at least one simulated calibration object into a portion of the scene in the image according to the image. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and Y of the image. In some embodiments, operation 204 may include inserting at least one simulated calibration object into a portion of the scene in the image according to depth cues. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and the Y of the image and the Z of the depthcues. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors.

[0092] In some embodiments, operation 204 may be performed a plurality of times and a representation (e.g., average) may be used.

[0093] An operation 206 may include deriving color conditions, lighting conditions, or both, of the scene in the image based on the at least one simulated calibration object. This operation may be performed by a color and lighting conditions derivation module similar to or the same as color and lighting conditions derivation module 112, in accordance with one or more implementations. The derived color and lighting conditions quantify and represent the color and lighting conditions of the scene in the image. The derived color and lighting conditions may estimate the 3D lighting of the scene depicted in the image. Derived color and lighting conditions may be an image-based and may be encoded into derived color and lighting conditions in feature-space. The derived color and lighting conditions in feature-space may be represented as a vector.

[0094] In some embodiments, operation 206 may include unwrapping the at least one simulated calibration object, such as a mirror sphere or a gray sphere, to generate an equirectangular map (sometimes referred to as an environment map), and derive the color conditions, the lighting conditions, or both, of the scene in the image based on the equirectangular map. In some embodiments, operation 206 may include applying a high dynamic range image map onto the at least one simulated calibration object, such as a mirror sphere or a gray sphere, and deriving the color conditions, the lighting conditions, or both based on the high dynamic range image map. In some embodiments, operation 206 may include mapping the at least one simulated calibration object, such as a color chart, to a known calibration object, such as a known color chart, and deriving the color conditions, the lighting conditions, or both, based on the mapping. While specific techniques for deriving the color conditions, the lighting conditions, or both, are disclosed herein, one of ordinary skill in the art may appreciate other techniques, such as those used in visual effects, may be used.

[0095] The derived color conditions may include various parameters such as camera colors. These parameters provide a comprehensive representation of the color environment in the image. The derived lighting conditions may include various parameters such as light directionality, light intensity, and light temperature. These parameters provide a comprehensive representation of the lighting environment in the image.

[0096] In some embodiments, operation 206 may include deriving color conditions,lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image. In other words, operation 206 may include deriving color conditions, lighting conditions, or both, of a front half of the scene in the image where the front half of the scene is the portion visible by the camera associated with the image. In some embodiments, operation 206 may include deriving color conditions, lighting conditions, or both, of the front half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of the at least one simulated calibration object. In these embodiments, portions of the equirectangular map that correspond to portions of the scene not observed by the camera associated with the image may be filled in, for example using interpolation, based on portions of the equirectangular map that correspond to portions of the scene observed by the camera associated with the image.

[0097] In some embodiments, operation 206 may include deriving color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by the camera associated with the image and for portions of the scene not observed by the camera associated with the image. In other words, operation 206 may include deriving color conditions, lighting conditions, or both, of a front half and a back half of the scene in the image where the front half of the scene is the portion visible by the camera associated with the image and where the back half of the scene is the portion not visible by the camera. In some embodiments, operation 206 may include deriving color conditions, lighting conditions, or both, of the front half and the back half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of and substantially behind the at least one simulated calibration object. In these embodiments, the equirectangular map may correspond to the entire scene, or substantially the entire scene.

[0098] Operation 206 may include using the at least one simulated calibration object as a reference point to represent the color and lighting conditions of the scene in the image, for example rather than relying on pre-existing lighting models, which may be generic, and rather than manual adjustments, which may require expert intervention and may be time consuming.

[0099] Operation 204 may include inserting the at least one simulated calibration object into the portion of the scene in the image and operation 206 may include deriving a color and lighting condition of the scene in the image based on the at least one simulated calibration object. Operation 204 introduces an element (i.e., the at least one simulated calibration object) into the image and operation 206 relies on the at least one simulated calibration object ratherthan solely the image.

[0100] An operation 208 may include providing the image from operation 202 as input to one or more GML models. This operation may be performed by a generative machine learning module similar to or the same as generative machine learning module 114, in accordance with one or more implementations. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the input derived color and lighting conditions to adjust parameters or outputs of the GML model, ensuring that the generated visual data aligns with the image’s color and lighting environment.

[0101] An operation 210 may include controlling the one or more GML models based on the derived color and lighting conditions. This operation may be performed by a control module similar to or the same as control module 116, in accordance with one or more implementations. The derived color and lighting conditions from operation 206 may be provided as input to one or more control networks. The one or more GML models may be modulated by the one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. Controlling the one or more GML models based on the derived color and lighting conditions may ensure that changes made by the one or more controlled GML models are informed by the derived color and lighting conditions derived by operation 206 and of the scene in the image obtained by operation 202, for example such that the changes maintain consistency with respect to the derived color and lighting conditions from operation 206 and of the scene in the image obtained by operation 202.

[0102] The one or more control networks may be one or more neural network architectures that may be used to control the one or more stable diffusion models. The one or more neural network architectures may include one or more neural networks. The derived color and lighting conditions may be input into blocks of a neural network, where a network block refers to a set of neural layers that are put together to form a single unit of a neural network. The one or more control networks may be applied to one or more encoder / decoder levels of the one or more stable diffusion models.

[0103] The one or more control networks may add localized and task specific input conditions, such as the derived color and lighting conditions, to the one or more stable diffusion models. The one or more control networks may be used to control the image generation processof the one or more stable diffusion models. The one or more stable diffusion models may be image diffusion models that are trained to progressively add noise to images and denoise images to generate samples / outputs. The one or more control networks may control the noising and / or denoising processes, and thereby controlling the generated samples / outputs.

[0104] The one or more control networks may add a conditioning image to the one or more stable diffusion models. A control may be added to a conditional output of the one or more stable diffusion models and may be multiplied by a weight. The control may be an image-space control that is encoded into a feature-space control. For example, the derived color and lighting conditions in image-space may be encoded into the derived color and lighting conditions in feature-space.

[0105] An operation 212 may include generating, using the one or more controlled GML models, a first visualization comprising new visual data. This operation may be performed by a visualization generation module similar to or the same as visualization generation module 118, in accordance with one or more implementations. The new visual data may be integrated with the scene in the image obtained by operation 202 and maintain consistent color and lighting conditions. Referring briefly to FIG. 3C, it illustrates visualization 310 including motorcycle 312, according to some embodiments. Operation 212 may use the image and the control to generate the first visualization comprising the new visual data based on the image obtained in operation 202. The first visualization comprising the new visual data may be generated based on the scene in the image obtained in operation 202 may have the same color and lighting conditions as the scene in the image obtained in operation 202. In this way, the image and the control may be used to guide encoding / decoding latent data to replicate the scene in the image obtained in operation 202, specifically the color and lighting conditions of the scene in the image. Feature data of the control may therefore decrease the number of plausible inferences made by the network, resulting in more accurate and lower computationally expensive results.

[0106] In some embodiments, operation 212 may include generating, using the one or more controlled GML models, a first visualization comprising a new 3D object, wherein the first visualization includes a 2D representation of the new 3D object.

[0107] In some embodiments, operation 212 may include generating a 3D environment based on the derived color and lighting conditions, for example from operation 206. In some embodiments, the 3D environment may include placing the environment map in the 3D environment at a distance from a camera pose associated with the image, for example at fivemeters, at ten meters, at an infinite distance, etc.. In some embodiments, the 3D environment may be a spherical representation of the environment map where the environment map is wrapped along the inside of a sphere and 3D objects are placed within the sphere. In some embodiments, the 3D environment may be cube representation of the environment map where the environment map is wrapped along the inside of a cuboid and 3D objects are placed within the cuboid. In some embodiments, the 3D objects include a 3D model of the scene in the image. In some embodiments, the 3D model may be generated based on the image. In some embodiments, the 3D model may be generated based on a depth cue. The depth cue may be predicted using depth predictors such as single image depth predictors. The some embodiments the 3D representation may be generated based on one or more modeling techniques, such as primitive-based modeling.

[0108] In some embodiments, operation 212 may include placing 3D objects within the generated 3D environment. Placing the 3D objects within the generated 3D environment may result in the 3D objects being reflective of the generated 3D environment. In some embodiments, operation 206 may include ray-tracing the derived conditions onto the 3D objects in the 3D environment and observing the 3D objects according to a camera associated with the image. In some embodiments, this observation of the 3D objects within the generated 3D environment according to the camera of the image may be represented in 2D (e.g., as an image) or in 3D, and may be provided as input to the one or more controlled GML models. The one or more controlled GML models may use the observation of the 3D objects and the image to generate the first visualization comprising the 3D objects. In some embodiments, the one or more controlled GML models may composite the observation of the 3D objects and the image to generate the first visualization comprising the 3D objects.

[0109] In some embodiments, operation 212 may include accessing a 3D representation of the scene in the image and may ray -trace the derived conditions to calculate occlusions and shadows based on the calculated occlusions. In some embodiments, the 3D representation maybe based on a depth cue. The depth cue may be predicted using depth predictors such as single image depth predictors. The some embodiments the 3D representation may be generated based on one or more modeling techniques, such as primitive-based modeling.

[0110] In some embodiments, operation 212 may include changing the image obtained by operation 202 using the one or more controlled GML models. In some embodiments, operation 212 may include selectively changing the image such that not all pixels of the image are modified. In these embodiments, use of the one or more controlled GML models will result in the color and lighting conditions of the selectively changed portions of the scene in the imagematching the color and lighting conditions of the non-changed portions of the scene in the image.

[0111] In some embodiments, method 200 include deriving one or more parameters of one or more objects in the first visualization based on the derived color and lighting conditions. In some embodiments, the one or more parameters of the one or more objects may include color information. In some embodiments, the one or more parameters of the one or more objects may include material information.

[0112] In some embodiments, method 200 may be used in a variety7of real-world applications, such as architecture and real-estate visualizations. In some examples, such as interior applications, one or more constrained GML models may be used to generate modifications (e g., updates, replacements, new placements, etc.) of floors, walls, ceilings, doors, windows, and other elements, while maintaining consistent color and lighting conditions with existing or unchanged portions. In some examples, such as exterior applications, one or more constrained GML models may be used to generate modifications (e.g., updates, replacements, new placements, etc.) of landscaping, walls, roofing elements, doors, windows, and other elements, while maintaining consistent color and lighting conditions with existing or unchanged portions. Integration of interior and / or exterior capabilities provides comprehensive tools for architects, designers, real-estate professionals, craftspersons, and homeowners, to create realistic and accurate representations including both existing elements and modifications.

[0113] FIG. 4 illustrates a system 400 configured for deriving color and lighting conditions, in accordance with one or more implementations. In some implementations, system 400 may include one or more computing platforms 402. Computing platform(s) 402 may be configured to communicate with one or more remote platforms 404 according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. Remote platform(s) 404 may be configured to communicate with other remote platforms via computing platform(s) 402 and / or according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. Users may access system 400 via remote platform(s) 404.

[0114] Computing platform(s) 402 may be configured by machine-readable instructions 406. Machine-readable instructions 406 may include one or more instruction modules. The instruction modules may include computer program modules. The instruction modules may include one or more of image module 408, object insertion module 410, visualization rendition module 412, change detection module 414, color and lighting condition derivation module 416,and / or other instruction modules.

[0115] Image module 408 may be configured to provide, receive, capture, or otherwise obtain an image.

[0116] Object insertion module 410 may be configured to insert at least one simulated calibration object into a portion of the scene in the image. In some embodiments, the object insertion module 410 may be configured to extend a boundary portion of the scene in the image to create an empty crop portion. In these embodiments, the object insertion module 410 may be configured to insert the at least one simulated calibration object into the empty crop portion, or a portion thereof. In some embodiments, the at least one simulated calibration object includes a chromatic object, such as a mirror sphere, or a matte object, such as a gray sphere. In some embodiments, the at least one simulated calibration object is a color chart. The simulated calibration obj ect may be a three-dimensional (3D) obj ect inserted in two dimensions of the image.

[0117] In some embodiments, object insertion module 410 may be configured to extend a boundary of the image to create an empty crop portion and insert the at least one simulated calibration object into the empty crop portion, or portion thereof.

[0118] In some embodiments, object insertion module 410 may be configured to insert a simulated calibration object into a portion of scene in the image. In these embodiments, a modified version of the simulated calibration object may be used to derive global color conditions, lighting conditions, or both, of the scene in the image as a whole, for example by color and lighting condition derivation module 416. In these embodiments, while the derived global color conditions, lighting conditions, or both, may be representative of the scene in the image as a whole, they may be biased to the portion of the scene in the image the simulated calibration object is inserted into. In some embodiments, the derived global color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in the scene in the image according to the derived global color conditions, lighting conditions, or both, for the scene in the image.

[0119] In some embodiments, object insertion module 410 may be configured to insert a plurality of simulated calibration objects into various portions of the scene in the image. For example, the image may be divided into a grid and a simulated calibration object may be inserted into each cell of the grid. At the most granular level, a simulated calibration object may be inserted at each pixel location of the image. In these embodiments, modified versions of the plurality of simulated calibration object may be used to derive local color, conditions,lighting conditions, or both, of various portions of the scene in the image, for example by color and light condition detection module 416. In some embodiments, the derived local color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in one or more portions of the scene in the image according to the derived local color conditions, lighting conditions, or both, for the one or more portions of the scene in the image.

[0120] In some embodiments, object insertion module 410 may be configured to insert at least one simulated calibration object into a portion of the scene in the image according to the image. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and Y of the image. In some embodiments, obj ect insertion module 410 may be configured to insert at least one simulated calibration object into a portion of the scene in the image according to depth cues. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and the Y of the image and the Z of the depth cues. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors.

[0121] In some embodiments, object insertion module 410 may be configured to insert at least one simulated calibration object a plurality' of times and a representation (e.g., average) may be used.

[0122] Visualization rendition module 412 may be configured to render anew visualization of the scene in the image including a modified version of the at least one simulated calibration object. In some embodiments, the visualization rendition module 412 may be configured to render the new visualization with generative machine learning, including stable diffusion.

[0123] In some embodiments, visualization rendition module 412 may be configured to render the new visualization of the scene in the image including a modified version of the at least one simulated calibration object may be with respect to 2D space. In these embodiments, the image from image module 408, an inpaint mask, and a representation of the at least one simulated calibration object may be provided as input to one or more GML models. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the input mask to adjust parameters or outputs of the one or more GML models and the one or more GML models may use the image and the representation of the at least one simulated calibration object, ensuringthat the new visualization includes a modified version of the at least one simulated calibration object according to the inpaint mask and the representation of the at least one simulated calibration object. In these embodiments, the modified version of at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image.

[0124] In some embodiments, visualization rendition module 412 may be configured to render the new visualization of the scene in the image including a modified version of the at least one simulated calibration object may be with respect to 2D and 3D space. In these embodiments, the image from image module 408, an inpaint mask, a depth cue, and a representation of the at least one simulated calibration object may be provided as input to one or more GML models. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors. In some embodiments, the inpaint mask may be superimposed with the depth cue to generate a depth cue with an inpaint mask. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the depth cue with the inpaint mask and the one or more GML models may use the image and the representation of the at least one simulated calibration object to adjust parameters or outputs of the one or more GML model, ensuring that the new visualization includes a modified version of the at least one simulated calibration object according to the depth cue with the inpaint mask. In some embodiments, the one or more control networks may be one or more depth conditioned control networks.

[0125] In some embodiments, the modified version of the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image and for portions of the scene not observed by the camera associated with the image. For the portions of the scene that are observed by the camera associated with the image, the modified version of the at least one simulated calibration object may be generated based on the depth cue and portions of the image that are in front of the at least one simulated calibration obj ect in 3D space. For the portions of the scene not observed by the camera associated with the image, the modified version of the at least one simulated calibration object may be generated based on the depth cue and portions of the scene in the image that are behind the at least onesimulated calibration object in 3D space.

[0126] Change detection module 414 may be configured to detect a change between the at least one simulated calibration object in the image and the modified version of the at least one simulated calibration object in the new visualization. In some embodiments, change detection module 414 may be configured to compare one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization.

[0127] Color and lighting condition derivation module 416 may be configured to derive a color and lighting condition of the new visualization based on the detected change. The derived color and lighting condition quantifies and represents the color and lighting condition of the scene in the new visualization. The derived color and lighting condition may estimate the 3D lighting of the scene depicted in the new visualization. In some embodiments, the color and lighting condition includes light directionality, light intensity, light temperature, color intensity, color temperature, and the like, or a combination thereof.

[0128] In some embodiments, change detection module 414 may be optional. In these embodiments, color and lighting condition derivation module 416 may be configured to derive a color and lighting condition of the new visualization based on the new visualization. In some embodiments, deriving the color and lighting condition of the new visualization based on the new visualization may include unwrapping the modified version of the at least one simulated calibration object, such as a mirror sphere or a gray sphere, to generate an equirectangular map (sometimes referred to as an environment map), and deriving the color and lighting condition of the scene in the new visualization based on the equirectangular map. In some embodiments, this may include applying a high dynamic range image map onto the modified version of the at least one simulated calibration object, such as a mirror sphere or a gray sphere, and deriving the color and lighting condition based on the high dynamic range image map. In some embodiments, this may include mapping the modified version of the at least one simulated calibration object, such as a color chart, to a known calibration object, such as a known color chart, and deriving the color and lighting condition based on the mapping. While specific techniques for deriving the color and lighting condition are disclosed herein, one of ordinary skill in the art may appreciate other techniques, such as those used in visual effects, may be used.

[0129] In some embodiments, one or more existing or additional modules of the system 400 may be configured to derive a camera condition of the new visualization based on themodified version of the at least one simulated calibration object. In some embodiments, the camera condition may include camera extrinsics, camera intrinsics such as image sensor properties, and the like, or a combination thereof.

[0130] In some embodiments, color and lighting condition derivation module 416 may be configured to derive the color and lighting condition of the scene in the new visualization for portions of the scene observed by a camera associated with the new visualization. In other words, color and lighting condition derivation module 416 may be configured to derive a color and lighting condition of a front half of the scene in the new visualization where the front half of the scene is the portion visible by the camera associated with the new visualization. In some embodiments, color and lighting condition derivation module 416 may be configured to derive color conditions, lighting conditions, or both, of the front half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of the at least one simulated calibration object. In these embodiments, portions of the equirectangular map that correspond to portions of the scene not observed by the camera associated with the new visualization may be filled in, for example using interpolation, based on portions of the equirectangular map that correspond to portions of the scene observed by the camera associated with the new visualization.

[0131] In some embodiments, color and lighting condition derivation module 416 may be configured to denve a color and lighting condition of the scene in the new visualization for portions of the scene observed by the camera associated with the new visualization and for portions of the scene not observed by the camera associated with the new visualization. In other words, color and lighting condition derivation module 416 may be configured to derive a color and lighting condition of a front half and a back half of the scene in the new visualization where the front half of the scene is the portion visible by the camera associated with the new visualization and where the back half of the scene is the portion not visible by the camera. In some embodiments, color and lighting condition derivation module may be configured to derive color conditions, lighting conditions, or both, of the front half and the back half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of and substantially behind the at least one simulated calibration object. In these embodiments, the equirectangular map may correspond to the entire scene, or substantially the entire scene.

[0132] In some embodiments, one or more existing modules, such as the object insertion module 410 and / or the visualization rendition module 412, or one or more additional modules of the system 400 may be configured to adjust the new visualization based on the color andlighting condition. In some embodiments, adjusting the new visualization may include inserting a new object in the new visualization. In some embodiments, inserting the new object may include rendering the new object based on the color and lighting condition.

[0133] In some embodiments, one or more existing modules, such as the color and lighting condition derivation module 416, or one or more additional modules of the system 410 may be configured to derive an color and lighting condition of the scene in the image based on the detected change.

[0134] In some implementations, computing platform(s) 402, remote platform(s) 404, and / or external resources 420 may be operatively linked via one or more electronic communication links. For example, such electronic communication links may be established, at least in part, via a network such as the Internet and / or other networks. It will be appreciated that this is not intended to be limiting, and that the scope of this disclosure includes implementations in which computing platform(s) 402. remote platform(s) 404, and / or external resources 420 may be operatively linked via some other communication media.

[0135] A given remote platform 404 may include one or more processors configured to execute computer program modules. The computer program modules may be configured to enable an expert or user associated with the given remote platform 404 to interface with system 400 and / or external resources 420, and / or provide other functionality attributed herein to remote platform(s) 404. By way of non-limiting example, a given remote platform 404 and / or a given computing platform 402 may include one or more of a server, a desktop computer, a laptop computer, a handheld computer, a tablet computing platform, aNetBook, a Smartphone, a gaming console, and / or other computing platforms.

[0136] External resources 420 may include sources of information outside of system 400, external entities participating with system 400, and / or other resources. In some implementations, some or all of the functionality attributed herein to external resources 420 may be provided by resources included in system 400.

[0137] Computing platform(s) 402 may include electronic storage 422, one or more processors 424. and / or other components. Computing platform(s) 402 may include communication lines, or ports to enable the exchange of information with a network and / or other computing platforms. Illustration of computing platform(s) 402 in FIG. 4 is not intended to be limiting. Computing platform(s) 402 may include a plurality of hardware, software, and / or firmware components operating together to provide the functionality attributed herein to computing platform(s) 402. For example, computing platform(s) 402 may be implementedby a cloud of computing platforms operating together as computing platform(s) 402.

[0138] Electronic storage 422 may comprise non-transitory storage media that electronically stores information. The electronic storage media of electronic storage 422 may include one or both of system storage that is provided integrally (i.e., substantially nonremovable) with computing platform(s) 402 and / or removable storage that is removably connectable to computing platform(s) 402 via, for example, a port (e.g., a USB port, a firewire port, etc.) or a drive (e.g., a disk drive, etc.). Electronic storage 422 may include one or more of optically readable storage media (e.g.. optical disks, etc.), magnetically readable storage media (e.g., magnetic tape, magnetic hard drive, floppy drive, etc.), electrical charge-based storage media (e.g., EEPROM, RAM, etc.), solid-state storage media (e.g., flash drive, etc.), and / or other electronically readable storage media. Electronic storage 422 may include one or more virtual storage resources (e.g., cloud storage, a virtual private network, and / or other virtual storage resources). Electronic storage 422 may store software algorithms, information determined by processor(s) 424, information received from computing platform(s) 402, information received from remote platform(s) 404, and / or other information that enables computing platform(s) 402 to function as described herein.

[0139] Processor(s) 424 may be configured to provide information processing capabilities in computing platform(s) 402. As such, processor(s) 424 may include one or more of a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information. Although processor(s) 424 is show n in FIG. 4 as a single entity, this is for illustrative purposes only. In some implementations, processor(s) 424 may include a plurality of processing units. These processing units may be physically located within the same device, or processor(s) 424 may represent processing functionality7of a plurality of devices operating in coordination. Processor(s) 424 may be configured to execute modules 408, 410, 412, 414. and / or 416, and / or other modules. Processor(s) 424 may be configured to execute modules 408, 410. 412, 414, and / or 416, and / or other modules by software; hardware; firmware; some combination of software, hardware, and / or firmware; and / or other mechanisms for configuring processing capabilities on processor(s) 424. As used herein, the term "module" may refer to any component or set of components that perform the functionality' attributed to the module. This may include one or more physical processors during execution of processor readable instructions, the processor readable instructions, circuitry, hardware, storage media, or any other components.

[0140] It should be appreciated that although modules 408, 410, 412, 414, and / or 416 areillustrated in FIG. 4 as being implemented within a single processing unit, in implementations in which processor(s) 424 includes multiple processing units, one or more of modules 408, 410, 412, 414, and / or 416 may be implemented remotely from the other modules. The description of the functionality provided by the different modules 408, 410, 412, 414, and / or 416 described below is for illustrative purposes, and is not intended to be limiting, as any of modules 408, 410, 412, 414, and / or 416 may provide more or less functionality than is described. For example, one or more of modules 408. 410, 412. 414, and / or 416 may be eliminated, and some or all of its functionality may be provided by other ones of modules 408, 410, 412, 414, and / or 416. As another example, processor(s) 424 may be configured to execute one or more additional modules that may perform some or all of the functionality attributed below to one of modules 408, 410, 412, 414, and / or 416.

[0141] FIG. 5 illustrates a method 500 for deriving color and lighting conditions, in accordance with one or more implementations. The operations of method 500 presented below are intended to be illustrative. In some implementations, method 500 may be accomplished with one or more additional operations not described, and / or without one or more of the operations discussed. Additionally, the order in which the operations of method 500 are illustrated in FIG. 5 and described below is not intended to be limiting.

[0142] In some implementations, method 500 may be implemented in one or more processing devices (e.g.. a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information). The one or more processing devices may include one or more devices executing some or all of the operations of method 500 in response to instructions stored electronically on an electronic storage medium. The one or more processing devices may include one or more devices configured through hardware, firmware, and / or software to be specifically designed for execution of one or more of the operations of method 500.

[0143] An operation 502 may include providing, receiving, capturing, or otherwise obtaining an image. In some embodiments, the image may be a two-dimensional (2D) view7of a three-dimensional (3D) model. Referring briefly to FIG. 6A. it illustrates an image 600, according to some embodiments. Operation 502 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to image module 408, in accordance with one or more implementations.

[0144] An operation 504 may include inserting at least one simulated calibration objectinto a portion of the scene in the image. Operation 506 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to object insertion module 410, in accordance with one or more implementations. In some embodiments, the at least one simulated calibration object includes a chromatic object, such as a mirror sphere, or a matte object, such as a gray sphere. In some embodiments, the at least one simulated calibration object is a color chart. The simulated calibration object may be a three-dimensional (3D) object inserted in two dimensions of the image.

[0145] In some embodiments, operation 504 may include extending a boundary of the image to create an empty crop portion and inserting the at least one simulated calibration object into the empty crop portion, or portion thereof. Referring briefly to FIG. 6B, it illustrates an image 600 including an empty crop portion 602, according to some embodiments. Referring briefly to FIG. 6C. it illustrates an image 603 including a mirror sphere 604A in a portion of an empty portion 602, according to some embodiments.

[0146] In some embodiments, operation 504 may include inserting a simulated calibration object into a portion of scene in the image. In these embodiments, a modified version of the simulated calibration object may be used to derive global color conditions, lighting conditions, or both, of the scene in the image as a whole, for example in operation 510. In these embodiments, while the derived global color conditions, lighting conditions, or both, may be representative of the scene in the image as a whole, they may be biased to the portion of the scene in the image the simulated calibration object is inserted into. In some embodiments, the derived global color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in the scene in the image according to the derived global color conditions, lighting conditions, or both, for the scene in the image.

[0147] In some embodiments, operation 504 may include inserting a plurality of simulated calibration objects into various portions of the scene in the image. For example, the image may be divided into a grid and a simulated calibration obj ect may be inserted into each cell of the grid. At the most granular level, a simulated calibration object may be inserted at each pixel location of the image. In these embodiments, modified versions of the plurality of simulated calibration object may be used to derive local color, conditions, lighting conditions, or both, of various portions of the scene in the image, for example in operation 510. In some embodiments, the derived local color conditions, lighting conditions, or both, may be used to generate a new visualization including new visual data in one or more portions of the scene in the image according to the derived local color conditions, lighting conditions, or both, for the one or moreportions of the scene in the image.

[0148] In some embodiments, operation 504 may include inserting at least one simulated calibration object into a portion of the scene in the image according to the image. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and Y of the image. In some embodiments, operation 504 may include inserting at least one simulated calibration object into a portion of the scene in the image according to depth cues. In these embodiments, the at least one simulated calibration object may be inserted into the scene according to the X and the Y of the image and the Z of the depth cues. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors.

[0149] In some embodiments, operation 504 may be performed a plurality of times and a representation (e.g., average) may be used.

[0150] An operation 506 may include rendering a new visualization of the scene in the image including a modified version of the at least one simulated calibration object. In some embodiments, the new visualization is rendered with generative machine learning. In some embodiments, generative machine learning includes stable diffusion. Operation 508 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to visualization rendition module 414, in accordance with one or more implementations. Referring briefly to FIG. 6D, it illustrates a new visualization 606 of an image 600 including a modified mirror sphere 604B, according to some embodiments. The modified mirror sphere 604B is a modified version of mirror sphere 604A of FIG. 6C.

[0151] In some embodiments, rendering the new visualization of the scene in the image including a modified version of the at least one simulated calibration object may be with respect to 2D space. In these embodiments, the image from operation 502, an inpaint mask, and a representation of the at least one simulated calibration object may be provided as input to one or more GML models. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the input mask to adjust parameters or outputs of the one or more GML models and the one or more GML models may use the image and the representation of the at least one simulated calibration object, ensuring that the new visualization includes a modified version of the atleast one simulated calibration object according to the inpaint mask and the representation of the at least one simulated calibration object. In these embodiments, the modified version of at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image.

[0152] In some embodiments, rendering the new visualization of the scene in the image including a modified version of the at least one simulated calibration object may be with respect to 2D and 3D space. In these embodiments, the image from operation 502. an inpaint mask, a depth cue, and a representation of the at least one simulated calibration object may be provided as input to one or more GML models. In some embodiments, the depth cue may be predicted using depth predictors such as single image depth predictors. In some embodiments, the inpaint mask may be superimposed with the depth cue to generate a depth cue with an inpaint mask. In some embodiments, the one or more GML models may comprise one or more stable diffusion models. The one or more GML models may be modulated by one or more control networks. For example, the one or more stable diffusion models may be modulated by the one or more control networks. The one or more control networks may use the depth cue with the inpaint mask and the one or more GML models may use the image and the representation of the at least one simulated calibration object to adjust parameters or outputs of the one or more GML model, ensuring that the new visualization includes a modified version of the at least one simulated calibration object according to the depth cue with the inpaint mask. In some embodiments, the one or more control networks may be one or more depth conditioned control networks.

[0153] In some embodiments, the modified version of the at least one simulated calibration object may be reflective / representative of the color conditions, lighting conditions, or both, of the scene in the image for portions of the scene observed by a camera associated with the image and for portions of the scene not observed by the camera associated with the image. For the portions of the scene that are observed by the camera associated with the image, the modified version of the at least one simulated calibration object may be generated based on the depth cue and portions of the image that are in front of the at least one simulated calibration object in 3D space. For the portions of the scene not observed by the camera associated with the image, the modified version of the at least one simulated calibration object may be generated based on the depth cue and portions of the scene in the image that are behind the at least one simulated calibration object in 3D space.

[0154] An operation 508 may include detecting a change between the at least one simulatedcalibration object in the image and the modified version of the at least one simulated calibration object in the new visualization. In some embodiments, the change may be a change in color and / or lighting condition. Color and / or lighting condition may include light directionality, light intensity, light temperature, color temperature, color intensity, and the like, or a combination thereof. In some embodiments, the change may be a change in texture condition. Operation 508 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to change detection module 414. in accordance with one or more implementations. In some embodiments, detecting the change may include comparing one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization. Referring briefly to FIGS. 6C and 6D, a change between the minor sphere 604A in the image 600 and the mirror sphere 604B in the new visualization 606 of the scene in the image 600 may be detected.

[0155] An operation 510 may include deriving a color and lighting condition of the new visualization based on the detected change. In some embodiments, the color and lighting condition includes light directionality, light intensity, light temperature, color intensity, color temperature, and the like, or a combination thereof. The derived color and lighting condition quantifies and represents the color and lighting condition of the scene in the new visualization. The derived color and lighting condition may estimate the 3D lighting of the scene depicted in the new visualization. Operation 510 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to color and lighting condition derivation module 416, in accordance with one or more implementations.

[0156] In some embodiments, operation 508 may be optional. In these embodiments, operation 510 may include deriving a color and lighting condition of the new visualization based on the new visualization. In some embodiments, deriving the color and lighting condition of the new visualization based on the new visualization may include unwrapping the modified version of the at least one simulated calibration obj ect, such as a mirror sphere or a gray sphere, to generate an equirectangular map (sometimes referred to as an environment map), and deriving the color and lighting condition of the scene in the new visualization based on the equirectangular map. In some embodiments, this may include applying a high dynamic range image map onto the modified version of the at least one simulated calibration object, such as a mirror sphere or a gray sphere, and deriving the color and lighting condition based on the high dynamic range image map. In some embodiments, this may include mapping the modifiedversion of the at least one simulated calibration object, such as a color chart, to a known calibration object, such as a known color chart, and deriving the color and lighting condition based on the mapping. While specific techniques for deriving the color and lighting condition are disclosed herein, one of ordinary skill in the art may appreciate other techniques, such as those used in visual effects, may be used.

[0157] In some embodiments, the method 500 may further include deriving a camera condition of the new visualization based on the modified version of the at least one simulated calibration object. The camera condition may include camera extrinsics. camera intrinsics such as image sensor properties, and the like, or a combination thereof.

[0158] In some embodiments, operation 510 may include deriving the color and lighting condition of the scene in the new visualization for portions of the scene observed by a camera associated with the new visualization. In other words, operation 510 may include deriving a color and lighting condition of a front half of the scene in the new visualization where the front half of the scene is the portion visible by the camera associated with the new visualization. In some embodiments, operation 510 may include deriving color conditions, lighting conditions, or both, of the front half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of the at least one simulated calibration object. In these embodiments, portions of the equirectangular map that correspond to portions of the scene not observed by the camera associated with the new visualization may be filled in, for example using interpolation, based on portions of the equirectangular map that correspond to portions of the scene observed by the camera associated with the new visualization.

[0159] In some embodiments, operation 510 may include deriving a color and lighting condition of the scene in the new visualization for portions of the scene observed by the camera associated with the new visualization and for portions of the scene not observed by the camera associated with the new visualization. In other words, operation 510 may include deriving a color and lighting condition of a front half and a back half of the scene in the new visualization where the front half of the scene is the portion visible by the camera associated with the new visualization and where the back half of the scene is the portion not visible by the camera. In some embodiments, operation 510 may include deriving color conditions, lighting conditions, or both, of the front half and the back half of the scene in the image based on a 3D position of the at least one simulated calibration object in the scene in the image and scene data that is substantially in front of and substantially behind the at least one simulated calibration object. In these embodiments, the equirectangular map may correspond to the entire scene, orsubstantially the entire scene.

[0160] In some embodiments, the method 500 may further include adjusting the new visualization based on the color and lighting condition. In some embodiments, adjusting the new visualization may include inserting a new object in the new visualization. Inserting the new object may include rendering the new object based on the color and lighting condition. Referring briefly to FIG. 6E it illustrates an adjusted new visualization 608 including a car 610, according to some embodiments. The car 610 is rendered based on the color and lighting condition.

[0161] In some embodiments, the method 500 may further include deriving an color and lighting condition of the scene in the image based on the detected change.

[0162] The following are example embodiments.

[0163] In some examples, a method of controlling color and light conditions of a generative machine learning model comprises: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; providing the image as input to a generative machine learning model; controlling the generative machine learning model based on the derived lighting condition; and generating, using the controlled generative machine learning model, a first visualization comprising new visual data.

[0164] In some examples, the method further comprises extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

[0165] In some examples, the at least one simulated calibration object comprises a chromatic object. In some examples, the chromatic object is a mirror sphere. In some examples, the at least one simulated calibration object comprises a matte object. In some examples, the matte object is a gray sphere. In some examples, the at least one simulated calibration object is a color chart.

[0166] In some examples, inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

[0167] In some examples, deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lightingcondition is based on the equirectangular map.

[0168] In some examples, deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

[0169] In some examples, the lighting condition comprises light directionality. In some examples, the lighting condition comprises light intensity. In some examples, the lighting condition comprises light temperature.

[0170] In some examples, the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

[0171] In some examples, the image is input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

[0172] In some examples, the method further comprising deriving one or more parameters of one or more objects in the first visualization based on the derived lighting condition.

[0173] In some examples, the one or more parameters of the one or more objects in the first visualization comprise color information. In some examples, the one or more parameters of the one or more objects in the first visualization comprise material information.

[0174] In some examples, a system including one or more processors, and one or more non-transitory computer-readable storage media stonng instructions, which when executed by the one or more processors, causes the system to carry out the method of any one of the preceding examples.

[0175] In some examples, one or more non-transitory computer-readable storage media carrying machine-readable instructions, which when executed by one or more processors of one or more machines, cause the one or more machines to cany’ out the method of any one of the preceding examples.

[0176] In some examples, a method of controlling color and lighting conditions of a generative machine learning model comprises: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; generating a view of a 3D object within a 3D environment comprising a representation of the derived lighting condition; providing the image and the view of the 3D object as inputs to a generative machine learning model; and generating, using the generative machine learning model, a firstvisualization comprising new visual data.

[0177] In some examples, the method further comprises extending a boundary' of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

[0178] In some examples, the at least one simulated calibration object comprises a chromatic object. In some examples, the chromatic object is a mirror sphere. In some examples, the at least one simulated calibration object comprises a matte object. In some examples, the matte object is a gray sphere. In some examples, the at least one simulated calibration object is a color chart.

[0179] In some examples, inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

[0180] In some examples, deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

[0181] In some examples, deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

[0182] In some examples, the lighting condition comprises light directionality. In some examples, the lighting condition comprises light intensity. In some examples, the lighting condition comprises light temperature.

[0183] In some examples, the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

[0184] In some examples, the image and the view of the 3D object are input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

[0185] In some examples, the 3D environment comprising the representation of the derived lighting condition comprises a representation of the equirectangular map. In some examples, the representation of the equirectangular map is at a distance from a position of the 3D object. In some examples, the distance is substantially infinity.

[0186] In some examples, the view of the 3D object within the 3D environment is from acamera pose associated with the image.

[0187] In some examples, the view of the 3D object is a 3D view. In some examples, the view of the 3D object is a 2D view.

[0188] In some examples, a system including one or more processors, and one or more non-transitory computer-readable storage media storing instructions, which when executed by the one or more processors, causes the system to carry out the method of any one of the preceding examples.

[0189] In some examples, one or more non-transitory computer-readable storage media carrying machine-readable instructions, which when executed by one or more processors of one or more machines, cause the one or more machines to earn' out the method of any one of the preceding examples.

[0190] In some examples, a method of deriving lighting conditions comprises: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; rendering a new visualization of the scene in the image including a modified version of the at least one simulated calibration object; detecting a change between the at least one simulated calibration object in the image and the modified version of the at least one simulated calibration object in the new visualization; and deriving a lighting condition of the new visualization based on the detected change.

[0191] In some examples, the method further comprises extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

[0192] In some examples, the new' visualization is rendered with generative machine learning.

[0193] In some examples, generative machine learning includes stable diffusion.

[0194] In some examples, the at least one simulated calibration object comprises a chromatic object. In some examples, the chromatic object is a mirror sphere. In some examples, the at least one simulated calibration object comprises a matte object. In some examples, the matte object is a gray sphere. In some examples, the at least one simulated calibration object is a color chart.

[0195] In some examples, the at least one simulated calibration object comprises a three- dimensional object projected in two dimensions of the image. In some examples, the three-dimensional object is a sphere. In some examples, the three-dimensional object is a mirror sphere. In some examples, the three-dimensional object is a gray sphere.

[0196] In some examples, the lighting condition comprises light directionality. In some examples, the lighting condition comprises light intensity. In some examples, the lighting condition comprises light temperature.

[0197] In some examples, detecting the change comprises comparing one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization.

[0198] In some examples, the method further comprises deriving a camera condition of the new visualization based on the modified version of the at least one simulated calibration object.

[0199] In some examples, the camera condition comprises camera extrinsics. In some examples, the camera condition comprises camera intrinsics. In some examples, the camera intrinsics comprise image sensor properties.

[0200] In some examples, the method further comprises adjusting the new visualization based on the lighting condition.

[0201] In some examples, adjusting the new visualization comprises inserting anew object in the new visualization.

[0202] In some examples, inserting the new object comprises rendering the new object based on the lighting condition.

[0203] In some examples, the method further comprise deriving a lighting condition of the scene in the image based on the detected change.

[0204] In some examples, a system including one or more processors, and one or more non-transitory computer-readable storage media storing instructions, which when executed by the one or more processors, causes the system to carry out the method of any one of the preceding examples.

[0205] In some examples, one or more non-transitory computer-readable storage media carrying machine-readable instructions, which when executed by one or more processors of one or more machines, cause the one or more machines to carry out the method of any one of the preceding examples.

[0206] All of the processes described herein may be embodied in, and fully automated, via software code modules executed by a computing system that includes one or more computersor processors. The code modules may be stored in any type of non-transitory computer- readable medium or other computer storage device. Some or all the methods may be embodied in specialized computer hardware.

[0207] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence or can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, for example, through multi -threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.

[0208] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions. In some embodiments, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, one or more microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.

[0209] Conditional language such as. among others, "can," "could." "might" or "may," unless specifically stated otherwise, are understood within the context as used in general toconvey that certain embodiments include, while other embodiments do not include, certain features, elements and / or steps. Thus, such conditional language is not generally intended to imply that features, elements and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements and / or steps are included or are to be performed in any particular embodiment.

[0210] Disjunctive language such as the phrase "at least one of X, Y, or Z," unless specifically stated otherwise, is understood with the context as used in general to present that an item, term, etc., may be either X, Y, or Z, or any combination thereof (for example, X, Y, and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y, or at least one of Z to each be present.

[0211] Any process descriptions, elements or blocks in the flow diagrams described herein and / or depicted in the attached figures should be understood as potentially representing modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or elements in the process. Alternate implementations are included within the scope of the embodiments described herein in which elements or functions may be deleted, executed out of order from that shown, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved as would be understood by those skilled in the art.

[0212] Unless otherwise explicitly stated, articles such as "a" or "an" should generally be interpreted to include one or more described items. Accordingly, phrases such as "a device configured to" are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, "a processor configured to carry out recitations A, B and C" can include a first processor configured to carry out recitation A working in conjunction with a second processor configured to carry out recitations B and C.

[0213] The technology7as described herein may have also been described, at least in part, in terms of one or more embodiments, none of which is deemed exclusive to the other. Various configurations may omit, substitute, or add various procedures or components as appropriate. For instance, in alternative configurations, the methods may be performed in an order different from that described, or combined with other steps, or omitted altogether. This disclosure is further non-limiting and the examples and embodiments described herein does not limit thescope of the invention.

[0214] It should be emphasized that many variations and modifications may be made to the above-described embodiments, the elements of which are to be understood as being among other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure.

[0215] Although the present technology has been described in detail for the purpose of illustration based on what is currently considered to be the most practical and preferred implementations, it is to be understood that such detail is solely for that purpose and that the technology is not limited to the disclosed implementations, but. on the contrary , is intended to cover modifications and equivalent arrangements that are within the spirit and scope of the appended claims. For example, it is to be understood that the present technology contemplates that, to the extent possible, one or more features of any implementation can be combined with one or more features of any other implementation.

[0216] Furthermore, to the extent that the terms "includes," "including," "has," "contains," variants thereof, and other similar words are used in either the detailed description or the claims, these terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition word without precluding any additional or other elements.

Claims

CLAIMSWhat is claimed is:

1. A method of controlling color and light conditions of a generative machine learning model, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; providing the image as input to a generative machine learning model; controlling the generative machine learning model based on the derived lighting condition; and generating, using the controlled generative machine learning model, a first visualization comprising new visual data.

2. The method of claim 1. further comprising extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty' crop portion.

3. The method of claim 1, wherein the at least one simulated calibration object comprises a chromatic object.

4. The method of claim 3, wherein the chromatic object is a mirror sphere.

5. The method of claim 1, wherein the at least one simulated calibration object comprises a matte object.

6. The method of claim 5, wherein the matte object is a gray sphere.

7. The method of claim 1, wherein the at least one simulated calibration object is a color chart.

8. The method of claim 1, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

9. The method of claim 1, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

10. The method of claim 1, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

11. The method of claim 1, wherein the lighting condition comprises light directionality.

12. The method of claim 1, wherein the lighting condition comprises light intensity.

13. The method of claim 1, wherein the lighting condition comprises light temperature.

14. The method of claim 1, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

15. The method of claim 14, wherein the image is input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

16. The method of claim 1, further comprising deriving one or more parameters of one or more objects in the first visualization based on the derived lighting condition.

17. The method of claim 16, wherein the one or more parameters of the one or more obj ects in the first visualization comprise color information.

18. The method of claim 16, wherein the one or more parameters of the one or more objects in the first visualization comprise material information.

19. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method of controlling color and light conditions of a generative machine learning model, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; providing the image as input to a generative machine learning model; controlling the generative machine learning model based on the derived lighting condition; and generating, using the controlled generative machine learning model, a first visualization comprising new visual data.

20. The one or more non-transitory computer-readable media of claim 19, further comprising extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

21. The one or more non-transitory computer-readable media of claim 19, wherein the at least one simulated calibration object comprises a chromatic object.

22. The one or more non-transitory computer-readable media of claim 21. wherein the chromatic object is a mirror sphere.

23. The one or more non-transitory computer-readable media of claim 19, wherein the at least one simulated calibration object comprises a matte object.

24. The one or more non-transitory computer-readable media of claim 23, wherein the matte object is a gray sphere.

25. The one or more non-transitory computer-readable media of claim 19, wherein the at least one simulated calibration object is a color chart.

26. The one or more non-transitory computer-readable media of claim 19, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

27. The one or more non-transitory computer-readable media of claim 19, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

28. The one or more non-transitory computer-readable media of claim 19, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

29. The one or more non-transitory computer-readable media of claim 19, wherein the lighting condition comprises light directionality .

30. The one or more non-transitory computer-readable media of claim 19. wherein the lighting condition comprises light intensity.

31. The one or more non-transitor ' computer-readable media of claim 19, wherein the lighting condition comprises light temperature.

32. The one or more non-transitory computer-readable media of claim 19, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

33. The one or more non-transitory7computer-readable media of claim 32, wherein the image is input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

34. The one or more non-transitory computer-readable media of claim 19, further comprising deriving one or more parameters of one or more objects in the first visualization based on the derived lighting condition.

35. The one or more non-transitory7computer-readable media of claim 34, wherein the one or more parameters of the one or more objects in the first visualization comprise color information.

36. The one or more non-transitory computer-readable media of claim 34, wherein the one or more parameters of the one or more objects in the first visualization comprise material information.

37. A system for controlling color and light conditions of a generative machine learning model, the system comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the system to: provide an image; insert at least one simulated calibration object into a portion of a scene in the image; derive a lighting condition of the scene in the image based on the at least one simulated calibration object; provide the image as input to a generative machine learning model; control the generative machine learning model based on the derived lighting condition; and generate, using the controlled generative machine learning model, a first visualization comprising new visual data.

38. The system of claim 37, wherein the instructions further cause the system to extend a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

39. The system of claim 37, wherein the at least one simulated calibration object comprises a chromatic object.

40. The system of claim 39, wherein the chromatic object is a mirror sphere.

41. The system of claim 37, wherein the at least one simulated calibration object comprises a matte object.

42. The system of claim 41, wherein the matte object is a gray sphere.

43. The system of claim 37, wherein the at least one simulated calibration object is a color chart.

44. The system of claim 37, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

45. The system of claim 37, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

46. The system of claim 37, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

47. The system of claim 37, wherein the lighting condition comprises light directionality.

48. The system of claim 37, wherein the lighting condition comprises light intensity.

49. The system of claim 37, wherein the lighting condition comprises light temperature.

50. The system of claim 37, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

51. The system of claim 50, wherein the image is input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

52. The system of claim 37, wherein the instructions further cause the system to derive one or more parameters of one or more obj ects in the first visualization based on the derived lighting condition.

53. The system of claim 52, wherein the one or more parameters of the one or more objects in the first visualization comprise color information.

54. The system of claim 52, wherein the one or more parameters of the one or more objects in the first visualization comprise material information.

55. A method of controlling color and lighting conditions of a generative machine learning model, the method comprising: providing an image: inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; generating a view of a 3D object within a 3D environment comprising a representation of the derived lighting condition; providing the image and the view of the 3D object as inputs to a generative machine learning model; and generating, using the generative machine learning model, a first visualization comprising new visual data.

56. The method of claim 55, further comprising extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into theportion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

57. The method of claim 55, wherein the at least one simulated calibration object comprises a chromatic object.

58. The method of claim 57, wherein the chromatic object is a mirror sphere.

59. The method of claim 55, wherein the at least one simulated calibration object comprises a matte object.

60. The method of claim 59, wherein the matte object is a gray sphere.

61. The method of claim 55, wherein the at least one simulated calibration object is a color chart.

62. The method of claim 55, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

63. The method of claim 55, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

64. The method of claim 55, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

65. The method of claim 55, wherein the lighting condition comprises light directionality.

66. The method of claim 55. wherein the lighting condition comprises light intensity.

67. The method of claim 55, wherein the lighting condition comprises light temperature.

68. The method of claim 55, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

69. The method of claim 68, wherein the image and the view of the 3D object are input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

70. The method of claim 63, wherein the 3D environment comprising the representation of the derived lighting condition comprises a representation of the equirectangular map.

71. The method of claim 70. wherein the representation of the equirectangular map is at a distance from a position of the 3D object.

72. The method of claim 71, wherein the distance is substantially infinity.

73. The method of claim 55, wherein the view of the 3D object within the 3D environment is from a camera pose associated with the image.

74. The method of claim 55, wherein the view of the 3D object is a 3D view.

75. The method of claim 55, wherein the view of the 3D object is a 2D view.

76. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method of controlling color and lighting conditions of a generative machine learning model, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; deriving a lighting condition of the scene in the image based on the at least one simulated calibration object; generating a view of a 3D object within a 3D environment comprising a representation of the derived lighting condition; providing the image and the view of the 3D object as inputs to a generative machine learning model; and generating, using the generative machine learning model, a first visualization comprising new visual data.

77. The one or more non-transitory computer-readable media of claim 76, further comprising extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty7crop portion.

78. The one or more non-transitory computer-readable media of claim 76, wherein the at least one simulated calibration object comprises a chromatic object.

79. The one or more non-transitory' computer-readable media of claim 78, wherein the chromatic object is a mirror sphere.

80. The one or more non-transitory computer-readable media of claim 76, wherein the at least one simulated calibration object comprises a matte object.

81. The one or more non-transitory computer-readable media of claim 80, wherein the matte object is a gray sphere.

82. The one or more non-transitory7computer-readable media of claim 76, wherein the at least one simulated calibration object is a color chart.

83. The one or more non-transitory computer-readable media of claim 76, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

84. The one or more non-transitory computer-readable media of claim 76, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

85. The one or more non-transitory computer-readable media of claim 76, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

86. The one or more non-transitory computer-readable media of claim 76, wherein the lighting condition comprises light directionality.

87. The one or more non-transitory computer-readable media of claim 76. wherein the lighting condition comprises light intensity.

88. The one or more non-transitory computer-readable media of claim 76, wherein the lighting condition comprises light temperature.

89. The one or more non-transitory computer-readable media of claim 76, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

90. The one or more non-transitory computer-readable media of claim 89, wherein the image and the view of the 3D object are input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

91. The one or more non-transitory computer-readable media of claim 84, wherein the 3D environment comprising the representation of the derived lighting condition comprises a representation of the equirectangular map.

92. The one or more non-transitory computer-readable media of claim 91. wherein the representation of the equirectangular map is at a distance from a position of the 3D object.

93. The one or more non-transitory' computer-readable media of claim 92, wherein the distance is substantially infinity.

94. The one or more non-transitory computer-readable media of claim 76, wherein the view of the 3D object within the 3D environment is from a camera pose associated with the image.

95. The one or more non-transitory' computer-readable media of claim 76, wherein the view of the 3D object is a 3D view.

96. The one or more non-transitory computer-readable media of claim 76, wherein the view of the 3D object is a 2D view.

97. A system for controlling color and lighting conditions of a generative machine learning model, the system comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the system to: provide an image;insert at least one simulated calibration object into a portion of a scene in the image; derive a lighting condition of the scene in the image based on the at least one simulated calibration object; generate a view of a 3D object within a 3D environment comprising a representation of the derived lighting condition; provide the image and the view of the 3D object as inputs to a generative machine learning model; and generate, using the generative machine learning model, a first visualization comprising new visual data.

98. The system of claim 97, wherein the instructions further cause the system to extend a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration obj ect into the empty crop portion.

99. The system of claim 97, wherein the at least one simulated calibration object comprises a chromatic object.

100. The system of claim 99, wherein the chromatic object is a mirror sphere.

101. The system of claim 97, wherein the at least one simulated calibration obj ect comprises a matte object.

102. The system of claim 101, wherein the matte object is a gray sphere.

103. The system of claim 97, wherein the at least one simulated calibration object is a color chart.

104. The system of claim 97, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inpainting the at least one simulated calibration object into the portion of the scene in the image.

105. The system of claim 97, wherein deriving the lighting condition comprises unwrapping the at least one simulated calibration object to generate an equirectangular map, wherein the lighting condition is based on the equirectangular map.

106. The system of claim 97, wherein deriving the lighting condition comprises applying a high dynamic range image map onto the at least one simulated calibration object.

107. The system of claim 97, wherein the lighting condition comprises light directionality.

108. The system of claim 97, wherein the lighting condition comprises light intensity.

109. The system of claim 97, wherein the lighting condition comprises light temperature.

110. The system of claim 97, wherein the generative machine learning model comprises one or more stable diffusion models modulated by one or more control networks.

111. The system of claim 110, wherein the image and the view of the 3D object are input into one or more of the one or more stable diffusion models, and wherein the derived lighting condition is input to one or more of the one or more control networks.

112. The system of claim 105, wherein the 3D environment comprising the representation of the derived lighting condition comprises a representation of the equirectangular map.

113. The system of claim 112, wherein the representation of the equirectangular map is at a distance from a position of the 3D object.

114. The system of claim 113, wherein the distance is substantially infinity.

115. The system of claim 97, wherein the view of the 3D object within the 3D environment is from a camera pose associated with the image.

116. The system of claim 97, wherein the view of the 3D object is a 3D view.

117. The system of claim 97, wherein the view of the 3D object is a 2D view.

118. A method of deriving lighting conditions, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; rendering a new visualization of the scene in the image including a modified version of the at least one simulated calibration object; detecting a change between the at least one simulated calibration object in the image and the modified version of the at least one simulated calibration object in the new visualization; and deriving a lighting condition of the new visualization based on the detected change.

119. The method of claim 118. further comprising extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

120. The method of claim 118, wherein the new visualization is rendered with generative machine learning.

121. The method of claim 120, wherein generative machine learning includes stable diffusion.

122. The method of claim 118, wherein the at least one simulated calibration object comprises a chromatic object.

123. The method of claim 122, wherein the chromatic object is a mirror sphere.

124. The method of claim 118, wherein the at least one simulated calibration object comprises a matte object.

125. The method of claim 124, wherein the matte object is a gray sphere.

126. The method of claim 118, wherein the at least one simulated calibration object is a color chart.

127. The method of claim 118, wherein the at least one simulated calibration object comprises a three-dimensional object projected in two dimensions of the image.

128. The method of claim 127, wherein the three-dimensional object is a sphere.

129. The method of claim 128, wherein the three-dimensional object is a mirror sphere.

130. The method of claim 128, wherein the three-dimensional object is a gray sphere.

131. The method of claim 118, wherein the lighting condition comprises light directionality'.

132. The method of claim 118, wherein the lighting condition comprises light intensity'.

133. The method of claim 118, wherein the lighting condition comprises light temperature.

134. The method of claim 118, wherein detecting the change comprises comparing one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization.

135. The method of claim 118. further comprising deriving a camera condition of the new visualization based on the modified version of the at least one simulated calibration object.

136. The method of claim 135, wherein the camera condition comprises camera extrinsics.

137. The method of claim 135, wherein the camera condition comprises camera intrinsics.

138. The method of claim 137, wherein the camera intrinsics comprise image sensor properties.

139. The method of claim 118, further comprising adjusting the new visualization based on the lighting condition.

140. The method of claim 139, wherein adjusting the new visualization comprises inserting a new object in the new visualization.

141. The method of claim 140, wherein inserting the new object comprises rendering the new object based on the lighting condition.

142. The method of claim 118, further comprising deriving a lighting condition of the scene in the image based on the detected change.

143. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method of deriving lighting conditions, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image;rendering a new visualization of the scene in the image including a modified version of the at least one simulated calibration object; detecting a change between the at least one simulated calibration object in the image and the modified version of the at least one simulated calibration object in the new visualization; and deriving a lighting condition of the new visualization based on the detected change.

144. The one or more non-transitory computer-readable media of claim 143, further comprising extending a boun ary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

145. The one or more non-transitory computer-readable media of claim 143, wherein the new visualization is rendered with generative machine learning.

146. The one or more non-transitory computer-readable media of claim 145, wherein generative machine learning includes stable diffusion.

147. The one or more non-transitory computer-readable media of claim 143. wherein the at least one simulated calibration object comprises a chromatic object.

148. The one or more non-transitory computer-readable media of claim 147, wherein the chromatic object is a mirror sphere.

149. The one or more non-transitory computer-readable media of claim 143. wherein the at least one simulated calibration object comprises a matte object.

150. The one or more non-transitory computer-readable media of claim 149, wherein the matte object is a gray sphere.

151. The one or more non-transitory computer-readable media of claim 143. wherein the at least one simulated calibration object is a color chart.

152. The one or more non-transitory computer-readable media of claim 143, wherein the at least one simulated calibration object comprises a three-dimensional object projected in two dimensions of the image.

153. The one or more non-transitory computer-readable media of claim 152, wherein the three-dimensional object is a sphere.

154. The one or more non-transitory computer-readable media of claim 153, wherein the three-dimensional object is a mirror sphere.

155. The one or more non-transitory computer-readable media of claim 153, wherein the three-dimensional object is a gray sphere.

156. The one or more non-transitory computer-readable media of claim 143, wherein the lighting condition comprises light directionality.

157. The one or more non-transitory computer-readable media of claim 143. wherein the lighting condition comprises light intensity.

158. The one or more non-transitory computer-readable media of claim 143. wherein the lighting condition comprises light temperature.

159. The one or more non-transitory computer-readable media of claim 143, wherein detecting the change comprises comparing one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization.

160. The one or more non-transitory computer-readable media of claim 143, further comprising deriving a camera condition of the new visualization based on the modified version of the at least one simulated calibration object.

161. The one or more non-transitory computer-readable media of claim 160, wherein the camera condition comprises camera extrinsics.

162. The one or more non-transitory computer-readable media of claim 160. wherein the camera condition comprises camera intrinsics.

163. The one or more non-transitory computer-readable media of claim 162. wherein the camera intrinsics comprise image sensor properties.

164. The one or more non-transitory computer-readable media of claim 143, further comprising adjusting the new visualization based on the lighting condition.

165. The one or more non-transitory computer-readable media of claim 164. wherein adjusting the new visualization comprises inserting a new object in the new visualization.

166. The one or more non-transitory computer-readable media of claim 165, wherein inserting the new object comprises rendering the new object based on the lighting condition.

167. The one or more non-transitory computer-readable media of claim 143, further comprising deriving a lighting condition of the scene in the image based on the detected change.

168. A system for deriving lighting conditions, the system comprising: one or more processors; and memory’ storing instructions that, when executed by the one or more processors, cause the system to perform a method of deriving lighting conditions, the method comprising: providing an image; inserting at least one simulated calibration object into a portion of a scene in the image; rendering a new visualization of the scene in the image including a modified version of the at least one simulated calibration object;detecting a change between the at least one simulated calibration object in the image and the modified version of the at least one simulated calibration object in the new visualization; and deriving a lighting condition of the new visualization based on the detected change.

169. The system of claim 168, wherein the method further comprises extending a boundary of the image to create an empty crop portion, wherein inserting the at least one simulated calibration object into the portion of the scene in the image comprises inserting the at least one simulated calibration object into the empty crop portion.

170. The system of claim 168, wherein the new visualization is rendered with generative machine learning.

171. The system of claim 170, wherein generative machine learning includes stable diffusion.

172. The system of claim 168, wherein the at least one simulated calibration object comprises a chromatic object.

173. The system of claim 172, wherein the chromatic object is a mirror sphere.

174. The system of claim 168, wherein the at least one simulated calibration object comprises a matte object.

175. The system of claim 174, wherein the matte object is a gray sphere.

176. The system of claim 168. wherein the at least one simulated calibration object is a color chart.

177. The system of claim 168, wherein the at least one simulated calibration object comprises a three-dimensional object projected in two dimensions of the image.

178. The system of claim 177, wherein the three-dimensional object is a sphere.

179. The system of claim 178, wherein the three-dimensional object is a mirror sphere.

180. The system of claim 178, wherein the three-dimensional object is a gray sphere.

181. The system of claim 168, wherein the lighting condition comprises light directionality.

182. The system of claim 168, wherein the lighting condition comprises light intensity.

183. The system of claim 168, wherein the lighting condition comprises light temperature.

184. The system of claim 168, wherein detecting the change comprises comparing one or more pixels of the at least one simulated calibration object in the image to corresponding one or more pixels of the modified version of the at least one simulated calibration object in the new visualization.

185. The system of claim 168, wherein the method further comprises deriving a camera condition of the new visualization based on the modified version of the at least one simulated calibration object.

186. The system of claim 185, wherein the camera condition comprises camera extrinsics.

187. The system of claim 185, wherein the camera condition comprises camera intrinsics.

188. The system of claim 187, wherein the camera intrinsics comprise image sensor properties.

189. The system of claim 168, wherein the method further comprises adjusting the new visualization based on the lighting condition.

190. The system of claim 189, wherein adjusting the new visualization comprises inserting a new object in the new visualization.

191. The system of claim 190, wherein inserting the new obj ect comprises rendering the new object based on the lighting condition.

192. The system of claim 168, wherein the method further comprises deriving a lighting condition of the scene in the image based on the detected change.