Relighting outdoor images using machine learning

The method addresses unrealistic outdoor scenes by segmenting and modifying lighting using a diffusion model, ensuring accurate and high-quality adjustments in outdoor images.

JP7785236B2Active Publication Date: 2025-12-12GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025503429
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-05-09
Filing Date
2024-05-08
Publication Date
2025-12-12
Estimated Expiration
2044-05-08

AI Technical Summary

Technical Problem

Machine learning models often generate unrealistic outdoor scenes, particularly with improper representation of lighting conditions and details of subjects like people, leading to inconsistencies such as indoor-like appearances under outdoor backgrounds.

Method used

A method utilizing a diffusion model to modify the lighting of an input image by segmenting sky and object areas, applying local color transformations, and fusing the images while preserving subject and sky details, using bilateral grid upsampling and super-resolution techniques to enhance image quality.

Benefits of technology

The method effectively adjusts lighting conditions in outdoor scenes, ensuring realistic and high-quality output images that accurately reflect the requested lighting changes, such as transitioning from sunny to moonlit environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785236000001
    Figure 0007785236000001
  • Figure 0007785236000002
    Figure 0007785236000002
  • Figure 0007785236000003
    Figure 0007785236000003
Patent Text Reader

Abstract

The media application provides an initial image and a request to modify the lighting in the initial image as input to the diffusion model, the initial image including an object and a sky. The media application uses the diffusion model to output an output image that meets the request. The media application determines sky segments and object segments from the initial image. The media application generates a sky mask corresponding to the sky segment and an object mask corresponding to the object segment. The media application modifies the coloring of the initial image to match that of the output image. The media application fuses the modified initial image with the output image to form a fused image, using the object mask to prevent modification of the object from the modified initial image and the sky mask to prevent modification of the sky from the output image during fusing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 465,224, entitled "Relighting of Outdoor Images Using Machine Learning," filed May 9, 2023, which is incorporated herein in its entirety. [Background technology]

[0002] While machine learning models can generate outdoor scenes, the scenes are often unrealistic. For example, machine learning models can generate nighttime scenes with shadows. Furthermore, machine learning models can generate unrealistic outdoor scenes of people, where the finer details of the people may be improperly represented. For example, people may appear to have been photographed indoors while the background coincides with a sunset.

[0003] The description of the background art provided herein is intended to provide a general context for the present disclosure. To the extent described in this background art section, the work of the currently named inventors, and aspects of this specification that may not otherwise qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention

[0004] The computer-implemented method includes providing an initial image and a request to modify the lighting in the initial image as input to a diffusion model, the initial image including an object and a sky. The method further includes using the diffusion model to output an output image that meets the request. The method further includes determining a sky segment and an object segment from the initial image. The method further includes generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment. The method further includes modifying a coloration of the initial image to match a coloration of the output image. The method further includes fusing the modified initial image with the output image to form a fused image, using the object mask to prevent modification of the object from the modified initial image during fusing and the sky mask to prevent modification of the sky from the output image.

[0005] In some embodiments, modifying the coloring of the initial image includes performing bilateral grid upsampling (BGU) to identify local color transformations between the initial image and the output image and applying the local color transformations to the initial image while preventing sky coloration using a sky mask. In some embodiments, the method further includes generating a super-resolved version of at least a portion of the output image from the output image, and fusing the modified initial image with the output image includes fusing the super-resolved version of at least a portion of the output image while using an object mask to prevent object modifications from the modified initial image and using a sky mask to prevent sky modifications from the super-resolved version of at least a portion of the output image during fusing. In some embodiments, the output image includes one or more shadows corresponding to one or more objects in the output image, and the method further includes determining, from the output image, shadow segments corresponding to the one or more shadows in the output image and generating a shadow mask corresponding to the shadow segments, and fusing the output image with the modified initial image includes using the shadow mask to prevent the one or more shadow modifications from the output image during fusing.

[0006] In some embodiments, the request to change the lighting includes the user providing a text request including attributes selected from the group of light level, amount of cloud cover in the sky, sky color, and combinations thereof. In some embodiments, the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, a global preset, a menu of options, a library of pre-made text requests, and combinations thereof. In some embodiments, the method further includes determining that the initial image includes an outdoor scene and providing the user with suggestions to modify the lighting.

[0007] A non-transitory computer-readable medium is provided having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations including providing an initial image and a request to modify illumination in the initial image as input to a diffusion model, the initial image including an object and a sky, using the diffusion model to output an output image that satisfies the request, determining sky segments and object segments from the initial image, generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment, modifying coloring of the initial image to match that of the output image, and fusing the modified initial image with the output image to form a fused image, using the object mask to prevent modification of the object from the modified initial image and the sky mask to prevent modification of the sky from the output image during fusing.

[0008] In some embodiments, modifying the coloring of the initial image includes performing a BGU that identifies local color transformations between the initial image and the output image and applying the local color transformations to the initial image while preventing sky coloration using a sky mask. In some embodiments, the operations further include generating a super-resolution version of at least a portion of the output image from the output image, and fusing the modified initial image with the output image includes fusing the super-resolution version of at least a portion of the output image while using an object mask to prevent object modifications from the modified initial image and using a sky mask to prevent sky modifications from the super-resolution version of at least a portion of the output image during fusing. In some embodiments, the output image includes one or more shadows corresponding to one or more objects in the output image, and the operations further include determining, from the output image, shadow segments corresponding to the one or more shadows in the output image and generating a shadow mask corresponding to the shadow segments, and fusing the output image with the modified initial image includes using the shadow mask to prevent the one or more shadow modifications from the output image during fusing. In some embodiments, the request to change the lighting includes the user providing a text request including attributes selected from the group of light level, amount of clouds in the sky, sky color, and combinations thereof. In some embodiments, the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, a global preset, a menu of options, a library of pre-made text requests, and combinations thereof. In some embodiments, the operation further includes determining that the initial image includes an outdoor scene and providing the user with suggestions to modify the lighting.

[0009] The system includes a processor and a memory coupled to the processor having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including providing an initial image and a request to modify illumination in the initial image as input to a diffusion model, the initial image including an object and a sky, using the diffusion model to output an output image that satisfies the request, determining a sky segment and an object segment from the initial image, generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment, modifying a coloration of the initial image to match a coloration of the output image, and fusing the modified initial image with the output image to form a fused image, using the object mask to prevent modification of the object from the modified initial image and the sky mask to prevent modification of the sky from the output image during fusing.

[0010] In some embodiments, modifying the coloring of the initial image includes performing a BGU that identifies local color transformations between the initial image and the output image and applying the local color transformations to the initial image while preventing sky coloration using a sky mask. In some embodiments, the operations further include generating a super-resolution version of at least a portion of the output image from the output image, and fusing the modified initial image with the output image includes fusing the super-resolution version of at least a portion of the output image while using an object mask to prevent object modifications from the modified initial image and using a sky mask to prevent sky modifications from the super-resolution version of at least a portion of the output image during fusing. In some embodiments, the output image includes one or more shadows corresponding to one or more objects in the output image, and the operations further include determining, from the output image, shadow segments corresponding to the one or more shadows in the output image and generating a shadow mask corresponding to the shadow segments, and fusing the output image with the modified initial image includes using the shadow mask to prevent the one or more shadow modifications from the output image during fusing. In some embodiments, the request to change the lighting includes a user providing a text request including attributes selected from the group of light level, amount of clouds in the sky, sky color, and combinations thereof. In some embodiments, the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, global presets, a menu of options, a library of pre-made text requests, and combinations thereof. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3]FIG. 1 illustrates an exemplary user interface including options for selecting different aspects of an image to modify, a field for providing text, and an exemplary output image, according to some embodiments described herein. [Figure 4A] 4A and 4B illustrate an example user interface including options for selecting different types of modifications for changing the lighting of an image, according to some embodiments described herein. [Figure 4B] 4A and 4B illustrate an example user interface including options for selecting different types of modifications for changing the lighting of an image, according to some embodiments described herein. [Figure 5] FIG. 1 is a block diagram of an example architecture for generating a fused image incorporating a text request, according to some embodiments described herein. [Figure 6] 1 is an exemplary flowchart of a method for generating an illumination-corrected fused image according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0012] While machine learning models can generate outdoor scenes, these scenes are often unrealistic. For example, machine learning models may generate nighttime scenes with shadows. Furthermore, machine learning models may generate unrealistic outdoor scenes with people, in which the finer details of people may be improperly represented. For example, people may appear to have been photographed indoors while the background coincides with a sunset.

[0013] The techniques described below describe media applications that advantageously modify the lighting of an input image by using a diffusion model to output an output image that meets the lighting modification request. For example, a user may provide a text request to change an image of a person captured outdoors on a sunny day to an image of the person under a moonlit sky.

[0014] The media application generates an output image. For example, the media application may generate a synthetic moonlit sky using a diffusion model. The media application determines sky segments and object segments from an initial image. The media application generates object masks corresponding to the object segments and sky masks corresponding to the sky segments.

[0015] The media application may modify the coloring of the initial image to match that of the output image. The coloring ensures that the coloring of the subject matches the modifications to the output image. For example, replacing a sunny image with a moonlit image will change the color of a person from including all colors to including primarily blue, purple, or black hues. The media application may modify the coloring of the initial image by performing bilateral grid upsampling (BGU) to identify local color transformations between the initial image and the output image. The media application may also generate a super-resolution version of at least a portion of the output image (e.g., a composite sky portion of the output image) from the output image. The super-resolution version of the output image advantageously extracts more detail from the low-resolution output image to improve the quality of the output image.

[0016] The media application fuses the modified initial image (i.e., the initial image modified to match the coloring of the output image) with a super-resolved version of at least a portion of the output image to form a fused image, using a subject mask to prevent subject modifications from the modified image and a sky mask to prevent sky modifications from the super-resolved version during fusion.

[0017] Exemplary Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, a user device 115a, and a user device 115n coupled to a network 105. Users 125a, 125n may be associated with each user device 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," denotes a reference to the element with that specific reference number. A reference number in text without a following letter, e.g., "115," denotes a general reference to an embodiment of the element bearing that reference number.

[0018] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits and receives data to and from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0019] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, the user preferences of the user 125, etc.

[0020] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0021] In the illustrated implementation, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or as media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. The user devices 115a and 115n in FIG. 1 are used as an example. While FIG. 1 shows two user devices 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.

[0022] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are executed on the media server 101 or the user device 115. In some embodiments, some of the operations may be executed on the media server 101 and some may be executed on the user device 115. Execution of the operations is subject to user settings. For example, the user 125a may specify that operations be executed on their corresponding device 115a and not on the media server 101. Such settings result in the operations described herein being executed entirely on the user device 115a and not on the media server 101. Furthermore, the user 125a may specify that the user's images and / or other data be stored locally only on the user device 115a and not on the media server 101. Such settings result in user data not being sent to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of operations on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of operations by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable use of the media server 101.

[0023] Machine learning models (e.g., neural networks or other types of models) are stored locally on the user device 115 and utilized for one or more operations with specific user permission. Server-side models are used only with user permission. Additionally, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed if authorized by the user 125. Updated model parameters may be sent to the media server 101 if authorized by the user 125, for example, to enable federated learning. The model parameters do not include any user data.

[0024] The media application 103 receives a request to change the lighting in an initial image. The request may include a text request, a selection of a suggestion, a selection of a preset, a selection of an option from a library of pre-made text requests, etc. The initial image includes a subject.

[0025] The media application 103 provides an initial image and a request as input to the diffusion model. The diffusion model outputs an output image that satisfies the request by including the features described in the request. For example, if the request calls for changing the input image from a rainy image with an overcast sky to a clear sky with a sunset, the output image includes a clear sky with a sunset. Thus, the output image corresponds to the initial image and represents a corrected initial image, where the correction of the initial image is performed in response to the request, i.e., in response to the changes requested in the request. Specifically, the output image is the initial image corrected by implementing the changes indicated in the request (e.g., changing the lighting in the initial image).

[0026] The media application 103 determines sky segments and object segments from the initial image. The media application 103 generates a sky mask corresponding to the sky segments and an object mask corresponding to the object segments. The media application 103 modifies the coloring of the initial image to match the coloring of the output image. The media application 103 fuses the modified initial image with the output image to form a fused image, using the object mask to prevent object modifications from the modified image and the sky mask to prevent sky modifications from the output image during fusing.

[0027] In some embodiments, media application 103 may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.

[0028] Exemplary Computing Device 200 2 is a block diagram of an exemplary computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to implement a media application 103a. In another example, the computing device 200 is a user device 115.

[0029] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0030] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuits for achieving functionality, dedicated processors for performing neural network model-based processing, neural circuits, systems with processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to produce a probabilistic output; for example, the output produced by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor in communication with a memory.

[0031] Memory 237 is typically provided in computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 235, suitable for storing instructions for execution by the processor or set of processors. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.

[0032] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app”) running on a mobile computing device, etc.

[0033] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), and the like.

[0034] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices for input devices (keyboards, pointing devices, touchscreens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).

[0035] Some examples of interfaced devices that can be connected to I / O interface 239 may include display 241, which can be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. For example, display 241 may be utilized to display a user interface, including a graphical guide, on a viewfinder. Display 241 can include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 can be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass form factor or headset device, or a monitor screen of a computing device.

[0036] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0037] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store training data sets including labeled images, machine learning models, output from machine learning models, etc.

[0038] FIG. 2 illustrates an exemplary media application 103 stored in memory 237 , including a user interface module 202 , a segmenter 204 , a diffusion module 206 , a resolution module 208 , a colorization module 210 , and a fusion module 212 .

[0039] The user interface module 202 generates graphical data for displaying a user interface including an image. In some embodiments, the user interface module 202 receives an initial image. The initial image may be received from the camera 243 of the computing device 200 or from the media server 101 via the I / O interface 239. The initial image includes an object, such as a person, an animal, or a tree. The object may include multiple objects, such as a person and a dog, a series of buildings, etc.

[0040] The user interface module 202 includes options for providing a request associated with the image. In some embodiments, the request is a text request received from a user. The user interface module 202 may include a text box in which the user enters the text request. For example, the text request may include a request to change the light level (e.g., brighten the sky, change the sky to a sunset, darken the sky, create a moonlit night, etc.), the amount of clouds in the sky (e.g., make the sky cloudy, remove clouds, indicate precipitation, etc.), and / or the color of the sky (e.g., add purple and blue, include light orange and dark yellow, create a rainbow in the sky, etc.).

[0041] In some embodiments, the user interface module 202 may provide the user with suggestions or presets that form part of a request to change the lighting. The suggestions may include region suggestions associated with one or more regions of the initial image. For example, the initial image may include different regions with different attributes, such as sky, clouds in the sky, horizon, water, etc.

[0042] The suggestions or presets may also include global presets (e.g., changes that affect the entire image), a menu of options (e.g., changes to different parts of the image, different subjects, etc.), and / or a library of pre-made text requests (e.g., sky options, golden hour options, etc.).

[0043] In some embodiments, the user interface module 202 may determine that the initial image includes an outdoor scene. For example, object recognition may be performed on the initial image to determine whether the initial image includes outdoor objects. The user interface module 202 provides suggestions to the user for modifying the lighting. The suggestions may take the form of a button that the user can select to generate a list of options, a menu of suggestions, a text field that appears where the user can directly enter their request, etc.

[0044] 3 shows exemplary user interfaces 300, 325, 350 including options for selecting different aspects of lighting to modify, fields for providing text, and exemplary output images, according to some embodiments described herein. Specifically, first user interface 300 automatically provides a global preset 305 for the user to select to modify lighting to appear as a sunset, nighttime, an overcast sky, etc. First user interface 300 also includes circles 310, 315, 320 representing region suggestions over different areas, allowing the user to specify modifications to be made only to the sky, fog, and water, respectively. For example, the user may tap circle 310 to modify the sky, circle 315 to modify the fog, or circle 320 to modify the water.

[0045] The second user interface 325 includes a text entry field 330 in which the user can specify the changes they want to make. The user can include a description that is specific enough to encompass the object they want to change (e.g., change water to make it smoother), or they can describe the specific changes to be made after selecting the object in the second user interface 325 that they want to change.

[0046] A third user interface 350 includes an output image in which the text request "make it almost night" has been fulfilled. The resulting image has both a darker background and a darker subject 355 because the subject 355 has been modified to appear consistent with the modified lighting.

[0047] The user interface module 202 generates graphical data for displaying the output image. In some embodiments, the user interface may also include options for editing the output image, sharing the output image, adding the output image to a photo album, etc. In some embodiments, the output image is marked to identify that artificial intelligence was used to generate the image.

[0048] 4A and 4B show example user interfaces 400, 425, 450 including options for selecting different types of modifications to change the lighting of an image, according to some embodiments described herein. A first user interface 400 includes options for modifying different aspects of an input image 402. The user interface module 202 provides options for generating different skies 405 or golden hours 410. Other variations are possible.

[0049] The second user interface 425 shows an output image 427 generated by the fusion module 212 in response to a user selecting the sky 405 option in the first user interface 400. The sky 405 options include different types of sky with variations in clouds, blue hues, light levels, etc. The user interface module 202 provides the user with one or more output images from which to select. The user may also select a button 430 to obtain a new set of results with different types of sky for the output image.

[0050] The third user interface 450 shows an output image 452 generated by the fusion module 212 in response to a user selecting the golden hour 410 option in the first user interface 400. The golden hour 410 option includes a daytime image just after sunrise or before sunset, when the sunlight is redder and softer than when the sun is higher in the sky. The user interface module 202 provides the user with one or more output images from which to choose. The user may also select a button 430 to obtain a new set of results of different types of golden hour output images.

[0051] In some embodiments, the user interface module 202 generates a user interface that includes options for modifying user preferences. For example, the user interface may include user preferences for specifying a level of probability, including the amount of noise the user wants to see in the output image (e.g., the degree to which the output image differs from the initial image) and the degree to which seeds are used in the output image (e.g., the degree to which the output image differs from the image captured by the camera). The level of probability and the degree to which seeds are used may be represented using radio buttons for different levels (e.g., low, medium, high), sliders for scales, text boxes for percentages, or other options.

[0052] The segmenter 204 segments the initial image. In some embodiments, the segmenter 204 determines sky segments and object segments. The segmenter 204 may also generate shadow segments corresponding to one or more shadows added to the output image that correspond to objects in the output image. The sky segments include pixels corresponding to sky locations in the initial image. The object segments include pixels corresponding to objects, where the object may be a person, a dog, a building, etc. The shadow segments include pixels associated with shadows in the output image. Additional segmentation may be applied, such as by generating one or more foreground segments, product segments (such as fabric, shoes, bags, etc.), power line segments, skin segments, etc. In some embodiments, the segmenter 204 generates a segmentation map that associates identification with each pixel in the initial image as belonging to the sky, one or more objects, shadow segments, etc.

[0053] Segmenter 204 may perform segmentation by performing object recognition on the initial image to identify objects in the initial image. For example, segmenter 204 may compare the sky and one or more objects in the initial image with prior information about objects, such as sky, people, shadows, vehicles, and buildings, to identify the expected shapes of objects and determine whether a pixel is associated with sky or an object in the initial image or a shadow in the output image. In some embodiments, segmenter 204 may divide the image into foreground and background before performing object recognition to aid in identifying the sky and the object, since sky is located in the background and objects are located in the foreground. Segmenter 204 may associate segments with locations in the initial image, such as a bounding box having an x-coordinate, a y-coordinate, and a scale, or the coordinates of the pixel associated with the segment.

[0054] The segmenter 204 generates one or more preservation masks. The segmenter 204 generates a sky mask that serves as a preservation mask corresponding to a sky segment. For example, the sky mask includes pixels that correspond to pixels of the sky segment in the initial image. The segmenter 204 generates an object mask that serves as a preservation mask corresponding to an object segment. For example, the object mask includes pixels that correspond to pixels of the object segment in the initial image. In some embodiments, the segmenter 204 generates a shadow mask that serves as a preservation mask corresponding to shadows in the output image.

[0055] In some embodiments, the storage mask is generated based on generating superpixels of the image and matching the centers of the superpixels to depth map values ​​(e.g., obtained by the camera 243 using a depth sensor or by deriving depth from pixel values) for depth-based cluster detection. More specifically, the depth values ​​of the masked area may be used to determine a depth range, and superpixels that fall within the depth range may be identified.

[0056] Another technique for generating the mask involves weighting the depth values ​​based on how close the depth values ​​are to a stored mask, where the weights are represented by a distance transform map.

[0057] In some embodiments, segmenter 204 uses a machine learning algorithm, such as a neural network, to segment the initial image and generate a storage mask. In some embodiments, segmenter 204 may specify circuitry (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that allows processor 235 to apply the machine learning model. In some embodiments, segmenter 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, segmenter 204 may provide an application programming interface (API) that can be used by operating system 262 and / or other applications 264 to invoke segmenter 204, for example, to apply the machine learning model to application data 266 and output a storage mask.

[0058] The segmenter 204 uses training data to generate a trained machine learning model. For example, the training data may include pairs of an initial image with a sky and an object, and an output image with a sky mask and an object mask.

[0059] The training data may come from any source, e.g., a data repository specifically marked for training, data that has been given permission to be used as training data for machine learning, etc. In some embodiments, training may occur on the media server 101, which provides the training data directly to the user device 115, training may occur locally on the user device 115, or a combination of both.

[0060] In some embodiments, segmenter 204 uses weights obtained and unedited / transferred from another application. For example, in these embodiments, a trained model may be generated, e.g., on a different device, and provided as part of segmenter 204. In various embodiments, the trained model may be provided as a data file that includes a model structure or format (e.g., defining the number and type of neural network nodes, the connectivity between the nodes, and the organization of the nodes into multiple layers) and associated weights. Segmenter 204 may read the trained model data file and implement a neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model.

[0061] The trained machine learning model may include one or more model forms or structures, such as a linear network, a deep learning neural network implementing multiple layers (e.g., "hidden layers" between an input layer and an output layer, each layer being a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from processing each tile), a sequence-to-sequence neural network (e.g., a network that takes sequential data as input, such as words in a sentence or frames in a video, and produces a resulting sequence as output), or any other type of neural network.

[0062] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., an input layer) may receive data as input or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used to analyze an initial image. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. For example, the first layer may output a segmentation between foreground and background. The final layer (e.g., an output layer) produces the output of the machine learning model. For example, the output layer may receive the segmentation of the initial image into foreground and background and output whether a pixel is part of a saved mask. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0063] In different embodiments, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to a model structure or format. In some embodiments, a node may be a memoryless computational node configured, for example, to process a unit of input and produce a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, a node may include memory, for example, capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node, which may use memory to maintain a "state" that allows the node to behave like a finite state machine (FSM).

[0064] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as multiple nodes organized into layers, as specified by the model format or model structure. At initialization, each weight may be applied to the nodes connected according to the model format, e.g., to the connection between each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained, e.g., using training data, to produce results.

[0065] Training may include applying supervised learning methods. In supervised learning, the training data may include multiple inputs (e.g., images, storage masks, etc.) and corresponding ground truth outputs for each input (e.g., an object ground truth mask that correctly identifies objects, a sky ground truth mask that correctly identifies sky in each image, a shadow ground truth mask that correctly identifies shadows in each image, etc.). Based on a comparison of the model's outputs and the ground truth outputs, the values ​​of the weights are automatically adjusted, for example, in a manner that increases the probability that the model will produce the ground truth output for the image.

[0066] In various embodiments, the trained model includes a set of weights or embeddings that correspond to the model structure. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings that correspond to the model structure. In embodiments where data is omitted, the segmenter 204 may generate a trained model that is based on prior training, e.g., by a developer of the segmenter 204, by a third party, etc. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.

[0067] In some embodiments, the trained machine learning model receives an initial image having a sky and one or more objects. In some embodiments, the trained machine learning model outputs a sky mask corresponding to the sky and one or more object masks corresponding to the one or more objects, where the sky mask and the object mask are preserved masks. In some embodiments, the trained machine learning model receives an output image having one or more shadows and outputs a shadow mask.

[0068] In some embodiments, the machine learning model outputs a confidence value for each saved mask output by the trained machine learning model. The confidence value may be expressed as a percentage, a number between 0 and 1, etc. For example, the machine learning model outputs a confidence value of 85% for the confidence that the saved mask correctly incorporates the subject and does not contain pixels from another person or object.

[0069] The diffusion module 206 receives as input the initial image and a request to change the lighting in the initial image. The diffusion module 206 outputs an output image that satisfies the request to change the lighting. In some embodiments, the diffusion module 206 performs text conditioning to generate an output image conditioned on the text requirement. For example, if the text requirement is of a sunset, the diffusion module 206 performs the text conditioning by generating a version of the initial image that is modified to appear to have been captured during a sunset. The diffusion model may perform diffusion until a balance between process efficiency and output image quality is achieved.

[0070] In some embodiments, the diffusion module 206 trains the diffusion model using two types of training data: The first type of training data includes pairs of images, which may include synthetic pairs generated through an inter-prompt generation machine learning model. The inter-prompt generation machine learning model is a diffusion model that receives a text prompt, uses cross-attention to extract keys and values ​​from the text prompt, and switches portions of the attention map previously generated for the first image based on the input text prompt to output a second image that matches the text prompt.

[0071] The second type of training data includes pairs of real and synthetic images. The real images are received by a diffusion model, such as a denoising diffusion implicit model (DDIM). The diffusion model outputs a synthetic image using an inversion method based on the real image and instructions on how to edit the input image. The diffusion module 206 trains the diffusion model to generate output images from requests using a forward process, in which the diffusion model adds noise to the data, and an inverse process, in which the diffusion model learns to recover the data from the noise.

[0072] The diffusion module 206 trains a diffusion model to maintain photorealism and preserve the identities of people depicted in images. During training, the diffusion model receives editing instructions and modifies the editing instructions to create corresponding prompts based on a language model, such as a large-scale language model. For example, the diffusion module 206 uses a language model to convert the editing instruction "make it sunset time" into prompts describing various outdoor scenes during the day and corresponding prompts for sunset. The diffusion model creates a set of input and output image pairs from the generated prompt pairs, where each prompt can generate N images (using a different seed). The diffusion module 206 filters certain images from the image pairs, such as image transformations that are inconsistent with the given editing instruction, image transformations that do not result in sufficiently aligned images, and mismatched pairs. In some embodiments, the diffusion module 206 also filters images based on an edit alignment score, which reflects the alignment between the image-to-image transformation and the original edited caption, and an image-text alignment score, which reflects the alignment between the input / output image and the corresponding input / output prompt. In some embodiments, the diffusion module 206 trains the diffusion model by generating one or more loss functions based on the filtered images from the image pairs.

[0073] Once the diffusion model is trained, it receives requests to change the lighting in the initial image. In some embodiments, the requests do not come directly from the user, but instead are pre-populated by the media application 103, such as requests selected by the user from a library of pre-made text requests.

[0074] In some embodiments, the resolution module 208 generates a super-resolved version of at least a portion of the output image. For example, the resolution module 208 may generate a super-resolved version from a portion of the output image that corresponds to a sky segment. In some embodiments, the diffusion model works best with a low-resolution output image. As a result, the resolution module 208 advantageously improves the quality of the sky segment by generating a super-resolved version of the output image.

[0075] In some embodiments, the resolution module 208 generates a super-resolved version of at least a portion of the output image using one or more of the following techniques: The resolution module 208 may perform pre-upsampling by upsampling a low-resolution output image to a coarse high-resolution image of a desired size using bicubic interpolation and provide the coarse high-resolution image as input to a deep convolutional neural network (CNN), which outputs the super-resolved version; The resolution module 208 may also perform post-upsampling by providing the low-resolution image to a CNN without increasing the resolution, with the upsampling layer being applied at the end of the CNN; In some embodiments, the resolution module 208 outputs the super-resolved version using a diffusion model instead of or in addition to a CNN.

[0076] The coloring module 210 modifies the coloring of the initial image to match the coloring of the output image. In some embodiments, the coloring module 210 modifies a portion of the initial image, such as an object, and does not modify the remainder of the initial image, because the remainder of the initial image is replaced with a super-resolved version of at least a portion of the output image during fusion. As a result, the coloring module 210 may modify the coloring of the portion of the initial image based on the content of the initial image (i.e., whether the initial image includes more than a sky segment and an object segment).

[0077] In some embodiments, the coloring module 210 determines a bilateral grid approximation, which is a three-dimensional array that combines the dimensions of a two-dimensional spatial region corresponding to an (x,y) location in the image plane with the dimensions of a one-dimensional region, typically image intensity. In some embodiments, the coloring module 210 performs bilateral grid upsampling (BGU), which determines a local color transformation (excluding sky portions) between the initial image and the output image by matching low-resolution versions of the input / output image pair and applies an affine model to the high-resolution input. The coloring module 210 applies the local color transformation to the initial image while preventing sky colors using a sky mask.

[0078] The fusion module 212 fuses the modified initial image with the output image to form a fused image, using an object mask to prevent modifications to the object from the modified initial image and a sky mask to prevent modifications to the sky from the output image during fusion. The object mask advantageously prevents fusion of the object with the modified initial image in the output image. Because the diffusion module 206 may generate an output image with a discernible distortion in the object, the object mask is used to ensure that a distorted version of the object is not mixed with the initial image.

[0079] Pixels in the sky mask correspond to sky pixels in the initial image, and vice versa. When the output images are fused, sky pixels in the modified initial image are merged or not modified, respectively. Pixels in the object mask correspond to object pixels in the initial image and pixels in the modified initial image, and vice versa. When the modified initial image is fused, object pixels in the modified initial image are merged or not modified, respectively.

[0080] In some embodiments in which a super-resolved version of the output image is generated, the fusion module 212 fuses the super-resolved version of at least a portion of the output image with the modified initial image, using a sky mask to prevent modification of the sky from the modified image during fusion. For example, the fusion module 212 may fuse a portion of the super-resolved version of the output image that corresponds to a sky segment with the modified image. The sky mask advantageously prevents the super-resolved version of the sky from being altered by the modified initial image. The fusion module 212 may fuse the super-resolved version of the portion of the output image after the output image is fused with the modified image, or may fuse all three images during the same fusion step.

[0081] In some embodiments, the diffusion module 206 generates an output image with one or more shadows on corresponding objects that match the illumination. For example, the shadows correspond to the direction of sunlight cast from the sun. In some embodiments, the segmenter 204 determines to output a shadow mask that is used to protect shadows on people and / or objects. The fusion module 212 can prevent retouching of the shadows while fusing the output image with the retouched image.

[0082] Although the diffusion module 206, resolution module 208, coloring module 210, and fusing module 212 are shown as separate components in Figure 2, one or more of the components may be combined. For example, the diffusion module 206 may also perform the fusing functions described with reference to the fusing module 212.

[0083] Exemplary Architecture 5 is a block diagram of an example architecture 500 for generating a fused image incorporating a request. An input image 505 is received by a media application, and segmentation 510 is performed on the input image to generate a segmentation mask 515. The segmentation mask includes a sky segment and an object segment. The input image 505 and a request to modify the lighting are received by a diffusion model 520, which generates an output image 525 that satisfies the text request. In this example, the request to modify the lighting is to make the input image 505 a night scene with a moon.

[0084] The media application performs a super-resolution 527 process on the sky portion of the output image 525 to improve the quality of the output image 525. As a result of the super-resolution 527 process, the media application outputs a super-resolution sky image 535.

[0085] The diffusion module applies BGU 530 to the initial image, so that BGU 530 matches the coloring of the output image. For example, modified image 540 has darker colors than initial image 505 because it appears to have been captured at night.

[0086] The media application fuses the rectified image 540 with the super-resolution sky image 535, using the segmentation mask 515 to prevent the sky in the super-resolution sky image 535 from being rectified by the rectified image 540 and to prevent objects in the rectified image 540 from being combined with potentially distorted versions of objects from the super-resolution sky image 535. The media application fuses the images to produce a final image 550.

[0087] Exemplary Flowchart 6 shows an example flowchart of a method 600 for generating a fused image with modified lighting. Method 600 may be performed by computing device 200 of FIG. 2. In some embodiments, method 600 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.

[0088] 6 may begin at block 602. In block 602, an initial image and a request to modify the lighting in the initial image are provided as input to a diffusion model, where the initial image includes an object and a sky.

[0089] A request to change the lighting may include a user providing a text request including attributes selected from the group of light level ("change this to a moonlit night"), amount of clouds in the sky ("change this to a clear sky"), and / or sky color ("change this to a red and orange sky"). A request to change the lighting may include a library of region suggestions, global presets, a menu of options, and / or pre-made text requests associated with one or more regions of the initial image.

[0090] In some embodiments, the initial image is determined to include an outdoor scene and suggestions to modify the lighting are provided to the user, where a request to change the lighting is received in response to providing the suggestions. Block 602 may be followed by block 604.

[0091] The diffusion model outputs a satisfying output image in block 604. Block 604 may be followed by block 606.

[0092] In block 606, sky segments and object segments are determined from the initial image. Block 606 may be followed by block 608.

[0093] A sky mask corresponding to the sky segment and an object mask corresponding to the object segment are generated in block 608. Block 608 may be followed by block 610.

[0094] At block 610, the coloration of the initial image is modified to match the coloration of the output image. In some embodiments, modifying the coloration of the initial image includes performing a BGU to identify a local color transformation between the initial image and the output image and applying the local color transformation to the initial image. Block 610 may be followed by block 612.

[0095] In block 612, an optional step includes generating a super-resolution version of at least a portion of the output image from the output image. Block 612 may be followed by block 614.

[0096] At block 614, the modified initial image is fused with the output image to form a fused image, while using the object mask to prevent object modifications from the modified initial image and the sky mask to prevent sky modifications from the output image during fusion. If a super-resolved version of at least a portion of the output image is generated, fusing includes fusing the super-resolved version of at least a portion of the output image while using the object mask to prevent object modifications from the modified initial image and the sky mask to prevent sky modifications from the super-resolved version of at least a portion of the output image during fusion.

[0097] In some embodiments, the output image includes one or more shadows corresponding to one or more objects in the output image, and the method further includes determining, from the initial image, shadow segments corresponding to the one or more shadows in the output image, and fusing the output image with the modified initial image includes preventing modification to the one or more shadows from the output image during fusing.

[0098] In addition to the above, a user may be provided with controls that allow the user to choose both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social networks, social actions, or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent from the server to the user. Furthermore, certain data may be processed in one or more ways such that personally identifiable information is removed before it is stored or used. For example, a user's identifying information may be processed so that personally identifiable information about the user cannot be determined, or if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, a user may control what information is collected about them, how that information is used, and what information is provided to them.

[0099] Thus, in accordance with the foregoing, a media application provides an initial image and a request to modify the lighting in the initial image, where the initial image includes an object and a sky, as input to a diffusion model. The media application uses the diffusion model to output an output image that meets the request. The media application determines sky segments and object segments from the initial image. The media application generates a sky mask corresponding to the sky segment and an object mask corresponding to the object segment. The media application modifies the coloring of the initial image to match that of the output image. The media application fuses the modified initial image with the output image to form a fused image, using the object mask to prevent modification of the object from the modified initial image and the sky mask to prevent modification of the sky from the output image during fusing.

[0100] In the above description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form in order to avoid obscuring the description. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments may apply to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0101] A reference herein to "some embodiments" or "some instances" means that a particular feature, structure, or characteristic described in connection with an embodiment or instance may be included in at least one implementation herein. The appearances of the phrase "in some embodiments" in various places in this specification are not necessarily all referring to the same embodiments.

[0102] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0103] It should be recognized, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the discussion that follows, throughout this specification, discussions utilizing terms including "processing" or "calculating" or "computing" or "determining" or "displaying" and the like will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is also represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display devices.

[0104]

[0013] Embodiments herein may also relate to a processor for performing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key with non-volatile memory, or any type of disk, including any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0105] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0106] Furthermore, this specification may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution apparatus, or instruction execution device.

[0107] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. 1. A computer-implemented method comprising: providing an initial image and a request to modify the lighting in the initial image as input to a diffusion model, the initial image including an object and a sky, the method further comprising: outputting an output image that satisfies the requirements using the diffusion model; determining sky segments and object segments from the initial image; generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment; modifying the coloration of the initial image to match the coloration of the output image; fusing the modified initial image with the output image to form a fused image while, during fusing, using the object mask to prevent modifications to the object from the modified initial image and using the sky mask to prevent modifications to the sky from the output image.

2. Modifying the coloring of the initial image comprises: performing bilateral grid upsampling (BGU) to identify local color transformations between the initial image and the output image; and applying the local color transformation to the initial image.

3. generating a super-resolution version of at least a portion of the output image from the output image; 2. The method of claim 1, wherein fusing the modified initial image with the output image comprises fusing the super-resolved version of at least the portion of the output image while using the object mask to prevent modifications to the object from the modified initial image and using the sky mask to prevent modifications to the sky from the super-resolved version of at least the portion of the output image during fusing.

4. the output image includes one or more shadows corresponding to one or more objects in the output image, and the method further comprises: determining, from the output image, shadow segments corresponding to the one or more shadows in the output image; generating a shadow mask corresponding to the shadow segment; 2. The method of claim 1, wherein fusing the output image with the modified initial image includes using the shadow mask to prevent modification to the one or more shadows from the output image during fusing.

5. 10. The method of claim 1, wherein the request to change the lighting comprises a user providing a text request including attributes selected from the group of light level, amount of clouds in the sky, color of the sky, and combinations thereof.

6. 2. The method of claim 1, wherein the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, global presets, a menu of options, a library of pre-made text requests, and combinations thereof.

7. determining that the initial image includes an outdoor scene prior to receiving the request to change the lighting in the initial image; The method of claim 1 , further comprising: providing a user with suggestions for modifying the lighting.

8. A program having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: providing an initial image and a request to modify illumination in the initial image as input to a diffusion model, the initial image including an object and a sky, the operations further comprising: outputting an output image that satisfies the requirements using the diffusion model; determining sky segments and object segments from the initial image; generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment; modifying the coloration of the initial image to match the coloration of the output image; fusing the modified initial image with the output image to form a fused image while, during the fusing, using the object mask to prevent modifications to the object from the modified initial image and using the sky mask to prevent modifications to the sky from the output image.

9. Modifying the coloring of the initial image comprises: performing bilateral grid upsampling (BGU) to identify local color transformations between the initial image and the output image; and applying the local color transformation to the initial image.

10. The operation is generating a super-resolution version of at least a portion of the output image from the output image; 9. The program of claim 8, wherein fusing the modified initial image with the output image comprises fusing the super-resolved version of at least the portion of the output image while using the object mask to prevent modifications to the object from the modified initial image and using the sky mask to prevent modifications to the sky from the super-resolved version of at least the portion of the output image during fusing.

11. the output image includes one or more shadows corresponding to one or more objects in the output image, and the operation comprises: determining, from the output image, shadow segments corresponding to the one or more shadows in the output image; generating a shadow mask corresponding to the shadow segment; 9. The program of claim 8, wherein fusing the output image with the modified initial image includes using the shadow mask to prevent modification to the one or more shadows from the output image during fusing.

12. 9. The program of claim 8, wherein the request to change the lighting comprises a user providing a text request including attributes selected from the group of light level, amount of clouds in the sky, color of the sky, and combinations thereof.

13. 9. The program of claim 8, wherein the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, global presets, a menu of options, a library of pre-made text requests, and combinations thereof.

14. The operation is determining that the initial image includes an outdoor scene prior to receiving the request to change the lighting in the initial image; The program of claim 8 , further comprising providing a user with suggestions for modifying the lighting.

15. 1. A system comprising: a processor; a memory coupled to the processor and having instructions stored thereon, the instructions, when executed by the processor, causing the processor to perform operations, the operations including: performing operations including providing an initial image and a request to modify illumination in the initial image as input to a diffusion model, the initial image including an object and a sky, the operations further comprising: outputting an output image that satisfies the requirements using the diffusion model; determining sky segments and object segments from the initial image; generating a sky mask corresponding to the sky segment and an object mask corresponding to the object segment; modifying the coloration of the initial image to match the coloration of the output image; fusing the modified initial image with the output image to form a fused image while, during the fusing, using the object mask to prevent modifications to the object from the modified initial image and using the sky mask to prevent modifications to the sky from the output image.

16. Modifying the coloring of the initial image comprises: performing bilateral grid upsampling (BGU) to identify local color transformations between the initial image and the output image; and applying the local color transformation to the initial image.

17. The operation is generating a super-resolution version of at least a portion of the output image from the output image; 16. The system of claim 15, wherein fusing the modified initial image with the output image includes fusing the super-resolved version of at least the portion of the output image while using the subject mask to prevent modifications to the subject from the modified initial image and using the sky mask to prevent modifications to the sky from the super-resolved version of at least the portion of the output image during the fusing.

18. the output image includes one or more shadows corresponding to one or more objects in the output image, and the operation comprises: determining, from the output image, shadow segments corresponding to the one or more shadows in the output image; generating a shadow mask corresponding to the shadow segment; 16. The system of claim 15, wherein fusing the output image with the modified initial image includes using the shadow mask to prevent modification to the one or more shadows from the output image during fusing.

19. 16. The system of claim 15, wherein the request to change the lighting comprises a user providing a text request including attributes selected from the group of light level, amount of clouds in the sky, color of the sky, and combinations thereof.

20. 16. The system of claim 15, wherein the request to change the lighting is selected from the group of area suggestions associated with one or more areas of the initial image, global presets, a menu of options, a library of pre-made text requests, and combinations thereof.

Citation Information

Patent Citations

  • Image processing method and device, electronic device, and storage medium

    JP2021530056A