A media application provides, as input to a
diffusion model, an initial image and a request to change a lighting in the initial image, wherein the initial image includes a subject and a
sky. The media application outputs, with the
diffusion model, an output image that satisfies the request. The media application determines, from the initial image, a
sky segment and a subject segment. The media application generates a
sky mask that corresponds to the sky segment and a subject
mask that corresponds to the subject segment. The media application modifies a coloring of the initial image to match a coloring of the output image. The media application blends the modified initial image with the output image to form a blended image while using the subject
mask to prevent modification to the subject from the modified initial image and the sky mask to prevent modification to the sky from the output image during the blending.