Method and device for generating shadows in a media

WO2026164335A1PCT designated stage Publication Date: 2026-08-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-06-02
Publication Date
2026-08-06

Smart Images

  • Figure KR2025007580_06082026_PF_FP_ABST
    Figure KR2025007580_06082026_PF_FP_ABST
Patent Text Reader

Abstract

A method (300) and an electronic device (100) is disclosed. The method (300) includes obtaining information associated with a media based on analyzing the media. Further, the method (300) includes determining a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media. The method (300) further includes identifying at least one object shadow to be generated based on one or more present existing objects in the media. Furthermore, the method (300) includes generating a composite prompt using the identified at least one object shadow. Moreover, the method (300) includes generating one or more shadows in the media based on processing the composite prompt and the obtained information.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND DEVICE FOR GENERATING SHADOWS IN A MEDIA

[0001] The present disclosure relates to media enhancement, and more particularly, to a method and an electronic device for generating shadows in a media.

[0002] The information in this section merely provides background information related to the present disclosure and may not constitute prior art(s) for the present disclosure.

[0003] Media are also often posted to various internet sites, such as web pages, social networking services, etc. for users and others to view. A user commonly edits the media such as images and video to enhance the visual appearance of the media.

[0004] Recent years have seen significant advancements in software platforms for performing media editing tasks. The media editing is used to enhance the visual appearance of the media. The media editing includes a generation of shadows in the media for enhancing the visual appearance of the media. More specifically, the shadows are used for the conversion of simple or everyday scenes of the media into aesthetically interesting media.

[0005] Currently, there are editing tools that are used to generate the shadows in the media. However, these editing tools rely on physical objects to generate the shadows in the media. For instance, the shadows which are generated are of the physical objects present in the media.

[0006] Further, the shadows rely heavily on light sources, and without proper lighting, the shadows may not be distinct enough to capture effectively. Furthermore, the generation of the shadows require expertise planning and availability of the physical objects to introduce the required shadows in the media.

[0007] Additionally, in current solutions, the generation of the shadows require manual efforts for selecting the physical objects which are optimum for generating the shadows. The selection of the physical objects which are not optimum for a scene present in the media, leads to degrade the visual appearance of the media. Additionally, in the current solutions, the generation of the shadows is very time-consuming process, which degrades the user experience.

[0008] Therefore, there is a need for an alternative solution that may overcome the above-discussed limitations and provide an improved method and a system for generating one or more shadows in a media.

[0009] The drawbacks / difficulties / disadvantages / limitations of the related techniques explained in the background section are just for exemplary purposes and the disclosure would never limit its scope only such limitations. A person skilled in the art would understand that this disclosure and below mentioned description may also solve other problems or overcome the other drawbacks / disadvantages.

[0010] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the disclosure. This summary is neither intended to identify essential inventive concepts of the disclosure nor is it intended for determining the scope of the disclosure.

[0011] According to an aspect of the present disclosure, a method is disclosed. The method includes obtaining information associated with a media based on analyzing the media. Further, the method includes determining a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media. The method further includes identifying at least one object shadow to be generated based on one or more present existing objects in the media. Furthermore, the method includes generating a composite prompt using the identified at least one object shadow. Moreover, the method includes generating one or more shadows in the media based on processing the composite prompt and the obtained information.

[0012] According to an aspect of the present disclosure, an electronic device is disclosed. The electronic device includes memory storing instructions. The electronic device further includes at least one processor comprising processing circuitry, in communication with the memory, wherein the instructions, when executed by the at least one processor individually or collectively. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to obtain information associated with a media based on analyzing the media. Further, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to determine a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify at least one object shadow to be generated based on one or more present existing objects in the media. Furthermore, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate a composite prompt using the identified at least one object shadow. Moreover, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate one or more shadows in the media based on processing the composite prompt and the obtained information.

[0013] According to an aspect of the present disclosure, a machine readable medium comprising instructions that when executed by at least one processor individually or collectively, cause an electronic device to perform the method provided.

[0014] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting to its scope. The disclosure will be described and explained with additional specificity and detail in the accompanying drawings.

[0015] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0016] Figure 1 illustrates a schematic block diagram of a system for generating one or more shadows in a media, in accordance with an embodiment of the present disclosure;

[0017] Figure 2 illustrates a block diagram depicting an exemplary process-flow for generating the one or more shadows in the media, in accordance with an embodiment of the present disclosure;

[0018] Figure 3 illustrates a flowchart depicting an exemplary method for generating the one or more shadows in the media, in accordance with an embodiment of the present disclosure;

[0019] Figure 4 illustrates a flowchart depicting sub-operations for determining the a correlated color temperature CCT, in accordance with an embodiment of the present disclosure;

[0020] Figure 5 illustrates a flowchart depicting sub-operations for generating a scene graph, in accordance with an embodiment of the present disclosure;

[0021] Figure 6 illustrates an example representation of the generated scene graph, in accordance with an embodiment of the present disclosure;

[0022] Figure 7 illustrates a flowchart depicting sub-operations for generating the composite prompt, in accordance with an embodiment of the present disclosure; and

[0023] Figures 8A-8B illustrate an exemplary use case depicting the generation of the shadow / s in the media, in accordance with an embodiment of the present disclosure;

[0024] Figures 9A-9B illustrate an exemplary use case depicting the generation of the shadow / s in the media in accordance with an embodiment of the present disclosure;

[0025] Figures 10A-10B illustrate an exemplary use case depicting the generation of the shadow / s in the media, in accordance with an embodiment of the present disclosure;

[0026] Figures 11A-11B illustrate an exemplary use case depicting the generation of the shadow / s in the media, in accordance with an embodiment of the present disclosure;

[0027] Figures 12A-12B illustrate an exemplary use case depicting the generation of the shadow / s in the media, in accordance with an embodiment of the present disclosure; and

[0028] Figures 13-15 illustrate exemplary images depicting generated shadows, in accordance with an embodiment of the present disclosure.

[0029] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0030] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It should be understood at the outset that although illustrative implementations of the embodiments of the present disclosure are illustrated below, the present disclosure may be implemented using any number of techniques, whether currently known or in existence. The present disclosure is not necessarily limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but may be modified within the scope of the present disclosure.

[0031] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the disclosure and are not intended to be restrictive thereof.

[0032] Reference throughout this specification to "an aspect", "another aspect" or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrase "in an embodiment", "in another embodiment" and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0033] It is to be understood that as used herein, terms such as, "includes," "comprises," "has," etc. are intended to mean that the one or more features or elements listed are within the element being defined, but the element is not necessarily limited to the listed features and elements, and that additional features and elements may be within the meaning of the element being defined. In contrast, terms such as, "consisting of" are intended to exclude features and elements that have not been listed.

[0034] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term "or" as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.

[0035] As is traditional in the field, embodiments may be described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which may be referred to herein as units or modules or the like, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.

[0036] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.

[0037] Figure 1 illustrates a schematic block diagram of a system generating one or more shadows in a media, in accordance with an embodiment of the present disclosure. In an embodiment, the one or more shadows alternatively may be referred to as shadows for the sake of brevity. In an embodiment, the shadows herein refer to shadows of one or more objects such as a tree, person, or similar entities. In an embodiment, the media may include but is not limited to videos, images, and graphics interchange format (GIP) of a real-world scene.

[0038] In an embodiment, the electronic device 100 may include a memory 102 including a database 104. The electronic device 100 may further include a processor 106 communicatively coupled with the memory 102. Further, the electronic device 100 may include an Input / Output (I / O) interface 110, and a plurality of modules 120. In an embodiment, the electronic device 100 may be implemented at a user equipment (UE). In an example, the UE may be a smartphone, a laptop computer, a desktop computer, a personal computer (PC), a notebook, a tablet, or a smartwatch.

[0039] In an embodiment, the electronic device 100 may be a cloud-based system, that may include a server, specifically a cloud server. In an embodiment, the electronic device 100 may be implemented by a combination of the UE and the server. In an embodiment, one or more steps may be performed in the UE and the remaining steps may be performed by the server.

[0040] In an embodiment, the electronic device 100 may be implemented using an artificial intelligence (AI) model which may include a plurality of neural network layers. Examples of neural networks include but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), and restricted boltzmann machine (RBM). A function associated with the AI module may be performed through the non-volatile memory, the volatile memory, and the processor 106.

[0041] In one embodiment, the memory 102 is configured to store instructions executable by the processor 106. In one embodiment, the memory 102 communicates via a bus within the electronic device 100. The memory 102 includes but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like. In one example, the memory includes a cache or random-access memory (RAM) for the processor 106. In alternative examples, the memory 102 is separate from the processor 106 such as a cache memory of a processor, the system memory, or other memory. The memory 102 is an external storage device or the memory 102 is for storing data. The memory 102 is operable to store instructions executable by the processor 106. The functions, acts, or tasks illustrated in the figures or described are performed by the programmed processor for executing the instructions stored in the memory 102. The functions, acts, or tasks are independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code, and the like, operating alone or in combination. Likewise, processing strategies include multiprocessing, multitasking, parallel processing, and the like.

[0042] In an embodiment, the processor 106 may be a single processing unit or a set of units each including multiple computing units. The processor 106 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions (computer-readable instructions) stored in the memory 102. Among other capabilities, the processor 106 may be configured to fetch and execute computer-readable instructions and data stored in the memory 102. The processor 106 includes one or a plurality of processors. The plurality of processors is further implemented as a general-purpose processor. The processor 106 may be disposed in communication with one or more input / output (I / O) devices via the input / output (I / O) interface 110. The I / O interface 110 employs communication code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, and the like, etc. In an embodiment of the present disclosure, the I / O interface 110 employs ethernet, industrial wireless local area network (LAN), process field bus (PROFIBUS), actuator sensor (AS) Interface, and the like.

[0043] In one embodiment, the plurality of modules 120 may include the one or more instructions (stored in a memory 102) that may be executed to cause the electronic device 100, in particular, the processor 106 of the electronic device 100, to perform the one or more functions / methods, as discussed here in the present disclosure. In one embodiment, the plurality of modules 120 may be implemented at least in part as hardware, which may work in conjunction with the instructions to perform the functions / methods discussed herein. At least one of the plurality of modules may be implemented through an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system. The AI model may consist of a plurality of neural network layers. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0044] The plurality of modules 120 may include an estimating module 122, a determining module 124, and a generating module 126. In an embodiment, the estimating module 122, the determining module 124, and the generating module 126 may be in communication with each other. The working of the plurality of modules 120 may be explained in conjunction with Figure 2.

[0045] Figure 2 illustrates a flow diagram 200 depicting a block diagram depicting an exemplary process-flow, in accordance with an embodiment of the present disclosure. The estimating module 122 may be configured to estimate(e.g., obtain) information associated with the media 201 based on analyzing the media 201. In an embodiment, the information may include but is not limited to depth map and / or a light characteristic such as a lighting sphere. The light characteristics herein refer to behavior and interaction of light with one or more objects in a scene. For example, the objects may herein refer to person, animals, tree, etc. The scene herein refers to a visual content or environment depicted within the media 201. Further, The determining module 124 may be configured to determine a correlated color temperature (CCT) corresponding to one or more light sources illuminating the scene based on the media 201. The scene herein refers to a visual content or environment depicted within the media 201. The CCT may herein refer to a measure of a colour appearance of the one or more light sources, that is expressed in Kelvins (K). More specifically, the colour appearance may herein refers to an appearance of a colour emitted by the one or more light sources to a human eye. In an embodiment, the CCT may describe the colour of the light emitted and a degree of "warmness" or "coolness" in the light appearance. Further, The generating module 126 may be configured to generate a scene graph associated with the media 201 based on the light characteristic. In an embodiment, the scene graph may indicate a visual representation of a relation between one or more present existing objects in the media 201. Furthermore, The generating module 126 may be configured to generate a composite prompt using at least one of the scene graph, the light characteristic, and the CCT. The composite prompt may herein refer to a prompt of the shadows 201a to be generated. Moreover, The generating module 126 may be configured to generate the one or more shadows 201a in the media 201 using an artificial intelligence (AI) model based on processing the composite prompt and the estimated(e.g., obtained) information. In an embodiment, at block 210a, the obtained information such as the depth map and the light characteristic is fed to an adaptor that may pass the obtained information to the AI model such as the diffusion model. The depth map and the light characteristic may ensure the preservation of the shadow's structural integrity during a diffusion process. The shadow's structural integrity may herein refer to how well the shadow / s are represented and structured in the media 201. Thus, based on incorporating the depth map, the adaptor may guide the diffusion model to follow a disparity and scale of the scene associated with the media 201, resulting in a more accurate generation of the shadows 201a. Additionally, the adaptor may incorporate the information associated with the light characteristic to control the shadow's direction and intensity. In an advantageous aspect, thus the adapter enables the diffusion model to generate the shadows 201a that may be consistent with lighting conditions in the media 201.

[0046] Now, the present disclosure is explained in detail in reference to a method 300 disclosed in Figure 3. More specifically, referring to Figures 1-3 in combination, the various operations of the method 300 as described hereinafter may be executed in the electronic device 100, or specifically in the processor 106 of the electronic device 100, for generating the shadows 201a in the media 201, wherein the shadows 201a are generated in the media to ultimately produce a final composited media 203. In an embodiment, the shadows 201a may be generated in the media 201. The electronic device may generate the final composited media 203 by generating 201a in the media 201. The final composited media 203 may be obtained by combining the media 201 and the shadows 201a.

[0047] Figure 3 illustrates a flowchart representation of the method 300 for generating the shadows 201a in the media 201, in accordance with an embodiment of the present disclosure. In an embodiment, the method 300 is a computer-implemented method that is explained in detail in the below paragraphs.

[0048] In an embodiment, at operation 302, the method 300 may include estimating(e.g., obtaining) the information associated with the media 201. In an embodiment, the information may be estimated(e.g., obtained) based on analyzing the media 201. In an embodiment, the media 201 may be received in a red-green-blue (RGB) format. For example, the media 201, i.e., the RGB image is received to obtain the information thereof. In an embodiment, the information may be the depth map and / or the light characteristic associated with the media 201. In an advantageous aspect, the depth maps may be utilized to gain insights into the relative distances and sizes of objects within the scene of the media 201, thereby enabling the shadows 201a to align seamlessly with a surrounding environment. In an advantageous effect, the present disclosure may provide a more natural and visually pleasing outcome, as the shadows 201a appear consistent with an overall depth perception of the media 201.

[0049] In an exemplary scenario, a teacher model generates pseudo-labels for the real images. Further, a student model is trained in the pseudo-labeled real images to ensure robust generalization. Thereafter, an encoder utilizes different scales of a vision model to extract features from the real images. Further, a depth decoder based on a dense prediction transformer (DPT), converts the extracted features into the depth map.

[0050] In one embodiment, the light characteristic such as a lighting sphere may be estimated(e.g., obtained) based on the media 201. In an embodiment, shelf models like but not limited to a diffusion light may be utilized to precisely estimate(e.g., obtain) the lighting in the media 201. In an exemplary scenario, the diffusion light leverages existing knowledge of diffusion models to obtain the lighting in the media 201. In an embodiment, a control-net and the adapter are trained using a high-quality synthetic media 201. In an embodiment, an input to the control-net is the depth map estimated(e.g., obtained) using an off-the-shelf estimator such as a depth-any model. In an embodiment, the media 201 is combined with an in-painting model trained using a low-rank adapter (LoRA) to paint the lighting sphere (chrome ball) into a centre of the media 201. The in-painting may enable to achieve accurate and high-quality lighting estimating(e.g., obtaining) for the media 201 such as real-world images.

[0051] Further, at operation 304, the method 300 may include determining the correlated color temperature (CCT) corresponding to the one or more light sources illuminating the scene based on the media 201. In an embodiment, the determination of the CCT may be explained in conjunction with Figure 4.

[0052] Figure 4 illustrates a flowchart depicting sub-operations for determining the CCT, in accordance with an embodiment of the present disclosure. In an embodiment, the determining module 124 may be configured to identify the colour temperature of light illuminating the objects at different parts of the scene.

[0053] At sub-operation 304a, the operation 304 may include splitting a window associated with the media 201 into a plurality of sub-windows based on a resolution of the media 201. In an exemplary scenario, the window herein refers to a region of the media 201 which is selected for processing.

[0054] Further, at sub-operation 304b, the operation 304 may include computing a mean value associated with each sub-window. In an embodiment, the CCT determining module may be configured to calculate an average colour value (mean value) for each sub-window of the media 201, for example, the RGB image using equation (1) as below:

[0055] .....equation (1)

[0056] Where n is a total number of pixels;

[0057] represents a value of a red pixel at ithlocation;

[0058] represents a value of a green pixel at ithlocation; and

[0059] represents a value of a blue pixel at ithlocation.

[0060] Furthermore, at sub-operation 304c, the operation 304 may include identifying the corresponding one or more light sources illuminating the scene based on the computed mean value, thereby determining the CCT corresponding to the one or more light sources at sub-operation 304d. More specifically, the corresponding one or more light sources are identified based on pre-set values of R / G and B / G values for different illuminants.

[0061] In one embodiment, the determining module 124 may be configured to compute a confidence for the determined CCT. The confidence may refer to how reliably and accurately a light source's colour temperature may be measured or predicted. In an embodiment, the confidence may be computed based on the difference between actual values and the determined mean value which is shown using equation (2) below:

[0062] ......equation (2)

[0063] Again, referring to Figure 3, at operation 306, the method 300 may include generating the scene graph associated with the media 201 based on the obtained light characteristic. In an embodiment, the generation of the scene graph may be explained in conjunction with Figure 5.

[0064] Figure 5 illustrates a flowchart depicting sub-operations for generating the scene graph, in accordance with an embodiment of the present disclosure. At sub-operation 306a, the operation 306 may include identifying the one or more present existing objects in the media 201 based on the obtained information associated with the media 201.

[0065] Further, at sub-operation 306b, the operation 306 may include determining a spatial relationship between the identified one or more objects. In an embodiment, the spatial relationship may indicate a relative position, a containment, an interaction, an orientation, and a distance of each object with respect to other objects. In an exemplary scenario, a region-based convolutional neural network (R-CNN) may be utilized to identify contextual information in the images (media 201) and use this information to identify and generate the relation between different objects in the images (media 201) using recurrent neural networks (RNN).

[0066] Furthermore, at sub-operation 306c, the operation 306 may include generating the scene graph using the AI model based on the determined spatial information. An example representation of the generation of the scene graph is illustrated in Figure 6.

[0067] Referring to Figure 6, at block 602, the AI model may analyze the image (media 201) to identify the objects in the media 201 and the relation between the objects in the media 201. Thereafter, at block 604, the scene graph is generated, providing a representation of the visual representation of the relation between the one or more present existing objects in the media 201 in a hierarchical manner. For example, there is an afternoon scene in the media 201, green grass, big wall in the park and a cycle standing behind the big wall.

[0068] Again, referring to Figure 3, at operation 308, the method 300 may include generating the composite prompt using the scene graph, the light characteristic, and / or the CCT. In an embodiment, the generation of the composite prompt is explained in conjunction with Figure 7.

[0069] Figure 7 illustrates a flowchart depicting sub-operations for generating the composite prompt, in accordance with an embodiment of the present disclosure.

[0070] At sub-operation 308a, the operation 308 may include determining one or more object shadows to be generated based on the scene graph associated with the media 201.

[0071] At sub-operation 308b, the operation 308 may include identifying at least one of positive shadow objects and negative shadow objects associated with the determined one or more object shadows to be generated. In an embodiment, the positive shadow objects (e.g., prompts) may mean the list of objects that are present in general for the given input scene. In an embodiment, the negative shadow objects (e.g., prompts) mean the list of objects that are not seen in general for the given input scene. For example, given an input desert scene, the positive shadow objects can be camel, cactus, etc., as they are more likely to be present in the desert. Whereas negative objects can be a beach-shade, snowman, etc., which do not occur in general and can make scene look a dramatic.

[0072] In an embodiment, a network T may be trained on a large-scale dataset which may be a scene recognition dataset to generate a list of predefined shadow objects for the identified one or more objects and scenes in the media 201. Further, one or more shadow objects may be selected based on the list of predefined shadow objects. In an embodiment, the one or more shadow objects may indicate objects having a relation with the scene within the media 201. In an embodiment, a relevance value for each shadow object may be determined using the AI model.

[0073] In an embodiment, the AI model may be configured to identify the positive shadow objects and the negative shadow objects among the selected shadow objects based on comparing the determined relevance value with a predefined threshold value using equation (3) and (4) as below:

[0074] Positive shadow object : value ......equation (3)

[0075] Negative shadow object : value ......equation (4)

[0076] At sub-operation 308c, the operation 308 may include generating the composite prompt based on the identified positive shadow objects / negative shadow objects, the light characteristic, and the CCT. In an embodiment, the composite prompt may alternatively be referred to as composite tokens that may ensure the shadows 201a generated are in accordance with the media 201. In one embodiment, the media 201 tokens may be generated from the CCT and the RGB media 201 of the lighting sphere (light characteristic). Thereafter, the media 201 tokens may be concatenated and projected into a feature space that may be processed by a text encoder to generate the composite tokens (composite prompt).

[0077] In one embodiment, the positive shadow objects may be used for generation of the shadows 201a. In an embodiment, the negative shadow objects may also be used for generation of the shadows 201a as sometime contradicting shadows may make the scene more aesthetic.

[0078] In one embodiment, if user input is available, the information of list of positive prompts with shadow which can be cast is generated. Further, corresponding negative prompts are used for training the network based on the object relation. In an embodiment, the user input may refer to the option where user exclusively inputs the objects / prompts for which are indicating objects to generated the shadow the user wants to.

[0079] Further, at operation 310, the method 300 may include generating the shadows 201a in the media 201 using the AI model. In an embodiment, the shadows 201a may be generated based on processing the composite prompt and the obtained information. In one embodiment, the obtained information such as the depth map and the light characteristic may be fed to an adaptor that may pass the obtained information to the AI model such as the diffusion model. The depth map may ensure the preservation of the shadow's structural integrity during a diffusion process. Thus, based on incorporating the depth map, the adaptor may guide the diffusion model to follow a disparity and scale of the scene associated with the media 201, resulting in more accurate shadow generation. Additionally, the adaptor may incorporate the information associated with the light characteristic to control the shadow's direction and intensity. This enables the diffusion model to generate the shadows 201a that may be consistent with lighting conditions in the media 201.

[0080] Further, the composite tokens mentioned earlier include valuable information that guides the diffusion model in generating the shadow / s 201a of a desired object. The composite tokens may be utilized to specify a type of object shadow to be generated, along with additional latent information related to a chrome ball (lighting characteristic). Moreover, it includes the correlated colour temperature (CCT) value of the light source, which plays a crucial role in determining the overall appearance and tone of the shadow / s 201a to be generated. By providing this comprehensive set of instructions, the composite tokens ensure that the diffusion model generates the shadow / s 201a that align perfectly with the media 201.

[0081] Figures 8A-8B illustrate an exemplary use case depicting the generation of the shadow / s 801a in the media 801, in accordance with an embodiment of the present disclosure. Referring to Figure 8A, a first image is illustrated on which the shadow / s 801a needs to be generated. Further, referring to Figure 8B, the first image is processed, and the shadow 801a of a boy and a girl is generated in the first image to ultimately produce image 803.

[0082] Figures 9A-9B illustrate an exemplary use case depicting the generation of the shadow / s 901a in the media 901, in accordance with an embodiment of the present disclosure. Referring to Figure 9A, a second image is illustrated. Further, referring to Figure 9B, the second image is processed, and the shadow 901a of the man is generated in the second image to ultimately produce image 903.

[0083] Figures 10A-10B illustrate an exemplary use case depicting the generation of the shadow / s 1001a in the media 1001, in accordance with an embodiment of the present disclosure. Referring to Figure 10A, a third image depicting a teacup is illustrated. The third image is processed, and the shadow 1001a of the hand and the tea bag is generated in the third image to ultimately produce image 1003 as illustrated in Figure 10B.

[0084] Figures 11A-11B illustrate an exemplary use case depicting the generation of the shadow / s 1101a in the media 1101, in accordance with an embodiment of the present disclosure. Referring to Figure 11A, a fourth image of the desert and shadow of people is illustrated. Further, the fourth image is processed, and the shadow 1101a of the camels in the fourth image to ultimately produce image 1103 is illustrated in Figure 11B.

[0085] Figures 12A-12B illustrate an exemplary use case depicting the generation of the shadow / s 1201a in the media 1201, in accordance with an embodiment of the present disclosure. In an embodiment, a fifth image shows a low-light scene of a cycle with a wall. The fifth image is processed using the present disclosure and the shadow 1201a of the tree in the wall to ultimately produce image 1203 is generated as illustrated in Figure 12B. The generation of the shadow eliminates the need to introduce specific objects that are captured to enhance a visual appeal of the image. The modified image with the generated shadow enhances the visual appeal of the image and focuses on the shape and form of the shadow that is introduced in the image.

[0086] In one embodiment, the present disclosure may be used for various applications, which is discussed with reference to the images (e.g., media 1301, 1401, and 1501) with the generated shadows (e.g., 1301a, 1401a, and 1501a) illustrated in Figures 13-15. Referring to Figures 13-15, the present disclosure may be used in advertising and marketing, privacy and personalization, enhanced context and depth, and storytelling.

[0087] · Advertising and Marketing - Unique shadow images may enhance product photography, adding mood and atmosphere to advertisements.

[0088] · Privacy and Personalization - The shadows 201a may be used to create interesting silhouettes that add mystery and intrigue to the media 201. Also, this feature allows for the personalization of the shadows 201a, where the user may control the human silhouettes introduced in the media 201.

[0089] · Enhanced Context and Depth - The shadows 201a generated in the media 201 may add context, depth, drama, and emotions to the media 201.

[0090] · Storytelling - The shadows 201a of any fantasy subjects / mythical objects may be introduced in the media 201 to enhance the story of the media 201. These objects may not be captured and hence using shadows 201a is a very effective storytelling tool.

[0091] In an embodiment, the present disclosure at least provides the following advantages:

[0092] The present disclosure uses the generation of the shadows 201a without relying on physical objects to cast the shadows 201a in the media 201.

[0093] Further, the present disclosure also using the light characteristic and the CCT to generates the shadows 201a which are optimum for the media 201, thereby eliminating a need of spending time in selecting a best suitable object, lightings, angles, position, etc.

[0094] Furthermore, the present disclosure enables generation of the shadows 201a in the media 201 in a single process. More specifically, the present disclosure enables the generation of the shadow in the media 201 in one click.

[0095] Additionally, the present disclosure is an automated solution, thereby eliminating manual efforts required in experiencing a professional photography.

[0096] Moreover, the present disclosure enhances the aesthetics of a plain scene in the media 201 by adding a contextually suitable shadow / s 201a in the media 201.

[0097] In a nutshell, the present disclosure enables generation of the shadow / s of interesting object(s) in the media. Further, the present disclosure enables addition of a lighting direction, intensity, pose, etc. to the shadows, thereby enhancing the user experience.

[0098] The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements. The elements can be at least one of a hardware device or a combination of hardware devices and software modules.

[0099] It is understood that terms including "unit" or "module" at the end may refer to the unit for processing at least one function or operation and may be implemented in hardware, software, or a combination of hardware and software.

[0100] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.

[0101] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.

[0102] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.

[0103] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.

[0104] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of at least one embodiment, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope of the embodiments as described herein.

[0105] The specific examples provided to explain the embodiments according to the present disclosure are merely a combination of each standard, method, detail method, and operation, and the various embodiments described herein can be performed through a combination of at least two or more techniques among the various techniques described. In addition, at this time, it can be performed according to a method determined through a combination of one or at least two or more of the aforementioned techniques. For example, it may be possible to perform a combination of parts of the operation of one embodiment with parts of the operation of another embodiment.

[0106] According to an aspect of the present disclosure, a method is disclosed. The method includes obtaining information associated with a media based on analyzing the media. Further, the method includes determining a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media. The method further includes identifying at least one object shadow to be generated based on one or more present existing objects in the media. Furthermore, the method includes generating a composite prompt using the identified at least one object shadow. Moreover, the method includes generating one or more shadows in the media based on processing the composite prompt and the obtained information.

[0107] In an embodiment, the information comprises at least one of a depth map, the light characteristic, or the one or more present existing objects associated with the media.

[0108] In an embodiment, the determining the CCT comprises splitting a window associated with the media into a plurality of sub-windows based on a resolution of the media. In an embodiment, the determining the CCT comprises computing a mean value associated with the plurality of sub-windows. In an embodiment, the determining the CCT comprises identifying the corresponding one or more light sources illuminating the scene based on the computed mean value. In an embodiment, the determining the CCT comprises determining the CCT corresponding to the identified one or more light sources, wherein the CCT indicates a measure of a colour appearance of the one or more light sources.

[0109] In an embodiment, the identifying at least one object shadow to be generated comprises determining a spatial relationship between the one or more present existing objects in the media, wherein the spatial relationship indicates a relative position, a containment, an interaction, an orientation, and a distance of at least one object with respect to other objects among the one or more present existing objects in the media. In an embodiment, the identifying at least one object shadow to be generated comprises generating the scene graph using the AI model based on the determined spatial information, wherein the scene graph indicates a visual representation of a relation between the one or more present existing objects in the media. In an embodiment, the identifying at least one object shadow to be generated comprises identifying at least one object shadow to be generated based on the scene graph.

[0110] In an embodiment, the generating the composite prompt comprises identifying at least one of positive shadow objects and negative shadow objects associated with the identified at least one object shadow to be generated. In an embodiment, the generating the composite prompt comprises generating the composite prompt based on the at least one of the identified positive shadow objects and the negative shadow objects, the light characteristics, or the CCT.

[0111] In an embodiment, the identifying the at least one of positive shadow objects and the negative shadow objects comprises selecting one or more shadow objects based on obtaining a list of defined shadow objects from a database, wherein the one or more shadow objects indicate objects having a relation with the scene within the media. In an embodiment, the identifying the at least one of positive shadow objects and the negative shadow objects comprises identifying the at least one of positive shadow objects and the negative shadow objects among the selected one or more shadow objects using the AI model.

[0112] In an embodiment, the identifying the at least one of positive shadow objects comprises using user input which is indicating objects so as to generate the shadow a user wants to.

[0113] According to an aspect of the present disclosure, an electronic device is disclosed. The electronic device includes memory storing instructions. The electronic device further includes at least one processor comprising processing circuitry, in communication with the memory, wherein the instructions, when executed by the at least one processor individually or collectively. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to obtain information associated with a media based on analyzing the media. Further, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to determine a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media. The instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify at least one object shadow to be generated based on one or more present existing objects in the media. Furthermore, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate a composite prompt using the identified at least one object shadow. Moreover, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate one or more shadows in the media based on processing the composite prompt and the obtained information.

[0114] In an embodiment, the information comprises at least one of the depth map, the light characteristic, or the one or more present existing objects associated with the media.

[0115] In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to split a window associated with the media into a plurality of sub-windows based on a resolution of the media. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to compute a mean value associated with the plurality of sub-windows. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify the corresponding one or more light sources illuminating the scene based on the computed mean value. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to determine the CCT corresponding to the identified one or more light sources, wherein the CCT indicates a measure of a colour appearance of the one or more light sources.

[0116] In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to determine a spatial relationship between the one or more present existing objects in the media, wherein the spatial relationship indicates a relative position, a containment, an interaction, an orientation, and a distance of at least one object with respect to other objects among the one or more present existing objects in the media. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate the scene graph using the AI model based on the determined spatial information, wherein the scene graph indicates a visual representation of a relation between one or more present existing objects in the media. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify at least one object shadow to be generated based on the scene graph.

[0117] In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify at least one of positive shadow objects and negative shadow objects associated with the identified at least one object shadow to be generated. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate the composite prompt based on at least one of the identified positive shadow objects and the negative shadow objects, the light characteristics, or the CCT.

[0118] In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to select one or more shadow objects based on obtaining a list of defined shadow objects from a database, wherein the one or more shadow objects indicate objects having a relation with the scene within the media. In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to identify at least one of the positive shadow objects and the negative shadow objects among the selected one or more shadow objects using the AI model.

[0119] In an embodiment, the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to use user input which is indicating objects so as to generate the shadow a user wants to.

Claims

1.A method (300) comprising:obtaining information associated with a media based on analyzing the media;determining a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media;identifying at least one object shadow to be generated based on one or more present existing objects in the media;generating a composite prompt using the identified at least one object shadow; andgenerating one or more shadows in the media based on processing the composite prompt and the obtained information.2.The method (300) of claim 1, wherein the information comprises at least one of a depth map, the light characteristic, or the one or more present existing objects associated with the media.3.The method (300) any one of claims 1 to 2, wherein the determining the CCT comprises:splitting a window associated with the media into a plurality of sub-windows based on a resolution of the media;computing a mean value associated with the plurality of sub-windows;identifying the corresponding one or more light sources illuminating the scene based on the computed mean value; anddetermining the CCT corresponding to the identified one or more light sources, wherein the CCT indicates a measure of a colour appearance of the one or more light sources.4.The method (300) any one of claims 1 to 3, wherein the identifying at least one object shadow to be generated comprises:determining a spatial relationship between the one or more present existing objects in the media, wherein the spatial relationship indicates a relative position, a containment, an interaction, an orientation, and a distance of at least one object with respect to other objects among the one or more present existing objects in the media;generating the scene graph using an artificial intelligence (AI) model based on the determined spatial information, wherein the scene graph indicates a visual representation of a relation between the one or more present existing objects in the media; andidentifying at least one object shadow to be generated based on the scene graph.5.The method (300) any one of claims 1 to 4, wherein the generating the composite prompt comprises:identifying at least one of positive shadow objects and negative shadow objects associated with the identified at least one object shadow to be generated; andgenerating the composite prompt based on the at least one of the identified positive shadow objects and the negative shadow objects, the light characteristics, or the CCT.6.The method (300) of claim 5, wherein the identifying the at least one of positive shadow objects and the negative shadow objects comprises:selecting one or more shadow objects based on obtaining a list of defined shadow objects from a database, wherein the one or more shadow objects indicate objects having a relation with the scene within the media; andidentifying the at least one of positive shadow objects and the negative shadow objects among the selected one or more shadow objects using an artificial intelligence (AI) model.7.The method (300) of claim 5, wherein the identifying the at least one of positive shadow objects comprises using user input which is indicating objects so as to generate the shadow a user wants to.8.An electronic device (100) comprising:memory (102) storing instructions; andat least one processor (106) comprising processing circuitry, in communication with the memory (102), wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to:obtain information associated with a media based on analyzing the media;determine a correlated color temperature (CCT) corresponding to one or more light sources illuminating a scene associated with the media;identify at least one object shadow to be generated based on one or more present existing objects in the media;generate a composite prompt using the identified at least one object shadow; andgenerate one or more shadows in the media based on processing the composite prompt and the obtained information.9.The electronic device (100) of claim 8, wherein the information comprises at least one of the depth map, the light characteristic, or the one or more present existing objects associated with the media;10.The electronic device (100) any one of claims 8 to 9, wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to:split a window associated with the media into a plurality of sub-windows based on a resolution of the media;compute a mean value associated with the plurality of sub-windows;identify the corresponding one or more light sources illuminating the scene based on the computed mean value; anddetermine the CCT corresponding to the identified one or more light sources, wherein the CCT indicates a measure of a colour appearance of the one or more light sources.11.The electronic device (100) any one of claims 8 to 10, wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to:determine a spatial relationship between the one or more present existing objects in the media, wherein the spatial relationship indicates a relative position, a containment, an interaction, an orientation, and a distance of at least one object with respect to other objects among the one or more present existing objects in the media;generate the scene graph using the AI model based on the determined spatial information, wherein the scene graph indicates a visual representation of a relation between one or more present existing objects in the media; andidentify at least one object shadow to be generated based on the scene graph.12.The electronic device (100) of claims 8 to 11, wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to:identify at least one of positive shadow objects and negative shadow objects associated with the identified at least one object shadow to be generated; andgenerate the composite prompt based on at least one of the identified positive shadow objects and the negative shadow objects, the light characteristics, or the CCT.13.The electronic device (100) of claim 12, wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to:select one or more shadow objects based on obtaining a list of defined shadow objects from a database, wherein the one or more shadow objects indicate objects having a relation with the scene within the media; andidentify at least one of the positive shadow objects and the negative shadow objects among the selected one or more shadow objects using the AI model.14.The electronic device (100) of claim 12, wherein the instructions, when executed by the at least one processor (106) individually or collectively, cause the electronic device (100) to use user input which is indicating objects so as to generate the shadow a user wants to.15.A machine readable medium comprising instructions that when executed by at least one processor individually or collectively, cause an electronic device to perform the method of any one of claims 1 to 7.