Method and electronic device for generating sequences of media
The electronic device generates personalized, high-quality short-form media by selecting transition objects and applying transformation effects, addressing the time and expertise challenges of traditional media creation, resulting in engaging and seamless media sequences.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-03-27
- Publication Date
- 2026-05-21
AI Technical Summary
Creating professional-quality short-form media is time-consuming and requires technical expertise, leading to generic and low-quality media output when automated tools are used, hindering user engagement and experience.
An electronic device and method that automatically generates seamless sequences of media by selecting transition objects, generating follow-up media based on spatial-temporal attributes, and applying transformation effects, using object detection and a generative model to enhance creativity and quality.
Enables the generation of personalized, high-quality, and engaging short-form media with seamless transitions, overcoming the limitations of manual and automated media creation tools.
Smart Images

Figure KR2025003946_21052026_PF_FP_ABST
Abstract
Description
METHOD AND ELECTRONIC DEVICE FOR GENERATING SEQUENCES OF MEDIA
[0001] The present disclosure relates to media, and more particularly, to a method and an electronic device for generating one or more sequences of media.
[0002] The information provided in this section is for background purposes only and is not necessarily considered as constituting prior art(s) to the present disclosure.
[0003] Media that is played on a device such as a television, a smartphone, or other electronic device may require the creation of the media. The creation of media may involve creating a script, shooting content, editing content, designing graphical elements, and other related tasks.
[0004] In today's digital world, users, such as creators, are always looking for new ways to create engaging media, such as short-form media that captivates an audience. From social media to educational platforms, captivating visuals are essential for grabbing attention and effectively conveying information. The audience is increasingly drawn towards the short-form media due to its engaging nature and fast-paced delivery of information. Many media platforms showcase the immense popularity of short-form media. This has led to the development of media editing tools that enable users to create the media.
[0005] In these media editing tools, users can add images, text, audio, animation, and video clips, arrange them on a timeline in a specific order, and apply effects such as narration, transitions between visual elements such as scenes or images, synchronization between audio and visuals, soundtrack creation, and more. However, these tasks are typically performed manually, and creating professional-quality media may require significant time, resources, and expertise. This presents a challenge for users and businesses in manually creating the media.
[0006] More specifically, creating media, particularly short-form media, using traditional media editing tools often requires a significant time investment, as scripting, filming, editing, and incorporating creative elements can be a lengthy and labor-intensive process.
[0007] Further, many users lack the technical expertise that is required to create creative media. This creates a barrier for users to leverage the power of the media but lack the necessary skills.
[0008] Therefore, nowadays users are switching to automated media tools, which help them create media much faster. However, media created using automated media tools is very generic and lacks creativity. Further, users need to compromise on the quality of the media created, which degrades user experience.
[0009] In a nutshell, an alternative solution is needed to overcome the above-discussed limitations and provide an improved method and an electronic device for generating one or more sequences of the media.
[0010] The drawbacks, difficulties, disadvantages, limitations of the conventional techniques explained in the background section are provided for exemplary purposes only, and the disclosure does not limit its scope to these limitations. A person skilled in the art would understand that this disclosure and the following description may also solve other problems or overcome additional drawbacks and disadvantages.
[0011] According to an aspect of the present disclosure, a method performed by an electronic device for generating one or more sequences of media is disclosed. The method may include obtaining ongoing media currently being played on the electronic device. The method may include selecting at least one transition object as a region of interest (ROI) in the ongoing media based on at least one of a user input or an object detection. The method may include generating follow-up media upon determining that the follow-up media is unavailable in a Database (DB), based on a description of the at least one selected transition object. The follow-up media may indicate a content sequence aligned with spatial-temporal attributes of the at least one selected transition object. The method may include generating a transformation effect based on motion information. The transformation effect may indicate a transition of one or more objects present in an outgoing frame of the ongoing media to one or more objects present in an incoming frame of the follow-up media. The method may include displaying the follow-up media with the transformation effect, thereby generating one or more sequences of the media with seamless crossfading.
[0012] According to an aspect of the present disclosure, an electronic device for generating one or more sequences of media is disclosed. The electronic device may include a memory configured to store at least one instruction of a computer program. The electronic device may include at least one processor in communication with the memory and configured to execute the at least one instruction to obtain ongoing media currently being played on the electronic device. The at least one processor is configured to select at least one transition object as a region of interest (ROI) in the ongoing media based on at least one of a user-input or an object detection. The at least one processor is configured to generate follow-up media upon determining that the follow-up media is unavailable in a Database (DB), based on a description of the at least one selected transition object. The follow-up media may indicate a content sequence aligned with spatial-temporal attributes of the at least one selected transition object. The at least one processor is configured to generate a transformation effect based on motion information. The transformation effect may indicate a transition of one or more objects present in an outgoing frame of the ongoing media to one or more objects present in an incoming frame of the follow-up media. The at least one processor is configured to display the follow-up media with the transformation effect, thereby generating one or more sequences of the media with seamless crossfading.
[0013] According to an embodiment, there is provided a computer readable storage medium having stored thereon a computer program that, when executed by at least one processor, performs the method.
[0014] To further clarify the advantages and features of the present invention, a more particular description of the disclosure will be rendered by reference to an embodiment thereof, which is illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered as limiting its scope.
[0015] In order to more clearly explain the technical solutions in an embodiment of the disclosure, the accompanying drawings to be used in the description of the embodiment of the disclosure will be briefly described below.
[0016] Figure 1 illustrates a schematic block diagram of an electronic device for generating one or more sequences of media, in accordance with an embodiment of the present disclosure;
[0017] Figure 2 illustrates a schematic block diagram depicting a plurality of modules included in the electronic device, in accordance with an embodiment of the present disclosure;
[0018] Figure 3illustrates a flowchart depicting a method for generating one or more sequences of media, in accordance with an embodiment of the present disclosure;
[0019] Figure 4 illustrates a flowchart depicting an operation for selecting at least one transition object based on a user input, in accordance with an embodiment of the present disclosure;
[0020] Figure 5 illustrates a flowchart depicting an operation for selecting at least one transition object based on object detection in accordance with an embodiment of the present disclosure;
[0021] Figure 6illustrates a schematic flow diagram depicting generation of an object-level description, in accordance with an embodiment of the present description;
[0022] Figure 7 illustrates a schematic flow diagram for generating a spatial-temporal description in accordance with an embodiment of the present disclosure;
[0023] Figure 8illustrates a flowchart depicting an operation for generating follow-up media, in accordance with an embodiment of the present disclosure;
[0024] Figure 9illustrates a schematic flow diagram depicting a generation of the action and an inferred object, in accordance with an embodiment of the present disclosure;
[0025] Figure 10illustrates a flowchart depicting an operation for generating a transformation effect, in accordance with an embodiment of the present disclosure;
[0026] Figures 11 illustrates a first exemplary use case depicting generation of follow-up media, in accordance with an embodiment of the present disclosure;
[0027] Figures 12A-12B illustrate a second exemplary use case depicting generation of follow-up media based on a user, in accordance with an embodiment of the present disclosure;
[0028] Figure 13illustrates a third exemplary use case depicting generation of one or more sequences of a screensaver, in accordance with an embodiment of the present disclosure; and
[0029] Figure 14illustrates a fourth exemplary use case depicting generation of follow-up media based on user preference, in accordance with an embodiment of the present disclosure.
[0030] Reference will be made in detail to example embodiments of the present disclosure, examples of which are illustrated in the drawings, wherein like reference numerals refer to like elements throughout the drawings.
[0031] The terms used in the present disclosure will be briefly described, and then embodiments of the present disclosure will be described in detail.
[0032] The terms used in the present disclosure are general terms as much as possible and have been widely used nowadays in consideration of the functions of the present disclosure, which, however, may be changed according to an intention of a technician in the art, a precedent, the advent of new technologies, or the like. Also, particular cases may include terms arbitrary selected by an applicant, and in this case, the meaning of the terms will be described in detail in the corresponding description. Therefore, the terms used in the present disclosure should be defined based on the meanings of the terms and the content throughout the present disclosure, rather than simply based on the titles of the terms.
[0033] It may be advantageous to set forth definitions of certain words and phrases used throughout this disclosure. Thus, the terms "transmit," "receive," and "communicate," as well as derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, have a relationship to or with, or the like.
[0034] Moreover, various functions described below may be implemented or supported by one or more computer programs, each of which is formed from a computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer readable program code. The term "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The term "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.
[0035] As used here, terms and phrases such as "have," "may have," "include," or "may include" a feature (like a number, function, operation, or component such as a part) indicate the existence of the feature and do not exclude the existence of other features. Also, as used here, the phrases "A or B," "at least one of A and / or B," or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B," "at least one of A and B," and "at least one of A or B" may indicate all of (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Further, as used here, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the order or importance of the devices. A first component may be denoted a second component and vice versa without departing from the scope of this disclosure.
[0036] It will be understood that, when an element (such as a first element) is referred to as being (operatively or communicatively) "coupled with / to" or "connected with / to" another element (such as a second element), the element can be coupled or connected with / to the other element directly or via a third element. In contrast, it will be understood that, when an element (such as a first element) is referred to as being "directly coupled with / to" or "directly connected with / to" another element (such as a second element), no other element (such as a third element) intervenes between the element and the other element.
[0037] As used here, the phrase "configured (or set) to" may be interchangeably used with the phrases "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of" depending on the circumstances. The phrase "configured (or set) to" does not essentially mean "specifically designed in hardware to." Rather, the phrase "configured to" may mean that a device can perform an operation together with another device or parts. For example, the phrase "processor configured (or set) to perform A, B, and C" may mean a generic-purpose processor (such as a CPU or application processor) that may perform the operations by executing one or more software programs stored in a memory device or a dedicated processor (such as an embedded processor) for performing the operations.
[0038] The terms and phrases as used here are provided merely to describe some embodiments of this disclosure but not to limit the scope of other embodiments of this disclosure. It is to be understood that the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. All terms and phrases, including technical and scientific terms and phrases, used here have the same meanings as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure belong. It will be further understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined here. In some cases, the terms and phrases defined here may be interpreted to exclude embodiments of this disclosure.
[0039] Examples of an "electronic device" according to embodiments of this disclosure may include at least one of a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothes, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the "electronic device" include a smart home appliance. Examples of the smart home appliance may include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washer, a dryer, an air cleaner, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC), a smart speaker or speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME), a gaming console, an electronic dictionary, an electronic key, a camcorder, or an electronic picture frame. Still other examples of the "electronic device" include at least one of various medical devices (such as diverse portable medical measuring devices (like a blood sugar measuring device, a heartbeat measuring device, or a body temperature measuring device), a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an automotive infotainment device, a sailing electronic device (such as a sailing navigation device or a gyro compass), avionics, security devices, vehicular head units, industrial or home robots, automatic teller machines (ATMs), point of sales (POS) devices, or Internet of Things (IoT) devices (such as a bulb, various sensors, electric or gas meter, sprinkler, fire alarm, thermostat, street light, toaster, fitness equipment, hot water tank, heater, or boiler). Other examples of the "electronic device" include at least one part of a piece of furniture or building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of this disclosure, the "electronic device" may be one or a combination of the above-listed devices. According to some embodiments of this disclosure, the "electronic device" may be a flexible electronic device. The electronic device disclosed here is not limited to the above-listed devices and may include new electronic devices depending on the development of technology.
[0040] In the following description, electronic devices are described with reference to the accompanying drawings according to various embodiments of this disclosure. As used here, the term "user" may denote a human or another device (such as an artificial intelligent electronic device) using the electronic device.
[0041] Definitions for other certain words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many if not most instances, such definitions apply to prior as well as future uses of such defined words and phrases.
[0042] Throughout the present disclosure, expressions such as "at least one of a, b or c", "at least one of a, b, or c", "at least one of a, b and c", "at least one of a, b, and c" may indicate "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof. Similarly, expressions such as "at least one of a or b", "at least one of a, or b", "at least one of a and b", "at least one of a, and b " may indicate "a," "b,", "a and b," or variations thereof.
[0043] In this disclosure, the expression "and / or" includes a combination of a plurality of described components or any component of the plurality of described components. In this disclosure, terms such as "1st," "2nd," "first," and "second" may be merely used to distinguish a corresponding component from other corresponding components and do not limit the corresponding components in terms of other aspects (for example, the degree of importance or the order).
[0044] Throughout the present disclosure, when a part "includes" or "comprises" an element, the part may further include other elements, not excluding the other elements, unless there is a particular description contrary thereto.
[0045] Also, the terms "portion," "module," etc. described in this disclosure denote a unit configured to process at least one function or operation, and the "portion," and the "module" may be realized as hardware or software, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), or a combination of the hardware and the software. The term "portion" used in an embodiment of the present disclosure does not have a meaning limited to software or hardware. A "portion" described in the present disclosure may be configured to be in a storage medium which may be addressed or may be configured to play one or more processors. According to an embodiment of the present disclosure, a "portion" may include components, such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, sub-routines, segments of a program code, drivers, firmware, a microcode, a circuit, data, a database, data structures, tables, arrays, and variables. Functions provided through a predetermined component or a predetermined "portion" may be combined to reduce the number of functions or may be divided into additional components. Also, according to an embodiment, a "portion" may include one or more processors.
[0046] According to an embodiment of the present disclosure, each of blocks of the flowcharts and combinations of the flowcharts may be performed by computer program instructions. The computer program instructions may be loaded on a general-purpose computer, a specialized computer, or a processor of other programmable data processing device. The instructions performed through the computer or the processor of the other programmable data processing device may generate a medium for performing the functions described in the flowchart block(s). The computer program instructions may also be stored in a computer-available or computer-readable memory oriented for the computer or the other programmable data processing device in order to realize the functions in a predetermined way. The instructions stored in the computer-available or computer-readable memory may also produce manufacturing items embedding an instruction medium for performing the functions described in the flowchart block(s). The computer program instructions may also be loaded on the computer or the other programmable data processing device.
[0047] Throughout this disclosure, the term "electronic device" may refer to a system that includes functions for generating one or more sequences of media. The term "media" as used herein, may include "short-form video media". The short-form video media may refer to media content that consists of or includes short-form videos, typically lasting a few seconds to a few minutes, but is not limited to.
[0048] The present disclosure may provide a method and an electronic device for generating one or more sequences of short-form video media with seamless crossfading between the one or more sequences of short-form video media. According to the present disclosure, personalized, high-quality, and engaging short-form video media can be automatically generated based on the media a user is viewing. The present disclosure may extract interesting segments from an original media, generate a new video clip (or snippet), and seamlessly connect them to ensure smooth transition. The present disclosure may automatically generate extended short-form video media based on the selected segments.
[0049] Figure 1 illustrates a schematic block diagram of an electronic device 100 for generating one or more sequences of media, in accordance with an embodiment of the present disclosure.
[0050] In an embodiment, the electronic device 100 may include a memory 102 including a Database (DB) 104, a processor 106 communicatively coupled with the memory 102, an Input / Output (I / O) interface 110, and a plurality of modules 120. In an embodiment, the electronic device 100 may be implemented by a User Equipment (UE). The UE may be a smartphone, a laptop computer, a desktop computer, a Personal Computer (PC), a notebook, a tablet, a smartwatch, or any other device in which a user may view an ongoing media, but is not limited thereto.
[0051] In an embodiment, the electronic device 100 may operate in conjunction with a cloud-based system, which may include a server, such as a cloud server. The electronic device 100 may perform all operations according to the present disclosure. The electronic device 100 may perform a portion of the operations, while the cloud server may perform the remaining operations according to the present disclosure. In an embodiment, the electronic device 100 may be include in a system that may include a server. The system may be implemented as a combination of the electronic device 100 and the server. The system may perform one or more operations according to the present disclosure using the electronic device 100, while the remaining operations may be performed by the server.
[0052] In an embodiment, the memory 102 may be configured to store at least one instruction of a computer program executable by the processor 106. In an embodiment, the memory 102 may communicate via a bus within the electronic device 100. The memory 102 may include but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like. For example, the memory 102 may include a cache or random-access memory (RAM) for the processor 106. For example, the memory 102 may be separate from the processor 106. For example, the memory 102 may be a cache memory of the processor 106, system memory, or other memory. The memory 102 may be an external storage device, which is used for storing data. The memory 102 may be operable to store the at least one instruction executable by the processor 106. The functions, acts, or tasks illustrated in the figures or described may be performed by the programmed processor for executing the at least one instruction stored in the memory 102. The functions, acts, or tasks may be independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code, and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing, and the like.
[0053] The processor 106 may be a single processing unit or a set of units each including multiple computing units. The processor 106 may be referred to as at least one processor or one or more processors. The processor 106 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions (computer-readable instructions or the at least one instruction) stored in the memory 102. The processor 106 may be configured to fetch and execute computer-readable instructions (or the at least one instruction) and data stored in the memory 102. The processor 106 may include one or a plurality of processors. The plurality of processors may be implemented as a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The plurality of processors may control the processing of the input data in accordance with a predefined operating rule or an artificial intelligence (AI) model stored in the memory 102. The predefined operating rule or the AI model may be provided through training or learning. The processor 106 according to an embodiment of the disclosure may include at least one circuitry, such as a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a many integrated core (MIC), a digital signal processor (DSP), or a neural processing unit (NPU).
[0054] The processor 106 may be communicate with one or more input / output (I / O) devices via the Input / Output (I / O) interface 110. The I / O interface 110 may employ communication technologies such as Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System for Mobile Communications (GSM), Long-Term Evolution (LTE), WiMax, and the like, etc. In an embodiment of the present invention, the I / O interface 110 may employ ethernet, industrial wireless Local Area Network (LAN), Process Field Bus (PROFIBUS), Actuator Sensor (AS) Interface, and the like. The electronic device 100 may be configured such that the modules 120 are included in the processor 106.
[0055] Figure 2 illustrates a schematic block diagram depicting the plurality of modules 120 included in the electronic device 100, in accordance with an embodiment of the present disclosure. The plurality of modules 120 may include the one or more instructions that may be executed to cause the electronic device 100, in particular, the processor 106 of the electronic device 100, to execute the one or more instructions.
[0056] The plurality of modules 120 may include an obtaining module 122, a selecting module 124, a generating module 126, a determining module 128, and a displaying module 130. In an embodiment, the obtaining module 122, the selecting module 124, the generating module 126, the determining module 128, and the displaying module 130 may communicate with each other. In an embodiment, the plurality of modules 120 may be configured to perform various operations or steps that may be discussed and explained in detail in conjunction with Figures 3-10.
[0057] A detailed explanation of various functions of the processor 106 and / or the plurality of modules 120 may be explained in view of Figures 3-10. Figures 3-10 may be performed by the electronic device 100 or by the processor 106 included in the electronic device 100.
[0058] Figure 3 illustrates a flowchart depicting a method 300 for generating the one or more sequences of the media, in accordance with an embodiment of the present disclosure. In an embodiment, the method 300 is a computer-implemented method 300 that is explained in detail in the below paragraphs.
[0059] Referring to Figure 3, the method 300 may begin with operation 302, which may include obtaining, via the obtaining module 122, the ongoing media currently being played on the electronic device 100. In an embodiment, the method 300 may begin with operation 302, which may include obtaining, via the obtaining module 122, the ongoing media currently being played on an external device. In an embodiment, the ongoing media may include a video, images, or the like. The media may be a short-form media or a long-form media.
[0060] At operation 304, the method 300 may include selecting, via the selecting module 124, at least one transition object as a region of interest (ROI) in the ongoing media. In an embodiment, the selecting module 124 may select the at least one transition object based on a user input which is explained in conjunction with Figure 4. In an embodiment, the selecting module 124 may select the at least one transition object based on an object detection which is explained in conjunction with Figure 5.
[0061] Figure 4 illustrates a flowchart depicting the operation for selecting the at least one transition object based on the user input, in accordance with an embodiment of the present disclosure.
[0062] At operation 304a, the operation 304 may include obtaining, via the obtaining module 122, the user input that indicates an identification of the at least one transition object as the ROI. the user may select the at least one transition object among one or more objects that may be present in the media.
[0063] At operation 304b, the operation 304 may include determining, via the determining module 128, the availability of the follow-up media in the DB 104.
[0064] At operation 304c, the operation 304 may include retrieving the follow-up media other than the ongoing media from the DB 104 based on the user input.
[0065] Figure 5 illustrates a flowchart depicting the operation for selecting the at least one transition object based on the object detection, in accordance with an embodiment of the present disclosure.
[0066] At operation 304-1, the operation 304 may include generating an object-level description based on detecting the one or more objects in each of one or more frames of the ongoing media. In an embodiment, the object-level description may herein indicate a caption sequence describing the one or more objects.
[0067] Figure 6 illustrates a schematic flow diagram 600 depicting a generation of the object-level description, in accordance with an embodiment of the present description.
[0068] In an embodiment, at block 602, the one or more frames may be extracted from the ongoing media at consistent intervals to balance temporal resolution and computational efficiency. Further, at block 604, the selecting module 124 may include a deep neural model, such as a Recurring Convolutional Neural Network (R-CNN), for detecting one or more objects and extracting features from the one or more frames. At block 606, a Region Proposal Network (RPN) is used to generate bounding boxes for object proposals within the one or more frames. In an embodiment, an object-level attribute classification is performed for each of the one or more objects detected at block 606 to obtain feature vectors, such as appearance features and geometric features associated with each of the one or more objects.
[0069] At block 608, the selecting module 124 may include an object-transforming model that may include an encoder and a decoder that may be composed of a stack of layers to process the feature vectors to generate a sequence of words as outputs, which is referred to as the object-level description. For example, if a cricket match is progress, the object-level description may include "batter ready to face the ball", "runner in position", "bowler in action", "umpire standing" or "cricket pitch".
[0070] In an embodiment, apart from the object-level description, a scene-level description may also be generated for the one or more objects. For example, if the cricket match is in progress and the batter is one of the one or more objects, then the scene-level description may be, for example, "the cricket match is in action, and the batter is ready to face the bowler.
[0071] At operation 304-2, the operation 304 may include generating, via the generating module 126, a spatial-temporal description based on the object-level description and spatial-temporal attributes of the one or more objects. In an embodiment, the spatial-temporal description may include one or more of the following: a dynamic behavior, interactions, and positional relationships of the one or more objects over time, incorporating motion patterns, attribute changes, and contextual spatial relationships.
[0072] Figure 7 illustrates a schematic flow diagram 700 for generating the spatial-temporal description, in accordance with an embodiment of the present disclosure. At block 702, the generating module 126 may include a spatial-temporal model that may obtain the object-level description and / or the scene-level description as input to generate the spatial-temporal description. The spatial-temporal model may analyze the object-level description and the scene-level description to obtain object attributes associated with the one or more objects. The object attributes may include but are not limited to a position, an appearance, an orientation, a scale, and an activity associated with the one or more objects. The object attributes may include an interaction of the one or more objects with each other. In a scenario, the determining module 128 may determine a position or a distance between the one or more objects (e.g., the object is on top of, next to, or behind another object). The scene-level description may provide additional information about the one or more objects such as the activity of the one or more objects. Based on the observation that one or more objects which may be close to each other are more likely to be correlated. This information may be incorporated into a spatial graph by connecting the one or more objects using a corresponding normalized Intersection over Union (IoU) value.
[0073] The spatial-temporal model may include a cross-modal multi-head attention that may capture a significance of the one or more objects in the one or more frames of the media and a context associated with the media. The spatial-temporal model may be adapted to introduce sentence-level and word-level consistency constraints to enhance the object-level description.
[0074] At block 704, the generating module 126 may include a description decoding model that may be used to update the object-level description. The description decoding model may refine the object-level description, thereby obtaining the spatial-temporal description.
[0075] At operation 304-3 of Figure 5, the operation 304 may include computing, via the determining module 128, an object-level score for each of the one or more objects. In an embodiment, the determining module 128 may correlate weighted factors associated with the spatial-temporal description and the spatial-temporal attributes to compute the object-level score for each of the one or more objects. In an embodiment, the spatial-temporal attributes may include an object size, an object appearance, a surrounding context, motion dynamics, and a semantic relevance. The computation of the object-level score is shown using equation (1) below:
[0076]
[0077] where Size represents a size of an object region,
[0078] Appearance refers to the visual appearance of each of the one or more objects across the one or more frames,
[0079] Context corresponds to a surrounding context of the object region,
[0080] Motion indicates movement or dynamics associated with the object region, and
[0081] Semantic relevance measures the semantic relevance of each object to the media.
[0082] The object-level score of each object may be compared with a predefined video transition threshold. The predefined video transition threshold may indicate a transition factor that is preset for a media transition associated with the media. Thereafter, the at least one transition object may be selected when the object-level score for the at least one transition object exceeds the predefined video transition threshold as shown in operation 304-4 of Figure 5. For example, when the predefined transition score for the media is 0.6, and object A and object B have the object-level scores of 0.5 and 0.7 respectively, then the object B is selected as the at least one transition object.
[0083] In an embodiment, the predefined video transition threshold may also be learned by a user-defined transition or a manually labeled dataset.
[0084] In an embodiment, the method 300 may include retrieving a follow-up media from the DB 104 based on the availability of the follow-up media in the DB 104. The follow-up media may herein refer to a content sequence aligned with the spatial-temporal attributes of the at least one selected transition object.
[0085] In an embodiment, the follow-up media may not be available on the DB 104. In this case, the method 300 may follow operation 306.
[0086] At operation 306, the method 300 may include generating, via the generating module 126, the follow-up media upon determining that the follow-up media is unavailable in the DB 104, based on a description of at least one selected transition object. The follow-up media may indicate a content sequence aligned with spatial-temporal attributes of the at least one selected transition object. In an embodiment, the generation of the follow-up media may be explained in conjunction with Figure 8.
[0087] Figure 8 illustrates a flowchart depicting operation for generating the follow-up media, in accordance with an embodiment of the present disclosure.
[0088] At operation 306a, the operation 306 may include generating, via the generating model 126, an action and an inferred object based on the at least one selected transition object. In an embodiment, the action may indicate herein an extrapolated interaction behavior. The inferred object may indicate a semantically consistent replacement of the at least one selected transition object.
[0089] Figure 9 illustrates a schematic flow diagram 900 depicting the generation of the action and the inferred object, in accordance with an embodiment of the present disclosure.
[0090] At block 902, the generating module 126 may use a region of interest (ROI) tracker for tracking a region of interest (RoI) associated with the at least one selected transition object. In an embodiment, the ROI tracker may use but is not limited to, a feature-based Kanade-Lucas-Tomasi (KLT) tracker or other techniques like SORT for tracking.
[0091] At block 904, the generating module 126 may use one or more methods such as optical flow-based methods or CNN-based methods for extracting action information in the ROI associated with the at least one selected transition object.
[0092] At block 906, the generating module 126 may utilize transition object extraction techniques for extracting the at least one selected transition object in the ROI based on the extracted action information.
[0093] At block 908, the generating module 126 may use but is not limited to a statistical analysis model and a law of dynamics for extrapolating the extracted action information to obtain the extrapolated action.
[0094] At block 910, the generating module 126 may obtain the inferred object based on the at least one selected transition object, from the DB 104. The inferred object may have the motion dynamics similar to the at least one selected transition object. In an embodiment, the similar motion dynamics may herein refer to the same action or movement. In an embodiment, the DB 104 may be a pre-trained DB or lookup DB. In an embodiment, the at least one selected transition object may be replaced with the inferred object in the follow-up media to be generated.
[0095] At block 912, the generating module 126 may create a text summary corresponding to the extrapolated action to be applied to the inferred object.
[0096] At operation 306b, the operation 306 may include generating, via the generating module 126, the description based on the action and the inferred object, using a generative model. In an embodiment, the description may herein refer to a storyline with one or more contextually aligned sequences for the follow-up media. The action herein may also be termed as the extrapolated action within the scope of the present disclosure.
[0097] The generative model may generate a prompt based on merging the action and the inferred object. The generative model may expand the generated prompt based on the one or more objects, the action, and the context associated with the ongoing media. The generative model may use a Large Language Model (LLM) to fine-tune the generated prompt. The generative model may use a Retrieval-Augment Generation (RAG) technique to retrieve the relevant information from the fine-tuned generated prompt and combine the relevant information with the text to generate the description. In an embodiment, the description in real-time may be generated as equation (2) below:
[0098]
[0099] where s(t) = the description to be generated in real time,
[0100] X (t-1) corresponds to a state of story capturing the generated prompt, the context, and the extrapolated action; and
[0101] A(t) corresponds to the extrapolated action.
[0102] At operation 306c, the operation 306 may include determining, via the determining module 128, a surprise element based on at least one of a user data, theme adjustments, and contextual understanding of the ongoing media. In an embodiment, the surprise element may herein refer to contextually relevant thematic changes to enhance user engagement.
[0103] The determining module 128 may include a token manipulation model that may generate a media token with one or more new entities that may associated with the user. For example, the user's favorite t-shirt may be added as a new entity for the follow-up media to be generated. In an embodiment, the one or more entities may be provided by the user based on his / her preference. In an embodiment, the electronic device 100 may automatically retrieve the one or more entities from the electronic device 100 or an external device associated with the user. For example, the electronic device 100 may automatically retrieve the one or more entities from a gallery of the smartphone associated with the user. The smartphone may be the electronic device 100 or the external device. In an embodiment, the one or more entities may be retrieved based on the frequency of occurrence of such entities. In an embodiment, each entity may be captured in various viewpoints to make an integration of the one or more entities more robust.
[0104] The generative model may include a guidance text-generating model that may select at least one theme from a set of predefined themes based on the context associated with the follow-up media to be generated. For example, the generative model may generate the follow-up media in a style of a painting, enhance blue hues in the follow-up media, or generate a daytime scene, etc. In an embodiment, the at least one theme may be selected based on the user preference. In an embodiment, the at least one theme may be selected automatically.
[0105] The generative model may include a loss function generator that may choose an appropriate loss function among a set of predefined loss functions that may be utilized to guide the generation process to fit the at least one selected theme in the follow-up media to be generated, thereby determining the surprise element.
[0106] Referring to Figure 8, at operation 306d, the operation 306 may include generating, via the generating module 126, the follow-up media based on merging the surprise element, parsing the description, and animating the parsed description.
[0107] The generating module 126 may include a story-generating model that may parse the generated description into a set of chunks. The story-generating model may convert each chunk to a corresponding embedding and generate one or more subsequent frames.
[0108] The generating module 126 may include an animation-generating model that may animate content associated with each of the one or more subsequent frames, thereby enabling the generation of the follow-up media.
[0109] Referring to Figure 3, at operation 308, the method 300 may include generating, via the generating module 126, a transformation effect based on motion information associated with the one or more objects. In an embodiment, the transformation effect may herein refer to a transition of the one or more objects present in an outgoing frame of the ongoing media to the one or more objects present in an incoming frame of the follow-up media. In an embodiment, the generation of the transformation effect may be explained in conjunction with Figure 10.
[0110] Figure 10 illustrates a flowchart depicting the operation for generating the transformation effect, in accordance with an embodiment of the present disclosure.
[0111] At operation 308a, the operation 308 may include capturing the motion information of the one or more objects in the outgoing frame of the ongoing media and the incoming frame of the follow-up media. In an embodiment, the motion information may indicate the movement and the motion dynamics of the one or more objects between the ongoing media and the follow-up media.
[0112] At operation 308b, the operation308 may include computing, via the determining module 128, a spatial-temporal difference between the outgoing frame of the ongoing media and the incoming frame of the follow-up media based on the motion information. The determining module 128 may compute denoising strength, a schedule, and a count of one or more intermediate frames associated with the follow-up media. In an embodiment, if the spatial-temporal difference increases, then the denoising strength may increase, the schedule may decrease and the count of the one or more intermediate frames may increase.
[0113] In an scenario, the computation of the denoising strength, the schedule, and the count of the one or more intermediate frames may be shown using equations (3)-(5) in the below paragraphs.
[0114] For example, assume that at least two objects among the one or more objects are in motion in both the incoming and the outgoing frames.
[0115]
[0116]
[0117]
[0118] Where = actual number of the objects that are in motion;
[0119] the noise corresponds to denoising strength; and
[0120] a number of denoising steps may be utilized to compute the schedule for denoising the noise from the follow-up media.
[0121] At operation 308c, the operation 308 may include generating the transformation effect based on the spatial-temporal difference such that the transformation effect seamlessly blends the content of the ongoing media and the follow-up media thereby ensuring visual continuity.
[0122] The noise parameter, the number of denoising steps, may be correlated to calculate the embedding of the outgoing frame (oe) and the embedding of the incoming frame (ie) using equation (6) as shown below:
[0123]
[0124] where corresponds to resultant embedding for intermediate frame
[0125] = embedding of the outgoing frame
[0126] = embedding of the incoming frame,
[0127] and are weight factors for intermediate frame corresponding to the incoming frame and the outgoing frame respectively,
[0128] Where,
[0129]
[0130] In an embodiment, a weightage value associated with the weight factors may be normalized between 0 and 1, which means that w1 and w2 may be between 0 to 1 and their total sum may be equal to 1. If w1 increases, w2 decreases accordingly.
[0131] The embedding of the outgoing frame and the incoming frame may be used to generate the one or more intermediate frames using an image diffusion technique, thereby generating the transformation effect.
[0132] At operation 310 of Figure 3, the method 300 may include displaying, via the displaying module 130, the follow-up media with the transformation effect, thereby generating the one or more sequences of the media with seamless crossfading.
[0133] In an embodiment, the method 300 may include obtaining, via the obtaining module 122, feedback based on the generated follow-up media from the user. In an embodiment, the feedback may be stored in the DB 104 which may be used later for the generation of the follow-up media.
[0134] In an embodiment, the generation of the follow-up media may get terminated if the user stops watching the ongoing media.
[0135] Figures 11 illustrates a first exemplary use case 1100 depicting the generation of the follow-up media, in accordance with an embodiment of the present disclosure. In the first exemplary use case 1100, the user is playing the ongoing video of the cricket match using the electronic device 100, in which at time T1, the batter hits the ball. In this case, the ball is selected as the transition object at time T6, when the ball starts falling, as illustrated in 1110 of Figure 11. The electronic device 100 analyzes the motion dynamics of the ball and generates a first follow-up media with the inferred object, i.e., a football, having the same motion dynamics, as illustrated in 1120 of Figure 11. At, time T5 of the first follow-up media (illustrated in 1120 of Figure 11), the running player is selected as the transition object, and a second follow-up media is generated after T7, as illustrated in 1130 of Figure 11. Referring to 1130 of Figure 11, the second follow-up media starts with a running rat. At time T7 of the second follow-up media, a falling cat is selected as the transition object, as illustrated in 1130 of Figure 11. A third follow-up media starts with a falling man, and the media is continued as illustrated in 1140 of Figure 11.
[0136] Figures 12A-12B illustrate a second exemplary use case 1200 depicting the generation of the follow-up media based on the user, in accordance with an embodiment of the present disclosure.
[0137] In the second exemplary use case 1200, the follow-up may be generated based on the user. In a scenario, the electronic device100 may identify the age of the user and generate the follow-up media based on the age of the user. Referring to Figure 12A, a user 1 is an adult, and the ongoing media is of a football match in which the player is running after the football. In this case, the follow-up media generated is of an athlete running. Referring to Figure 12B, a user 2 is a kid, and initially, the ongoing media is of the football match in which the player is running after the football. In this case, the follow-up media generated is of the running rat.
[0138] Figure 13 illustrates a third exemplary use case 1300 depicting the generation of the one or more sequences for the screensaver, in accordance with an embodiment of the present disclosure.
[0139] In the third exemplary use case 1300, the one or more sequences for the ongoing media such as screensaver are generated. In a scenario, the screen saver includes bubbles, in which the bubbles are selected as the transition object. In a first sequence (S1), the bubbles are replaced with the Earth. In the sequence (S1), the Earth is selected as the transition object which is replaced by the Sun in a second sequence (S2). In the second sequence (S2), the Sun is selected as the transition object and replaced by stars in a third sequence (S3). In the third sequence (S3), the stars are selected as the transition object and replaced by the galaxy in a fourth sequence (S4). In the fourth sequence (S4), the galaxy is selected as the transition object and replaced by buildings in a fifth sequence (S5).
[0140] Figure 14 illustrates a fourth exemplary use case 1400 depicting the generation of the follow-up media based on user preference, in accordance with an embodiment of the present disclosure.
[0141] In the fourth exemplary use case 1400, the user is a kid, and the ongoing media is played based on the user's preference such as an area of interest of the kid. In an embodiment, the kid may select a topic of number system among a plurality of topics. At first, the media is associated with natural numbers which is continued by the follow-up media of integers, rational numbers, real numbers, complex numbers, prime numbers, quaternion, octonion, cardinal numbers, exponential, infinity, and which is continued with the similar type of the follow-up media.
[0142] In an embodiment, the present disclosure prioritizes user preferences and enhances the auto-generation of the follow-up media for the ongoing media. The present disclosure may ensure a seamless transition between the ongoing media and the follow-up media, thereby enhancing user experience.
[0143] In an embodiment, the selecting (304) of the at least one transition object based on the user-input comprises obtaining (304a) the user input indicating an identification of the at least one transition object as the ROI, determining (304b) the availability of the follow-up media, and retrieving (304c) the follow-up media other than the ongoing media from the DB (104), based on the user input.
[0144] In an embodiment, the selecting (304) of the at least one transition object based on the object detection comprises generating (304-1) an object-level description based on detecting the one or more objects in each of one or more frames of the ongoing media, wherein the object-level description indicates a caption sequence describing the one or more objects, generating (304-2) a spatial-temporal description based on the object-level description and the spatial-temporal attributes of the one or more objects, computing (304-3) object-level score for each of the one or more objects based on correlating weighted factors associated with the spatial-temporal description and the spatial-temporal attributes, and selecting (304-4) the at least one transition object as the ROI among the one or more objects in the ongoing media, wherein the object-level score for the at least one selected transition object exceeds a predefined video transition threshold.
[0145] In an embodiment, the spatial-temporal attributes comprise a size of an object, an object appearance, a surrounding context, motion dynamics, and a semantic relevance, and the spatial-temporal description includes one or more of a dynamic behavior, interactions, and positional relationships of the detected one or more objects over time, incorporating motion patterns, attribute changes, and contextual spatial relationships.
[0146] In an embodiment, the generating (306) of the follow-up media comprises generating (306a) an action and an inferred object based on the at least one selected transition object, wherein the action indicates extrapolated interaction behavior, the inferred object indicates a semantically consistent replacement of the at least one transition object, generating (306b) the description based on the action and the inferred object, using a generative model, wherein the description indicates a storyline with contextually aligned sequence for the follow-up media, determining (306c) a surprise element based on at least one of user data, theme adjustments, and contextual understanding of the ongoing media, wherein the surprise element indicates a contextually relevant thematic change to enhance user engagement, and generating (306d) the follow-up media based on merging the surprise element, parsing the description, and animating the parsed description.
[0147] In an embodiment, the action is summarized in a text description of a predicted activity for the inferred object.
[0148] In an embodiment, the generating (308) of the transformation effect comprises capturing (308a) motion information of the one or more objects in the outgoing frame of the ongoing media and the incoming frame of the follow-up media, wherein the motion information indicates a movement and dynamics of the one or more objects between the ongoing media and the follow-up media, computing (308b) spatial-temporal difference between the outgoing frame of the ongoing media and the incoming frame of the follow-up media based on the motion information, and generating (308c) the transformation effect based on the spatial-temporal difference such that the transformation effect seamlessly blends the content of the ongoing media and the follow-up media thereby ensuring visual continuity.
[0149] In an embodiment, to select the at least one transition object based on the user-input the at least one processor (106) is configured to obtain the user input indicating an identification of the at least one transition object as the ROI, determine the availability of the follow-up media, and retrieve the follow-up media other than the ongoing media from the DB (104), based on the user input.
[0150] In an embodiment, to select the at least one transition object based on the object detection, the at least one processor (106) is configured to generate an object-level description based on detecting the one or more objects in each of one or more frames of the ongoing media, wherein the object-level description indicates a caption sequence describing the one or more objects, generate a spatial-temporal description based on the object-level description and the spatial-temporal attributes of the one or more objects, compute an object-level score for each of the one or more objects based on correlating weighted factors associated with the spatial-temporal description and the spatial-temporal attributes, and select the at least one transition object as the ROI among the one or more objects in the ongoing media, wherein the object-level score for the at least one selected transition object exceeds a predefined video transition threshold.
[0151] In an embodiment, to generate the follow-up media, the at least one processor (106) is configured to generate an action and an inferred object based on the at least one selected transition object, wherein the action indicates extrapolated interaction behavior, the inferred object indicates a semantically consistent replacement of the at least one transition object, generate the description based on the action and the inferred object, using a generative model, wherein the description indicates a storyline with contextually aligned sequence for the follow-up media, determine a surprise element based on at least one of, a user data, theme adjustments, and contextual understanding of the ongoing media, wherein the surprise element indicates a contextually relevant thematic change to enhance user engagement, and generate the follow-up media based on merging the surprise element, parsing the description, and animating the parsed description.
[0152] In an embodiment, to generate the transformation effect, the at least one processor (106) is configured to capture the motion information of the one or more objects in the outgoing frame of the ongoing media and the incoming media of the follow-up media, wherein the motion information indicates movement and dynamics of the one or more objects between the ongoing media and the follow-up media, compute spatial-temporal difference between the outgoing frame of the ongoing media and the incoming frame of the follow-up media based on the motion information, and generate the transformation effect based on the spatial-temporal difference such that the transformation effect seamlessly blends the content of the ongoing media and the follow-up media thereby ensuring visual continuity.
[0153] The present disclosure enables the user to tailor the content based on their preferences, thereby generating dynamic and contextually relevant content that attracts audiences and increases the overall viewing experience.
[0154] In a nutshell, the present disclosure may provide an efficient and cost-effective way of generating the one or more sequences of the media such as the short-form media.
[0155] The embodiment disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements. The elements can be at least one of a hardware device or a combination of hardware devices and software modules.
[0156] It is understood that terms including "unit" or "module" at the end may refer to the unit for processing at least one function or operation and may be implemented in hardware, software, or a combination of hardware and software.
[0157] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.
[0158] The drawings and the forgoing description give example of embodiment. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from an embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.
[0159] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.
[0160] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.
[0161] The foregoing description of an embodiment will fully reveal the general nature of the embodiment herein that others can, by applying current knowledge, readily modify and adapt such an embodiment for various applications without departing from the generic concept, and, therefore, such adaptations and modifications should be, and are intended to be, comprehended within the meaning and range of equivalents of the disclosed embodiment. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiment herein has been described in terms of at least one embodiment, those skilled in the art will recognize that the embodiment herein can be practiced with modification within the spirit and scope of the embodiment as described herein.
[0162]
Claims
1.A method (300) performed by an electronic device (100) for generating one or more sequences of media, the method (300) comprising:obtaining (302) ongoing media currently played on the electronic device (100);selecting (304) at least one transition object as a region of interest (ROI) in the ongoing media based on at least one of a user input or an object detection;generating (306) follow-up media upon determining that the follow-up media is unavailable in a Database (DB) (104), based on a description of the at least one selected transition object, wherein the follow-up media indicates a content sequence aligned with spatial-temporal attributes of the at least one selected transition object;generating (308) a transformation effect based on motion information, wherein the transformation effect indicates a transition of one or more objects present in an outgoing frame of the ongoing media to one or more objects present in an incoming frame of the follow-up media; anddisplaying (310) the follow-up media with the transformation effect thereby generating the one or more sequences of the media with seamless crossfading.2.The method (300) as claimed in claim 1, wherein the selecting (304) of the at least one transition object based on the user-input comprises:obtaining (304a) the user input indicating an identification of the at least one transition object as the ROI;determining (304b) the availability of the follow-up media; andretrieving (304c) the follow-up media other than the ongoing media from the DB (104), based on the user input.3.The method (300) as claimed in claim 1, wherein the selecting (304) of the at least one transition object based on the object detection comprises:generating (304-1) an object-level description based on detecting the one or more objects in each of one or more frames of the ongoing media, wherein the object-level description indicates a caption sequence describing the one or more objects;generating (304-2) a spatial-temporal description based on the object-level description and the spatial-temporal attributes of the one or more objects;computing (304-3) object-level score for each of the one or more objects based on correlating weighted factors associated with the spatial-temporal description and the spatial-temporal attributes; andselecting (304-4) the at least one transition object as the ROI among the one or more objects in the ongoing media, wherein the object-level score for the at least one selected transition object exceeds a predefined video transition threshold.4.The method (300) as claimed in claim 3,wherein the spatial-temporal attributes comprise a size of an object, an object appearance, a surrounding context, motion dynamics, and a semantic relevance; andwherein the spatial-temporal description includes one or more of a dynamic behavior, interactions, and positional relationships of the detected one or more objects over time, incorporating motion patterns, attribute changes, and contextual spatial relationships.5.The method (300) as claimed in any one of claims 1 to 4, wherein the generating (306) of the follow-up media comprises:generating (306a) an action and an inferred object based on the at least one selected transition object, wherein the action indicates extrapolated interaction behavior, the inferred object indicates a semantically consistent replacement of the at least one transition object;generating (306b) the description based on the action and the inferred object, using a generative model, wherein the description indicates a storyline with contextually aligned sequence for the follow-up media;determining (306c) a surprise element based on at least one of user data, theme adjustments, and contextual understanding of the ongoing media, wherein the surprise element indicates a contextually relevant thematic change to enhance user engagement; andgenerating (306d) the follow-up media based on merging the surprise element, parsing the description, and animating the parsed description.6.The method (300) as claimed in claim 5, wherein the action is summarized in a text description of a predicted activity for the inferred object.7.The method (300) as claimed in any one of claims 1 to 6, wherein generating (308) of the transformation effect comprises:capturing (308a) motion information of the one or more objects in the outgoing frame of the ongoing media and the incoming frame of the follow-up media, wherein the motion information indicates a movement and dynamics of the one or more objects between the ongoing media and the follow-up media;computing (308b) a spatial-temporal difference between the outgoing frame of the ongoing media and the incoming frame of the follow-up media based on the motion information; andgenerating (308c) the transformation effect based on the spatial-temporal difference such that the transformation effect seamlessly blends the content of the ongoing media and the follow-up media thereby ensuring visual continuity.8.An electronic device (100) for generating one or more sequences of media, the electronic device (100) comprising:a memory (102) configured to store at least one instruction of a computer program;at least one processor (106) configured to execute the at least one instruction to communicate with the memory (102),the at least one processor (106) is configured to:obtain ongoing media currently played on the electronic device (100);select at least one transition object as a region of interest (ROI) in the ongoing media based on at least one of a user input or an object detection;generate a follow-up media upon determining that the follow-up media is unavailable in a Database (DB) (104), based on a description of the at least one selected transition object, wherein the follow-up media indicates a content sequence aligned with spatial-temporal attributes of the at least one selected transition object;generate a transformation effect based on motion information, wherein the transformation effect indicates a transition of one or more objects present in an outgoing frame of the ongoing media to one or more objects present in an incoming frame of the follow-up media; anddisplay the follow-up media with the transformation effect thereby generating the one or more sequences of media with seamless crossfading.9.The electronic device (100) as claimed in claim 8, wherein to select the at least one transition object based on the user-input, the at least one processor (106) is configured to:obtain the user input indicating an identification of the at least one transition object as the ROI;determine the availability of the follow-up media; andretrieve the follow-up media other than the ongoing media from the DB (104), based on the user input.10.The electronic device (100) as claimed in claim 8, wherein to select the at least one transition object based on the object detection, the at least one processor (106) is configured to:generate an object-level description based on detecting the one or more objects in each of one or more frames of the ongoing media, wherein the object-level description indicates a caption sequence describing the one or more objects;generate a spatial-temporal description based on the object-level description and the spatial-temporal attributes of the one or more objects;compute an object-level score for each of the one or more objects based on correlating weighted factors associated with the spatial-temporal description and the spatial-temporal attributes; andselect the at least one transition object as the ROI among the one or more objects in the ongoing media, wherein the object-level score for the at least one selected transition object exceeds a predefined video transition threshold.11.The electronic device (100) as claimed in claim 10, wherein the spatial-temporal attributes comprise a size of an object, an object appearance, a surrounding context, motion dynamics, and a semantic relevance; andwherein the spatial-temporal description comprises one or more of dynamic behavior, interactions, and positional relationships of the detected one or more objects over time, incorporating motion patterns, attribute changes, and contextual spatial relationships.12.The electronic device (100) as claimed in any one of claims 8 to 11, wherein to generate the follow-up media, the at least one processor (106) is configured to:generate an action and an inferred object based on the at least one selected transition object, wherein the action indicates extrapolated interaction behavior, the inferred object indicates a semantically consistent replacement of the at least one transition object;generate the description based on the action and the inferred object, using a generative model, wherein the description indicates a storyline with contextually aligned sequence for the follow-up media;determine a surprise element based on at least one of, a user data, theme adjustments, and contextual understanding of the ongoing media, wherein the surprise element indicates a contextually relevant thematic change to enhance user engagement; andgenerate the follow-up media based on merging the surprise element, parsing the description, and animating the parsed description.13.The electronic device (100) as claimed in claim 12, wherein the action is summarized in a text description of a predicted activity for the inferred object.14.The electronic device (100) as claimed in any one of claims 8 to 13, wherein to generate the transformation effect, the at least one processor (106) is configured to:capture the motion information of the one or more objects in the outgoing frame of the ongoing media and the incoming media of the follow-up media, wherein the motion information indicates movement and dynamics of the one or more objects between the ongoing media and the follow-up media;compute a spatial-temporal difference between the outgoing frame of the ongoing media and the incoming frame of the follow-up media based on the motion information; andgenerate the transformation effect based on the spatial-temporal difference such that the transformation effect seamlessly blends the content of the ongoing media and the follow-up media thereby ensuring visual continuity.15.A computer readable storage medium having stored thereon a computer program that, when executed by at least one processor, performs the method of any one of claims 1-7.