System and method for generation of content-aware element interaction motion frames for electronic device display

The system generates content-aware element interaction motion frames for electronic device displays, addressing limitations in existing dynamic solutions by providing personalized, immersive, and connected experiences with reduced battery consumption.

WO2025104606A1PCT designated stage expired Publication Date: 2025-05-22SAMSUNG ELECTRONICS CO LTD

Patent Information

Application Number
PCT/IB2024/061267
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-11-13
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing solutions for dynamic display screens on electronic devices, such as live wallpapers and video wallpapers, provide limited continuity and immersive experience due to preprogrammed interactions and lack of personalization, while also being battery-intensive and limited to specific wallpaper options.

Method used

A system and method for generating content-aware element interaction motion frames that involve receiving an input image, extracting physics-informed motion features, determining spatial motion relations, estimating layers, combining objects into groups, and generating motion vectors to create personalized, immersive, and connected experiences on electronic device displays.

Benefits of technology

The solution provides enhanced personalization, immersion, and continuity across different states of an electronic device's display, such as lock screen and home screen, while reducing battery consumption and offering more aesthetic and interactive visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061267_22052025_PF_FP_ABST
    Figure IB2024061267_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein is a method for generating a content-aware element interaction motion frames for a state on a display of an electronic device. The method comprises receiving an input image having a plurality of objects and extracting physics-informed motion features of each object to obtain motion representation. The method comprises determining a spatial motion relation of each object based on the motion representation and a placement relation of each object with each of the plurality of objects. The method comprises estimating one or more layers and corresponding information of each object and combining one or more objects into a plurality of groups, using the information of one or more layers and the placement relation, suitable for the state. The method comprises generating one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation.
Need to check novelty before this filing date? Find Prior Art

Description

DESCRIPTIONTitle of InventionSYSTEM AND METHOD FOR GENERATION OF CONTENT-AWARE ELEMENT INTERACTION MOTION FRAMES FOR ELECTRONIC DEVICE DISPLAYTechnical Field

[0001] The present disclosure relates to image processing for display of electronic devices and more particularly, relates to a system and method for generation of content-aware element interaction motion frames for display on electronic device.Background Art

[0002] Nowadays, all smart devices have a provision for a user to personalize a display screen of the smart device. The personalized display may be a picture of the user or something that may feel relevant to the user. Additionally, in recent developments, animated display images are widely used and allow the smart device’s display interface to incorporate dynamic / moving elements. The dynamic elements are another way of incorporating personalizations into the smart devices to make the display more aesthetically pleasing and engaging. The dynamic elements may include animated wallpaper, interactive widgets, and animated icons.

[0003] Recently, transforming static images into dynamic, continuous experiences is gaining significant interest due to increase in use of gaming, digital life, virtual reality which has given the user an experience of immersive transitions and interactions.

[0004] The existing solutions such as live wallpapers and video wallpapers provide limited continuity and immersive experience since they provide limited, preprogrammed interactions only for pre-loaded wallpaper options for the smart devices. In addition, these solutions are usually paid applications and are battery intensive.

[0005] Figure 1 A-1 B illustrates an exemplary scenario 100 of wallpaper sets across various screens of an electronic device, in accordance with a prior art. Conventionally, effects in an always-on-display (AOD) offer a way of stylization for the AOD with a standard background blend transition. Further, pre-loaded wallpapers lack personalization but offer immersion and basic touch interactions.Moreover, the pre-loaded wallpapers and themes do not offer a complete aesthetic set of visuals and end-to-end effects. Furthermore, a lock screen interactivity is limited to swipe-unlock movements, does not incorporate stimulus-based interaction, and only has exclusive to limited, pre-programmed abstract options. Further, continuity from the user selected AOD effects is only limited to the lock screen and does not include continuity effects between the lock screen, and home screen for visuals chosen by the user.

[0006] Therefore, in view of the above-mentioned problems, it is advantageous to provide an improved system and method that can overcome the above-mentioned problems and limitations associated with animated, immersive, and interactive display screens.Technical Problem

[0007] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.

[0008] According to an embodiment of the present disclosure, a method for generating a content-aware element interaction motion frames for a state on a display of an electronic device is disclosed. The method comprises receiving an input image having a plurality of objects. Further, the method comprises extracting physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation. Furthermore, the method comprises determining a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects. Furthermore, the method comprises estimating one or more layers and corresponding information of each object of the plurality of objects in the input image. Still further, the method comprises combining one or more objects of the plurality of objects into a plurality of groups, using the information of one or more layers and the placement relation of the one or more objects, suitable for the state. Furthermore, the method comprises generating one or more motion vectors for each of the plurality of groups using the spatial motion relation and theplacement relation to obtain the content-aware element interaction motion frames for the state.

[0009] According to another embodiment, a system for generating a content-aware element interaction motion frames for a state on a display of an electronic device is disclosed. The system comprises one or more processors and a memory coupled with the one or more processors. The one or more processors are configured to receive an input image having a plurality of objects. Further, the one or more processors are configured to extract physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation. Further, the one or more processors are configured to determine a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects. Furthermore, the one or more processors are configured to estimate one or more layers and corresponding information of each object of the plurality of objects in the input image. Furthermore, the one or more processors are configured to combine one or more objects of the plurality of objects into a plurality of groups, using the information of one or more layers and the placement relation of the one or more objects, suitable for the state. Further, the one or more processors are configured to generate one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content- aware element interaction motion frames for the state.

[0010] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.Description of Drawings

[0011] The foregoing and other features of embodiments will become more apparent from the following detailed description of embodiments when read in conjunction withthe accompanying drawings. In the drawings, like reference numerals refer to like elements.

[0012] Figure 1A -1B illustrate an exemplary scenario of wallpaper sets across various screens of a electronic device, in accordance with a prior art.

[0013] Figure 2 illustrates a pictorial diagram depicting an exemplary environment for generating content-aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure. Figure 3 illustrates a block diagram of a system architecture for generating content-aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure.

[0014] Figure 4A-4D illustrate a schematic block diagram of working of the system and its modules for generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0015] Figure 5A-5C illustrate a pictorial depiction of a phase-wise working of modules for generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0016] Figures 6A-6C and Figures 7A-7B illustrate a pictorial depiction of first phase of generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0017] Figures 8 and Figures 9A-9F illustrate a pictorial depiction of second phase of generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0018] Figures 10A-10D, 11 and Figures 12A-12B illustrate a pictorial depiction of third phase of generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0019] Figure 13 illustrates a pictorial depiction of fourth phase of generating content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0020] Figure 14 illustrates a flow chart showing a method for generating a content- aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure.

[0021] Figures 15A-15B illustrate an exemplary scenario of personalized immersive screens of electronic devices, in accordance with an embodiment of the present disclosure.

[0022] Figures 16A-16C and Figure 17A-17B illustrate an exemplary scenario of personalized immersive screens on electronic devices based on different state and various factors, in accordance with an embodiment of the present disclosure.

[0023] Figure 18A-18C illustrates a pictorial depiction of category immersive screen display based on different electronic devices, in accordance with an embodiment of the present disclosure.Best Mode

[0024] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.

[0025] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.

[0026] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as “one or more features” or “one or more elements” or “at least one feature” or “at least one element.” Furthermore, the use of the terms“one or more” or “at least one” feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, “there needs to be one or more...” or “one or more elements is required.”

[0027] Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.

[0028] Use of the phrases and / or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “further embodiment”, “furthermore embodiment”, “additional embodiment” or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, or in the context of more than one embodiment, or in the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

[0029] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.

[0030] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include othersteps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.

[0031] Hereinafter, it is understood that terms including “unit” or “module” at the end may refer to the unit for processing at least one function or operation and may be implemented in hardware, software, or a combination of hardware and software.

[0032] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.

[0033] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term “or” as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.

[0034] As is traditional in the field, embodiments may be described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which may be referred to herein as units or modules or the like, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionallybe driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the invention. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the invention.

[0035] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.

[0036] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit “1” are shown at least in Figure 1 . Similarly, reference numerals starting with digit “2” are shown at least in Figure 2.

[0037] An object of the present disclosure is to provide an improved technique to overcome the above-described limitations associated with existing interactive or dynamic display screen and ease users into the idea of active wallpaper experience, rather than the current passive experience.

[0038] Another object of the present disclosure is to generate physics-aware transitions and animations from images provided by users.

[0039] Another object of the present disclosure is to provide personalized, immersive, and connected experience to the users when using the electronic device.

[0040] Yet another object of the present disclosure is to generate physics-aware animations with an understanding of the content of the image.

[0041] Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.

[0042] Figure 2 illustrates a pictorial diagram depicting an exemplary environment for generating content-aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure. As shown in the figure, a user 202 provides an image 206 as an input using electronic device 204 to a system 210. The input image 206 is sent via a network interface 208 to the system 210. The system 210 generates a content-aware element interaction motion frames for electronic device display, based on the input image 206 received from the user 202.

[0043] The system 210 may include a software, a hardware, a combination of software or hardware, an in-built application on the electronic device or an application to be installed and operated on the electronic device in communication with a network interface 208. The system 210 may also be available via cloud-based server and available remotely from the electronic device.

[0044] The network interface 208 may be configured to provide network connectivity and enable communication with paired devices such as the system 210. The network connectivity may be provided via a wireless connection or a wired connection. For example, the network connectivity may be provided via cellular technology, such as 3rd Generation (3G), 4th Generation (4G), 5th Generation (5G), pre-5G, 6th Generation (6G), or any other wireless communication technology such as Bluetooth.

[0045] Figure 3 illustrates a block diagram of a system architecture for generating a content-aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure.

[0046] The system 210 generates a content-aware element interaction motion frames for a display on the electronic device 204 based on the input image 206 received from the user 202 via network interface 208.

[0047] The system 210 may include a processor 302 which is communicatively coupled to a memory 304, one or more modules 306, and a data unit 308.

[0048] In an example, the processor 302 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 302 may be configured to fetch and execute computer-readable instructions and data stored in the memory 304. At this time, the processor 302 may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, and an Al-dedicated processor such as a neural processing unit (NPU). The processor 302 may control the processing of input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory, i.e. , the memory 304. The predefined operating rule or artificial intelligence model is provided through training or learning. Further, the processor 302 may be operatively coupled to each of the memory, the I / O Interface. The processor 302 may be configured to process, execute, or perform a plurality of operations described herein.

[0049] In an example, the memory 304 may include any non-transitory computer- readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. The memory 304 is communicatively coupled with the processor 302 to store processing instructions for completing the process. Further, the memory 304 may include an operating system for performing one or more tasks of the system, as performed by a generic operating system in a computing domain. The memory 304 is operable to store instructions executable by the processor 302.

[0050] In some embodiments, the one or more modules 306 may include a set of instructions that can be executed to cause the system 210 to perform any one or more of the methods disclosed. The system 210 may operate as a standalone device or may be connected, e.g., using a network, to other computer systems orperipheral devices. Further, while a single system 210 is illustrated, the term “system” shall also be taken to include any collection of systems or sub-systems that individually or jointly execute a set, or multiple sets, of instructions to perform one or more computer functions.

[0051] In an embodiment, the module(s) 306 may be implemented using one or more artificial intelligence (Al) modules that may include a plurality of neural network layers. Examples of neural networks include but are not limited to, Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), and Restricted Boltzmann Machine (RBM). Further, ‘learning’ may be referred to in the disclosure as a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to supervised learning, unsupervised learning, semisupervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the present subject matter’s mechanism through an Al model. A function associated with an Al module may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor, such as a neural processing unit (NPU). One or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0052] The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor such as a neural processing unit (NPU).

[0053] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0054] Here, being provided through learning means that, by applying a learning technique to a plurality of learning data, a predefined operating rule or Al model of a desired characteristic is made. The learning may be performed in a device itself in which Al according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0055] The Al model may consist of a plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0056] The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semisupervised learning, or reinforcement learning.

[0057] In some embodiments, the data unit 308 serves, amongst other things, as a repository for storing data processed, received, and generated by one or more of the modules 306.

[0058] According to embodiments of the present disclosure, the system 210 may include one or more modules 306, such as an input module 310, a feature extraction module 312, an element relationship module 314, a layerization module 316, and a content aware motion generation module 318. The input module 310, the feature extraction module 312, the element relationship module 314, the layerization module 316, and the content aware motion generation module 318 are communicably coupled with each other.

[0059] In an embodiment, the input module 310 may be configured to receive an input image having a plurality of objects.

[0060] In an embodiment, the feature extraction module 312 may be configured to extract physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation.

[0061] In an embodiment, to extract the physics-informed motion features, the feature extraction module 312 may be configured to obtain a spatial latent representation, a semantic map representation, and a depth map representation of the input image. The feature extraction module 312 may be configured to concatenate the spatial latent representation with the semantic map representation to obtain a spatial representation of elements of the plurality of objects in the input image. The feature extraction module 312 may be configured to multiply the spatial latent representation with the depth map representation and concatenating the multiplied representation with the semantic map representation to obtain a depthwise representation of the elements of the objects in the input image. The feature extraction module 312 may be configured to concatenate the spatial representation of the elements of the plurality of objects with the depth-wise representation of elements to obtain the motion representation features of elements of the plurality of objects in the input image by predicting physical properties of each of the plurality of objects.

[0062] In an embodiment, the element relationship module 314 may be configured to determine a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects. The spatial motion relation corresponds to the one or more elements present in the input image and a relation between the one or elements with respect to the motion representation and the placement relation.

[0063] In an embodiment, to determine a spatial motion relation of each object with each of the plurality of objects, the element relationship module 314 may be configured to segregate one or more elements of the plurality of objects from the motion representation features of elements. The element relationship module 314 may be configured to perform a cross-attention operation on the one or moreelements of the plurality of objects based on the spatial representation of elements, the depth-wise representation of elements, and the motion representation of elements. The element relationship module 314 may be configured to determine spatial motion relationship among the spatial representation, the depth-wise representation, and motion relationship among the spatial representation, depth-wise representation and the motion representation of the elements based on the crossattention operation on the one or elements of the plurality of objects to obtain an element relation graph and a plurality of bounding boxes of the elements of the plurality of objects.

[0064] In an embodiment, to determine the placement relation and estimating the one or more layers of each object with each of the plurality of objects, the element relationship module 314 may be configured to utilize the element relation graph, the plurality of bounding boxes of the elements of the plurality of objects and the segregated one or more elements of the plurality of objects to construct an edge distance-wise layers. The element relationship module 314 may be configured to identify the one or more layers by classifying each layer obtained from the edge distance-wise layers.

[0065] In an embodiment, the layerization module 316 may be configured to estimate one or more layers and corresponding information of each object of the plurality of objects in the input image. The layerization module 316 may be further configured to combine one or more objects of the plurality of objects into a plurality of groups, using the information of one or more layers and the placement relation of the one or more objects, suitable for the state.

[0066] In an embodiment, to identify the one or more layers for the state, the layerization module 316 may be configured to obtain one or more spatial features of the one or more layers using a predefined pre-training model. The layerization module 316 may be configured to determine the one or more elements in the one or more layers suitable for interactions and motion based on the one or more spatial features. The layerization module 316 may be configured to determine one or more components of a layout on the display of the electronic device using the pre-trained model. The layerization module 316 may be configured to determine if any of the oneor more components of the layout is obstructing the one or more elements for the motion to obtain one or more suitable components of the layout. The layerization module 316 may be configured to identifying the one or more layers for the state by determining the motion of the one or more elements suitable on the display of the electronic device based on the one or more suitable elements and the one or more suitable components of the layout.

[0067] In an embodiment, the content aware motion generation module 318 may be configured to generate one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content- aware element interaction motion frames for the state.

[0068] In an embodiment, to generate one or more motion vectors, the content aware motion generation module 318 may be configured to obtain a state specific element vector based on the identified one or more layers information for the state. The content aware motion generation module 318 may be configured to obtain motion features of the one or more suitable elements for the state from the segregated one or more elements of the plurality of objects. The content aware motion generation module 318 may be configured to obtain the textual features of the one or more suitable elements from the element relation graph. The content aware motion generation module 318 may be configured to obtain the state properties and contextual factors of the electronic device. The content aware motion generation module 318 may be configured to generate the motion frames using the motion features of the one or more suitable elements across temporal dimensions. The content aware motion generation module 318 may be configured to generate the interactive motion frames based on the generated motion frames and obtained the element relation graph. The content aware motion generation module 318 may be configured to generate the one or more motion vectors for the state based on the state specific element vector, the motion frames, the interactive motion frames, the textual features, the state properties and the contextual factors of the electronic device. The state properties include one or more of colour range, colour contrast, colour palette, average brightness, shadows, highlights, flat colour, true black, luminance, gradients and the contextual factors of the electronic device include oneor more of battery state, charging state, time of day, current weather, touch interactions, gyroscopic motion.

[0069] Figure 4A-4D illustrates a schematic block diagram of working of the system 210 and its modules 306 for generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0070] Initially, at operation 401 , the input module 310 is configured to receive an input image having a plurality of objects.

[0071] At operation 402, the feature extraction module 312 is configured to extract physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation.

[0072] At operation 402a, to extract the physics-informed motion features, the feature extraction module 312 is configured to obtain a spatial latent representation, a semantic map representation, and a depth map representation of the input image.

[0073] At operation 402b, the feature extraction module 312 is configured to concatenate the spatial latent representation with the semantic map representation to obtain a spatial representation of elements of the plurality of objects in the input image.

[0074] At operation 402c, the feature extraction module 312 is configured to multiply the spatial latent representation with the depth map representation and concatenating the multiplied representation with the semantic map representation to obtain a depth-wise representation of the elements of the objects in the input image.

[0075] At operation 402d, the feature extraction module 312 is configured to concatenate the spatial representation of the elements of the plurality of objects with the depth-wise representation of elements to obtain the motion representation features of elements of the plurality of objects in the input image by predicting physical properties of each of the plurality of objects.

[0076] At operation 403, the element relationship module 314 is configured to determine a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object witheach of the plurality of objects. The spatial motion relation corresponds to the one or more elements present in the input image and a relation between the one or elements with respect to the motion representation and the placement relation.

[0077] At operation 403-1 a, to determine a spatial motion relation of each object with each of the plurality of objects, the element relationship module 314 is configured to segregate one or more elements of the plurality of objects from the motion representation features of elements.

[0078] At operation 403-1 b, the element relationship module 314 is configured to perform a cross-attention operation on the one or more elements of the plurality of objects based on the spatial representation of elements, the depth-wise representation of elements, and the motion representation of elements.

[0079] At operation 403-1 c, the element relationship module 314 is configured to determine spatial motion relationship among the spatial representation, the depthwise representation, and motion relationship among the spatial representation, depth-wise representation and the motion representation of the elements based on the cross-attention operation on the one or elements of the plurality of objects to obtain an element relation graph and a plurality of bounding boxes of the elements of the plurality of objects.

[0080] At operation 403-2a, to determine the placement relation and estimating the one or more layers of each object with each of the plurality of objects, the element relationship module 314 is configured to utilize the element relation graph, the plurality of bounding boxes of the elements of the plurality of objects and the segregated one or more elements of the plurality of objects to construct an edge distance-wise layers.

[0081] At operation 403-2b, the element relationship module 314 is configured to identify the one or more layers by classifying each layer obtained from the edge distance-wise layers.

[0082] At operation 404, the layerization module 316 is configured to estimate one or more layers and corresponding information of each object of the plurality of objects in the input image.

[0083] At operation 405, The layerization module 316 is configured to combine one or more objects of the plurality of objects into a plurality of groups, using the information of one or more layers and the placement relation of the one or more objects, suitable for the state.

[0084] At operation 405a, to identify the one or more layers for the state, the layerization module 316 is configured to obtain one or more spatial features of the one or more layers using a predefined pre-training model.

[0085] At operation 405b, the layerization module 316 is configured to determine the one or more elements in the one or more layers suitable for interactions and motion based on the one or more spatial features.

[0086] At operation 405c, the layerization module 316 is configured to determine one or more components of a layout on the display of the electronic device using the pre-trained model.

[0087] At operation 405d, the layerization module 316 is configured to determine if any of the one or more components of the layout is obstructing the one or more elements for the motion to obtain one or more suitable components of the layout.

[0088] At operation 405e, the layerization module 316 is configured to identify the one or more layers for the state by determining the motion of the one or more elements suitable on the display of the electronic device based on the one or more suitable elements and the one or more suitable components of the layout.

[0089] At operation 406, the content aware motion generation module 318 is configured to generate one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content- aware element interaction motion frames for the state.

[0090] At operation 406a, to generate one or more motion vectors, the content aware motion generation module 318 is configured to obtain a state specific element vector based on the identified one or more layers information for the state.

[0091] At operation 406b, the content aware motion generation module 318 is configured to obtain motion features of the one or more suitable elements for the state from the segregated one or more elements of the plurality of objects.

[0092] At operation 406c, the content aware motion generation module 318 is configured to obtain the textual features of the one or more suitable elements from the element relation graph.

[0093] At operation 406d, the content aware motion generation module 318 is configured to obtain the state properties and contextual factors of the electronic device.

[0094] At operation 406e, the content aware motion generation module 318 is configured to generate the motion frames using the motion features of the one or more suitable elements across temporal dimensions.

[0095] At operation 406f, the content aware motion generation module 318 is configured to generate the interactive motion frames based on the generated motion frames and obtained the element relation graph.

[0096] At operation 406g, the content aware motion generation module 318 is configured to generate the one or more motion vectors for the state based on the state specific element vector, the motion frames, the interactive motion frames, the textual features, the state properties and the contextual factors of the electronic device. The state properties include one or more of colour range, colour contrast, colour palette, average brightness, shadows, highlights, flat colour, true black, luminance, gradients and the contextual factors of the electronic device include one or more of battery state, charging state, time of day, current weather, touch interactions, gyroscopic motion.

[0097] Figure 5A-5C illustrate a pictorial depiction of a phase-wise working of modules of the system 210 for generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0098] As shown in Figure 5A-5C, the system 210 receives the input image 206 via the input module 310 on which deep feature synthesis is performed via the feature extraction module 312. The module 312 extracts latent representation from image i.e. , spatial latent, semantic map, depth map and generates spatial and depth representation of elements with the latent representation using transformer encoder-decoder network. The module 312 then learns the elements characteristic motion features through a downstream task of predicting physics features such as friction, density, gravity, etc.

[0099] The output from the feature extraction module 312 is then processed by the element relationship module 314 for spatial motion relation understanding. The module 314 segregates the subject and object motion features from characteristic motion information obtained from deep feature synthesis and then performs cross attention in transformers on spatial, depth and motion representations with both subject and object element motion feature. The module 314 then utilizes the refined features to predict the identity of each element and relation among them to construct an elements relation graph.

[0100] Thereafter, the layerization module 316 performs relation aware elements taxonomization. This module 316 uses elements instance segmentation and the relation graph to merge multiple related elements to construct edge distance-wise layers (1 ,2..N distance layers). The module 316 then identifies the layers suitable for particular states (Ex: AOD / Lock Screen (LS) / Home Screen (HS) etc.) by classifying each layer obtained from elements layerization (each layer comprises of one or more element).

[0101] Lastly, the content aware motion generation module 318 generates context aware motion. The module 318 generates motion video for the elements which are identified suitable for each state using diffusion model (e.g. 3D ll-Net based video diffusion model) using following conditional inputs: elements motion feature encoding, state specific elements guidance, spatial-motion relations obtained from relation graph and contextual effects guidance.

[0102] The details of the functioning of the modules are discussed in the subsequent paragraphs with reference to drawings.

[0103] Figures 6A-6C and Figures 7A-7B illustrate a pictorial depiction of the first phase of generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0104] To specify, Figure 6A-6C illustrate deep feature synthesis performed via the feature extraction module 312. The objective of the module is to generate element representation corresponding to the spatial, depth and learn the characteristic motion representations. The module 312 receives latent representation of the image such as Spatial latent, Semantic map and depth map, as input.

[0105] As seen from the Figure 6A-6C, the module 312 comprises a sub-module a Spatial Feature Extractor which extracts elements spatial representation, a Depth Feature Extractor which extracts elements depth feature representation and Characteristic Motion Understanding module to learn the elements characteristic motion features.

[0106] The sub-module concatenates spatial features with the semantic map and pass it to the transformer encoder. The transformer then provides the weighted triplets (element-predicate-element) to the decoder as input (only during the transformer training phase) as guidance. Thereafter, the transformer decodes the context feature of the elements in the scene to provide elements spatial representation.

[0107] The Depth Feature Extractor multiplies the latent spatial feature with the depth map and then concatenates it with the sematic map before passing it to the transformer encoder. The transformer encoder provides the weighted triplets (element1-depth_predicate-element2) to the decoder as input (only during the transformer training phase) as guidance. The transformer then decodes the context feature of the elements in the scene to provide depth-wise elements representation.

[0108] The Characteristic Motion Understanding module concatenates spatial, and depth wise element representation obtained from the previous steps with the sematic map as input to the transformer encoder. The transformer encoder provides the triplets (element1-motion_predicate-element2) to the decoder as input (only during the transformer training phase) as guidance. To learn the characteristic motion representation of the entities, the latent is provided as input to the downstream task to predict physics features such as friction, density, gravity, etc. This provides end- to-end training thereby ensuring that the latent features embed the characteristics motion of the elements through prediction of physics features.

[0109] Thus, the module 312 generates elements feature representation (spatial, depth, motion) of dimension N x d where, N represents number of possible elements combinations and d is feature dimension, as output.

[0110] Figure 7A-7B illustrates an exemplary scenario for the Characteristic Motion Understanding module to learn the elements characteristic motion features, according to an embodiment of the present invention.

[0111] The Characteristic Motion Understanding module receives concatenated spatial and depth wise elements representation along with the sematic map, as input features. These input features are passed to the transformer encoder where first self-attention is applied to obtain the contextual embedding for the elements. Then corresponding triplet embedding is passed to the decoder (during training only). These contextual features and the decoded triplet embeddings are given as input to the cross-attention layers where Q, K are considered as the contextual embedding and V is considered as the self-attention on triplet embedding.

[0112] The transformer decoder outputs 1 x d dimension feature vector which is then used to estimate physical properties of the elements it represents as a downstream task. Then N such decoder blocks are used giving N x d dimension output holding all N elements characteristics motion representation.

[0113] Further, during inference, empty context token (<start>) is passed as triplet and extract only the decoder outputs which are characteristic motion features. The Characteristic Motion Understanding module generates element wise motion feature representation, as output.

[0114] Figure 8 and Figures 9A-9F illustrate a pictorial depiction of the second phase of generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0115] To specify, Figure 8 depicts the spatial motion relation understanding by the element relationship module 314. The objective of the module 314 is to identify all the elements present in the input image (subject-object separation) and relation between them concerning motion and placement predicates. The objective of themodule 314 is to identify the bounding boxes corresponding to the elements and construct a graph of the elements (as nodes) and predicates (as edges) using all predicted information.

[0116] The module 314 receives elements feature representation (spatial, depth, motion) from deep feature synthesis module 312, as input.

[0117] The module 314 comprises sub-modules: Motion Segregation module which segregates the subject and object motion features from characteristic motion information obtained from deep feature synthesis; Multimodal Cross Attention module to perform cross attention in transformers on spatial, depth and motion representations with both subject and object element motion feature; and Elements Relation Predictor which utilizes the refined features to predict the identity of each element and relation among them to construct an elements relation graph.

[0118] The output of the module 314 is a relation graph and element bounding boxes for the input image, segregated motion features for subject and object elements.

[0119] Figures 9A-9F illustrate detailed working of the sub-modules as discussed in Figure 8. At first, elements feature representation (spatial, depth, motion) of dimension N x d where, N represents number of possible elements combinations (triplets) and d is feature dimension, is received as input.

[0120] The Motion Segregation module segregates the subject and object motion features from characteristic motion information QM that encodes <subject-predicate- object>. Then subject positional embedding (Eps) and object positional embedding (Epo) with characteristic motion features (QM) are added independently. Thereafter, both embeddings are appended together and apply coupled self-attention to contextually separate subject and object from the characteristic motion information to obtain dissociated subject (EMs) and object (EMo) features.

[0121] The Multimodal Cross Attention module performs cross attention in transformers on spatial (QS), depth (QDp) and motion (QM) representations with both subject and object element motion feature (EMs, EMo). The module enables to attend different modalities corresponding to specific element motion features.

[0122] Where, query (Q) is element motion feature of subject / object (EMs I EMo), Key (K) and Value (V) are elements feature representation with respect to spatial (QS), depth (QDp) and motion (QM).

[0123] The generated spatial, depth and motion attention maps for subject and object elements hold relevant context to deduce the identity (e.g.,: fish, lake) and relation (placement and motion predicate).

[0124] The Elements Relation Predictor utilizes the refined features to predict the identity of each element and relation among them to construct an elements relation graph. In this module, each modality attention feature maps are down sampled using Conv2D independently for subject and object.

[0125] For Elements Identification, individually down sampled features are passed to MLP (Multilayer Perceptron) and classify its identity.

[0126] For Bounding box prediction, individually down sampled features are passed to CNN (convolutional neural network) and predicts bounding box for the corresponding element (subject and object)

[0127] For Placement Predicate, to deduce the relationship between subject and object, both features are fused and passed to MLP (Multilayer Perceptron) to identify the spatial relationship.

[0128] For Motion Predicate, to predict the motion relation, motion attention features is needed with additional knowledge of subject and object for which the motion is predicted. Hence, all down sampled features (spatial + depth + motion) are combined to predict the motion predicate.

[0129] For training, the model is trained end-to-end with multi-task learning as an objective function. The elements (Subject and object) Identification MLPs are classifiers and trained with cross-entropy loss.

[0130] For placement predicate, fusion of subject and object feature (e.g.: fish + lake) is used as anchor and apply triplet loss with ground truth predicate (e.g.: ‘In’) as positive sample.

[0131] For motion predicate, fusion of subject and object feature (spatial + depth + motion) is used as anchor and apply triplet loss with ground truth predicate (e.g.: ‘swim’) as positive sample.

[0132] Lastly, the Bounding box prediction is trained as a downstream task with standard object detection loss. Thus, the output is the relation graph and element bounding boxes for the input image.

[0133] Figures 10A-10D and Figures 12A-12B illustrate a pictorial depiction of third phase of generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention.

[0134] Figure 10A-10D illustrates performing relation aware elements taxonomization by the layerization module 316. The objective of the model is contextual segregation and elements layerization using relation graph. Further, the objective is to classify the layers (each layer comprises of one or more element) into its suitable state.

[0135] The module 316 receives relation graph, elements bounding boxes of the image and elements motion feature from previous module 314 as input.

[0136] The module 316 comprises sub-modules such as Relation-aware Elements Layerization module for Elements Instance Segmentation and utilizing the relation graph to merge multiple related elements to construct edge distance-wise layers; Elements Classification module to identify the layers suitable for particular states (Ex: AOD / LS / HS etc.) by classifying each layer obtained from relation-aware elements layerization.

[0137] Figure 11 illustrates detailed explanation of contextual segregation and layerization of each element and combination of related elements as layers, as discussed in Figure 10. The relation graph and element bounding boxes of the image are received as input.

[0138] First of all, Deep Occlusion-Aware Instance Segmentation which is Elements Instance Segmentation is performed. This is performed by use of Masked-RCNNbased mode to perform the instance segmentation. Then each element’s bounding box is parsed as individual input and occlusion-aware instance segmentation.

[0139] Then Contextual Element Blending is performed using the relationship graph to combine multiple related elements and merge them to construct edge distance wise layers.

[0140] For example: Two Distance Layers : (1) Lake -> fish, (2) Lake-> leaves

[0141] Thus, the spatial combination of elements based on their relations in the form of layers (N-distance layers) is obtained as output.

[0142] Figure 12A-12B illustrate layered elements classification as discussed in Figure 10. The objective is to identify the layers suitable and corresponding information for particular states (Ex: AOD / LS / HS etc.) by classifying each layer obtained from elements layerization (each layer comprises of one or more element).

[0143] At first, multi-level screen and element information such as all the layers and their combinations (N-distance layers), screen Layouts corresponding to each state, motion features of each elements in corresponding layer are received as input.

[0144] Each layer (i) of total N layers is processed one-by-one and classified into one of the states (k). Finally, for each states the most confident (prob > 0.70) layers are selected.

[0145] Layer Feature Understanding module understands the spatial context of elements and their relative positions. Conv2D based pre-trained model is used for obtaining spatial features. These features determine whether the elements present in the layer are suitable for interactions and motions.

[0146] Screen Layout Understanding module understands the spatial context of widgets / app icons and layout of the screen for each state. Conv2D based model is used to understand the layout features. These features determine if any components of the layout (Ex. Widgets, icons) are obstructing the elements for motion.

[0147] Motion features module understands the motion strength of each elements present in the layer. The motion features extracted in the previous module are used here. These determine if the motion strength of each element is suitable for a particular state.

[0148] For example: For AOD state minimal motion is sufficient.

[0149] Thus, the output is the layer chosen for each state(k) to generate motion.

[0150] Figure 13 illustrates a pictorial depiction of fourth phase of generating a content-aware element interaction motion frames for a state on a display of an electronic device, according to an embodiment of the present invention. The objective is to generate motion video for the elements which are identified suitable for each state using 3D ll-Net based video diffusion model.

[0151] The input includes: a) All the segmented elements from the image; b) State specific elements guidance (cs): Elements which are classified suitable for that state; c) Motion feature encoding (cm): Motion features of each element which are classified suitable for that state; d) Spatial-motion relation conditioning (cr): Relation among the elements via motion and placement predicate; e) Contextual effects guidance (ce): Using specific state properties (like: colour range, contrast, brightness, shadows) in conjunction with contextual factors from user mobile device for guidance (like: battery, charging state, time of day, weather) for exhibiting varying visual and motion effects.

[0152] This module 318 utilizes latent diffusion models (LDMs) as backbone for elements motion generation. From the segmented elements from the image in a state, the forward diffusion procedure introduces noise to the encoded latent z, thus producing a noisy latent vector zt. The 3D U-Net, parameterized by 9, denoises the noisy latent representation by predicting the noise. This denoising process is controlled using motion and contextual conditions through the temporal and crossattention mechanism. The overall learning objective of the network is:

[0153] S is the VAE encoder that transforms the images from pixel space to discrete latent space. ztis the latent code at time step t. At inference time, a random noise ztis sampled from / V(0, 1) and iteratively denoised by the ll-Net.

[0154] State specific elements guidance (cs): State specific elements vector (1 : element suitable for that state) is provided to multilayer perceptron (MLP) and is spatially incorporated into network’s convolutional layers. This effectively controls that only classified elements for a particular state get included for motion generation.

[0155] Motion feature encoding (cm): Corresponding motion features of the classified elements (i.e. one distance layered elements) are passed across the temporal dimension, to control the motion strength and generate realistic motion for the elements in the generated frames.

[0156] Spatial-motion relation conditioning (cr): Spatial-motion relations obtained from relation graph are encoded using the pre-trained CLIP text encoder, which is then injected into the Unet via cross-attention layers. This helps in capturing sematic relation features amongst the elements, facilitating a fine-grained and content-aware, interactive animation generation.

[0157] Contextual effects guidance (ce): Both specific state properties (like: colour range, contrast, brightness) and contextual factors obtained from device are fed to MLPs and obtain the final conditional features after concatenation. These features are conditioned to spatial convolution layer to guide the visual appearances of the generated frames.

[0158] Thus, the output of the module 318 is generated frames for every state consisting of animations of the element.

[0159] Figure 14 illustrates a flow chart showing a method for generating a content- aware element interaction motion frames for a state on a display of an electronic device, in accordance with an embodiment of the present disclosure.

[0160] At step 1402, the method 1400 comprises receiving an input image having a plurality of objects.

[0161] At step 1404, the method 1400 comprises extracting physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation.

[0162] At step 1406, the method 1400 comprises determining a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects.

[0163] At step 1408, the method 1400 comprises estimating one or more layers and corresponding information of each object of the plurality of objects in the input image.

[0164] At step 1410, the method 1400 comprises combining one or more objects of the plurality of objects into a plurality of groups, using the information of the one or more layers and the placement relation of the one or more objects, suitable for the state.

[0165] At step 1412, the method 1400 comprises generating one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content-aware element interaction motion frames for the state.

[0166] Figures 15A-15B illustrate an exemplary scenario of personalized immersive screens of electronic devices, in accordance with an embodiment of the present disclosure.

[0167] To specify, the Figures 15A-B shows contextual effects on the display based on non-tactile stimuli such as diegetic (i.e. in-world) scenarios from wallpaper impactors (e.g. dark mode, low battery, power saving modes) for complete immersion.

[0168] Figure 15A-A a wallpaper on the electronic device in lock screen mode.Figure 15A-B shows change in the contrast of the objects in the wallpaper when the mode changes to “dark mode LS”. Figure 15A-C illustrates the wallpaper on lock screen in “Low Battery / Power Saving mode”. The features are noticeable on the wallpaper: steam reduced: semantically indicating low battery, average brightnesslowered as a functional response to low battery, lowered brightness displayed inworld, e.g., through ‘cloudy’ scenario.

[0169] Figure 15B-A shows wallpaper during Day Time, Figure 15B-B shows change in the contrast as well as change from “Sun” to “Moon” on the wallpaper during Nigh time. Figure 15B-C illustrates transition of background color from dawn / dusk time. Figure 15B-D illustrates an exemplary wallpaper change when a charging cable is plugged into the electronic device.

[0170] Figures 16A-16C and Figures 17A-17B illustrate an exemplary scenario of personalized immersive screens on electronic devices based on different state and various factors, in accordance with an embodiment of the present disclosure.

[0171] Figure 16A-16C illustrates wallpaper personalization based on the state of the electronic device such as AOD, LS and HS.

[0172] For AOD in Figure 16A, composition is: element density lowered and use ‘night’ version for power saving. Further, in AOD state: minimal elements permitted, elements with large motions are not allowed, and when sensible, use ‘night’ version of scene (e.g.,: sun -> moon, daytime sky -> night sky).

[0173] Table 1 below shows the property, behaviour and exceptions allowed for AOD.Table 1

[0174] For LS in Figure 16B, composition is: maximum level of detail allowed in LS and avoid detail in Time block, Shortcut block. Further in LS, complete backgroundelements permitted, large number of elements permitted: (Animating and / or nonanimating) and minor environmental elements visible (moon, stars, etc.).

[0175] Table 2 below shows the property, behaviour and exceptions allowed for LS.Table 2

[0176] For HS in Figure 16C, the composition is elements number may be reduced compared to LS, element placement can vary from AOD / LS and composition primarily created from background elements. Further in HS, elements of min-max x-y sizes allowed to overlap with III elements and overlap with III elements primarily allowed in top half of screen.

[0177] Table 3 below shows the property, behaviour and exceptions allowed for HSS.Table 3

[0178] Figure 17A-17B illustrate an example use case of the home scree swipe experience, in accordance with an embodiment of the present disclosure. It shows parallax responses where applicable upon swipe across home screens contribute to world immersion.

[0179] Figure 17A shows a scene which contains multiple elements in different layers that are expanded in parallax. As seen, lily pads move a large distance ‘x’ and Fish move a smaller distance ‘y’ and home screen icons move entire screen’s width. Thus, objects in the distance remain largely stationary, objects in the middle and front ground move in accordance with the laws of parallax, in the direction of the swipe.

[0180] Figure 17B shows a scene that does not contain enough elements to create parallax effect. Thus, swiping reveals more of the canvas to the right or left of the main home screen.

[0181] Figure 18A-18C illustrates a pictorial depiction of category immersive screen display based on different electronic devices, in accordance with an embodiment of the present disclosure.

[0182] It shows that the present disclosure enables tailored experience to connect and work across various form factors. The experience is dynamic in terms of special wake animation and continuous elements.

[0183] For example, for differing screen sizes and cover / inner screen functions, the present disclosure facilitates simplified, or expanded, or re-composed version of the artwork, mini-interactions and animations on Flip cover, focus on one enlarged detail from the image, animations suitable to screen size and notifications may be indicated through certain event animations (e.g., Bird for new notification).

[0184] Thus, the present disclosure enables generation of motion video for the elements in the user image which are identified suitable for each state using: elements motion feature encoding, state specific elements guidance, Spatial-motionrelations obtained from relation graph and contextual effects guidance. The disclosure extract physics-aware characteristic motion feature extraction for the elements present in the input scene. The disclosure identifies the layers suitable for particular states (Ex: AOD / LS / HS etc.) by classifying each layer obtained from elements layerization (each layer comprises of one or more element)

[0185] In this application, unless specifically stated otherwise, the use of the singular includes the plural, and the use of “or” means “and / or.” Furthermore, use of the terms “including” or “having” is not limiting. Any range described herein will be understood to include the endpoints and all values between the endpoints. Features of the disclosed embodiments may be combined, rearranged, omitted, etc., within the scope of the invention to produce additional embodiments. Furthermore, certain features may sometimes be used to advantage without a corresponding use of other features.

[0186] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.

[0187] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.

[0188] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.

[0189] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.

Claims

Claims

1. A method (1400) for generating a content-aware element interaction motion frames for a state on a display of an electronic device, the method comprising: receiving (1402) an input image having a plurality of objects; extracting (1404) physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation; determining (1406) a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects; estimating (1408) one or more layers and corresponding information of each object of the plurality of objects in the input image; combining (1410) one or more objects of the plurality of objects into a plurality of groups, using the information of the one or more layers and the placement relation of the one or more objects, suitable for the state; and generating (1412) one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content-aware element interaction motion frames for the state.

2. The method (1400) as claimed in claim 1 , wherein extracting the physics-informed motion features comprises: obtaining a spatial latent representation, a semantic map representation, and a depth map representation of the input image; concatenating the spatial latent representation with the semantic map representation to obtain a spatial representation of elements of the plurality of objects in the input image; multiplying the spatial latent representation with the depth map representation and concatenating the multiplied representation with the semantic map representation to obtain a depth-wise representation of the elements of the objects in the input image; and concatenating the spatial representation of the elements of the plurality of objects with the depth-wise representation of elements to obtain the motion representation features of elements of the plurality of objects in the input image by predicting physical properties of each of the plurality of objects.

3. The method (1400) as claimed in claim 2, wherein determining the spatial motion relation of each object with each of the plurality of objects comprises: segregating one or more elements of the plurality of objects from the motion representation features of elements; performing a cross-attention operation on the one or more elements of the plurality of objects based on the spatial representation of elements, the depth-wise representation of elements, and the motion representation of elements; and determining spatial motion relationship among the spatial representation, the depth-wise representation, and motion relationship among the spatial representation, depth-wise representation and the motion representation of the elements based on the cross-attention operation on the one or elements of the plurality of objects to obtain an element relation graph and a plurality of bounding boxes of the elements of the plurality of objects.

4. The method (1400) as claimed in claim 3, wherein determining the placement relation and estimating the one or more layers of each object with each of the plurality of objects comprises:utilizing the element relation graph, the plurality of bounding boxes of the elements of the plurality of objects, and the segregated one or more elements of the plurality of objects to construct an edge distance-wise layers; and identifying the one or more layers by classifying each layer obtained from the edge distance-wise layers.

5. The method (1400) as claimed in claim 1 , wherein the spatial motion relation corresponds to the one or more elements present in the input image and a relation between the one or elements with respect to the motion representation and the placement relation.

6. The method (1400) as claimed in claim 1 , comprising: identifying the one or more layers for the state by: obtaining one or more spatial features of the one or more layers using a predefined pretraining model; determining the one or more elements in the one or more layers suitable for interactions and motion based on the one or more spatial features; determining one or more components of a layout on the display of the electronic device using the pre-trained model; determining if any of the one or more components of the layout is obstructing the one or more elements for the motion to obtain one or more suitable components of the layout; and identifying the one or more layers for the state by determining the motion of the one or more elements suitable on the display of the electronic device based on the one or more suitable elements and the one or more suitable components of the layout.

7. The method (1400) as claimed in claim 1 , wherein generating the one or more motion vectors comprises: obtaining a state specific element vector based on the identified one or more layers information for the state; obtaining motion features of the one or more suitable elements for the state from the segregated one or more elements of the plurality of objects; obtaining the textual features of the one or more suitable elements from the element relation graph; obtaining the state properties and contextual factors of the electronic device; generating the motion frames using the motion features of the one or more suitable elements across temporal dimensions; generating the interactive motion frames based on the generated motion frames and obtained the element relation graph; generating the one or more motion vectors for the state based on the state specific element vector, the motion frames, the interactive motion frames, the textual features, the state properties and the contextual factors of the electronic device.

8. The method (1400) as claimed in claim 7, wherein the state properties include one or more of colour range, colour contrast, colour palette, average brightness, shadows, highlights, flat colour, true black, luminance, gradients and the contextual factors of the electronic device include one or more of battery state, charging state, time of day, current weather, touch interactions, gyroscopic motion.

9. A system (210) for generating a content-aware element interaction motion frames for a state on a display of an electronic device, the system comprising: one or more processors (302); a memory (304) coupled with the one or more processors (302), wherein the one or more processors (302) are configured to: receive an input image having a plurality of objects; extract physics-informed motion features of each object of a plurality of objects present in the input image to obtain motion representation; determine a spatial motion relation of each object with each of the plurality of objects based on the motion representation and a placement relation of each object with each of the plurality of objects; estimate one or more layers and corresponding information of each object of the plurality of objects in the input image; combine one or more objects of the plurality of objects into a plurality of groups, using the information of one or more layers and the placement relation of the one or more objects, suitable for the state; and generate one or more motion vectors for each of the plurality of groups using the spatial motion relation and the placement relation to obtain the content-aware element interaction motion frames for the state.

10. The system (210) as claimed in claim 9, wherein to extract the physics-informed motion features, the one or more processors (302) are configured to: obtain a spatial latent representation, a semantic map representation, and a depth map representation of the input image; concatenate the spatial latent representation with the semantic map representation to obtain a spatial representation of elements of the plurality of objects in the input image; multiply the spatial latent representation with the depth map representation and concatenating the multiplied representation with the semantic map representation to obtain a depth-wise representation of the elements of the objects in the input image; and concatenate the spatial representation of the elements of the plurality of objects with the depth-wise representation of elements to obtain the motion representation features of elements of the plurality of objects in the input image by predicting physical properties of each of the plurality of objects.

11. The system (210) as claimed in claim 10, wherein to determine the spatial motion relation of each object with each of the plurality of objects, the one or more processors (302) are configured to: segregate one or more elements of the plurality of objects from the motion representation features of elements; perform a cross-attention operation on the one or more elements of the plurality of objects based on the spatial representation of elements, the depth-wise representation of elements, and the motion representation of elements; and determine spatial motion relationship among the spatial representation, the depth-wise representation, and motion relationship among the spatial representation, depth-wise representation and the motion representation of the elements based on the cross-attention operation on the one or elements of the plurality of objects to obtain an element relation graph and a plurality of bounding boxes of the elements of the plurality of objects.

12. The system (210) as claimed in claim 11 , wherein to determine the placement relation and estimating the one or more layers of each object with each of the plurality of objects, the one or more processors (302) are configured to: utilize the element relation graph, the plurality of bounding boxes of the elements of the plurality of objects and the segregated one or more elements of the plurality of objects to construct an edge distance-wise layers; and identify the one or more layers by classifying each layer obtained from the edge distance-wise layers.

13. The system (210) as claimed in claim 9, wherein the spatial motion relation corresponds to the one or more elements present in the input image and a relation between the one or elements with respect to the motion representation and the placement relation.

14. The system (210) as claimed in claim 9, wherein to identify the one or more layers for the state, the one or more processors (302) are configured to: obtain one or more spatial features of the one or more layers using a predefined pre-training model; determine the one or more elements in the one or more layers suitable for interactions and motion based on the one or more spatial features; determine one or more components of a layout on the display of the electronic device using the pre-trained model; determine if any of the one or more components of the layout is obstructing the one or more elements for the motion to obtain one or more suitable components of the layout; and identify the one or more layers for the state by determining the motion of the one or more elements suitable on the display of the electronic device based on the one or more suitable elements and the one or more suitable components of the layout.

15. The system (210) as claimed in claim 9, wherein to generate the one or more motion vectors, the one or more processors (302) are configured to: obtain a state specific element vector based on the identified one or more layers information for the state; obtain motion features of the one or more suitable elements for the state from the segregated one or more elements of the plurality of objects; obtaining the textual features of the one or more suitable elements from the element relation graph; obtain the state properties and contextual factors of the electronic device; generate the motion frames using the motion features of the one or more suitable elements across temporal dimensions; generate the interactive motion frames based on the generated motion frames and obtained the element relation graph; generate the one or more motion vectors for the state based on the state specific element vector, the motion frames, the interactive motion frames, the textual features, the state properties and the contextual factors of the electronic device.

Citation Information

Patent Citations

  • Apparatus and method for providing animation effect in portable terminal

    US20110216076A1

  • Action recognition method and apparatus, computer storage medium, and computer device

    US20220076002A1

  • Panoptic segmentation forecasting for augmented reality

    US20220319016A1

  • Spatial motion attention for intelligent video analytics

    US20230111865A1

  • Method and system of image processing for action classification

    US20230274580A1

Cited By

  • Fish school individual tracking and monitoring method, device and system

    CN120744453A