Methods and systems for transforming real-world entities shadow and reflection in extended reality device

The XR system transforms shadows and reflections of real-world entities by determining semantic information and generating virtual appearances, addressing the lack of contextual information in existing XR technologies and enhancing user immersion.

WO2026104944A1PCT designated stage Publication Date: 2026-05-21SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-11-05
Publication Date
2026-05-21

Smart Images

  • Figure IB2025061278_21052026_PF_FP_ABST
    Figure IB2025061278_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are systems (203) and methods (600) for transforming real-world entities shadow and reflection in extended reality device. Initially, semantic information associated with a field of view (FOV) of the XR device (201) is determined. Further, a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities are determined. Further, the one or more virtual appearances are generated. Finally, the one or more virtual appearances are overlaid on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.
Need to check novelty before this filing date? Find Prior Art

Description

DescriptionTitle of Invention :METHODS AND SYSTEMS FOR TRANSFORMING REAL-WORLD ENTITIES SHADOW AND REFLECTION IN EXTENDED REALITY DEVICETechnical Field

[0001] The present disclosure relates to the field of extended reality (XR) technology, and more particularly relates to methods and systems for transforming at least one of a shadow and a reflection of one or more real-world entities in an XR display device.Background Art

[0002] Technology pertaining to extended reality (XR) encompasses three main categories, i.e., virtual reality (VR), augmented reality (AR), and mixed reality (MR). Said technologies utilize wearables, such as headsets and glasses, to superimpose digital content onto the physical world.

[0003] Presently, existing XR technology is focused on offering an immersive experience by overlaying digital content on the physical world entities. An exemplary implementation of the existing XR technology is described below in conjunction with FIGS. 1 A and IB.

[0004] FIGS. 1 A and IB are schematic diagrams 100 depicting an exemplary implementation of existing XR technology, according to an existing art. As shown in FIG. 1 A, consider an exemplary scenario where user 101 and his friend 103 are present in room 105. Besides the user 101 and his friend 103, a refrigerator 109 may be present in one corner of the room 105. Further, the user 101 may be wearing an XR display device 107 to engage in an immersive experience with the surrounding environment.

[0005] Further, as shown in FIG. IB, as the user 101 interacts with the environment using the XR display device 107, a pair of virtual glasses 111 may be overlaid on the friend 103 such that the user 101 perceives the friend 103 as wearing glasses. Additionally, a contextual virtual object 113 may be overlaid on refrigerator 109 based on the context of the environment within room 105. Thus, according to the current implementation of the existing XR technology, even though the virtual content is overlaid on the real-world physical entities to provide an immersive experience, corresponding ground shadows of the real world physical entities remain the same as in the real world.

[0006] Further, the existing XR technology is extended to enable overlaying shadows and reflections of the digital content in the XR environment. Moreover, to enhance the immersive experience of the user, the existing XR technology may also enable transforming the shadows and reflections of the digital content.

[0007] Thus, the existing XR technology describes the transformation of the physical world entities and even the transformation of the shadows and reflections of the digital content in an XR environment. However, the original ground shadow and reflection of the physical world entities remain unchanged.

[0008] Notably, in the current implementation of existing XR solutions, the ground shadows and reflection remain the same as corresponding physical world entities. Further, an aspect of utilizing the original ground shadow and reflection properties of the physical world entities to convey contextual information has remained unexplored.

[0009] Further, current XR technologies lack comprehensive solutions for dynamically transforming ground shadows and reflections of physical world entities. While advancements in XR technologies have enabled overlaying virtual objects onto the real physical world entities, an aspect pertaining to interaction with ground shadows and reflections have remained under explored.

[0010] Moreover, existing XR technologies focus on overlaying digital content on the physical world entities while maintaining the original ground shadow and reflection properties without transformation. Thus, a significant gap lies in the existing XR technologies where the transformation of ground shadow and reflections to convey contextual information to user in an XR environment is unexplored.

[0011] Accordingly, there is a need to overcome the above-described limitations of the existing XR technology. Additionally, there is a need to provide a methodology for providing an immersive experience in an XR environment by utilizing the original ground shadow and reflection properties of the physical world entities.Solution to Problem

[0012] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.

[0013] According to one embodiment of the present disclosure, a method for transforming at least one of a shadow and a reflection of one or more real-world entities is disclosed. The method includes determining semantic information associated with the field of view (FOV) of the XR device (201). The method further includes determining a plurality of attributes for generating one or more virtual appearances corresponding to at least one of the shadows and the reflection of the one or more real world entities based on a context vector, the semantic information, and segmentation information associated with an image of the FOV. Furthermore, the method includes generating the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV. Moreover, the method includes overlaying the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.

[0014] According to another embodiment of the present disclosure, An extended reality(XR) display device for transforming at least one of a shadow and a reflection of one or more real-world entities is disclosed. The XR display device comprises a memory and a processor coupled with the memory. The processor is configured to determine semantic information associated with a field of view (FOV) of the XR device (201). Further, the processor is configured to determine a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities based on a context vector, the semantic information, and segmentation information associated with an image of the FOV. Furthermore, the processor is configured to generate the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV. Moreover, the processor is configured to overlays the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.

[0015] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.Brief Description of Drawings

[0016] The foregoing and other features of embodiments will become more apparent from the following detailed description of embodiments when read in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements.

[0017] FIG. 1A is a schematic diagram depicting an exemplary implementation of existing XR technology, according to an existing art;

[0018] FIG. IB is a schematic diagram depicting an exemplary implementation of existing XR technology, according to an existing art;

[0019] FIG. 2 is a block diagram depicting an exemplary extended reality (XR) device, according to embodiments of the present disclosure;

[0020] FIG. 3 is a block diagram depicting the plurality of modules of the system for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure;

[0021] FIG. 4 is a block diagram depicting an overall flow of operations among the plurality of modules, according to the embodiments of the present disclosure;

[0022] FIG. 5 is a schematic diagram depicting an exemplary scenario illustrating the implementation of the system for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure;

[0023] FIG. 6 is a flow diagram depicting a method for transforming at least one of a shadow and a reflection of one or more real-world entities, according to an embodiment of the present disclosure; and

[0024] FIG. 7A is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0025] FIG. 7B is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0026] FIG. 7C is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0027] FIG. 7D is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0028] FIG. 7E is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0029] FIG. 7F is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0030] FIG. 7G is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0031] FIG. 7H is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0032] FIG. 71 is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0033] FIG. 7J is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0034] FIG. 7K is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0035] FIG. 7L is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0036] FIG. 7M is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0037] FIG. 7N is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0038] FIG. 70 is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0039] FIG. 7P is an exemplary use case implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.Description of Embodiments

[0040] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.

[0041] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.

[0042] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as “one or more features” or “one or more elements” or “at least one feature” or “at least one element.” Furthermore, the use of the terms “one or more” or “at least one” feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, “there needs to be one or more...” or “one or more elements is required.”

[0043] Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfill the requirements of uniqueness, utility, and non-obviousness.

[0044] Use of the phrases and / or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “a further embodiment”, “furthermore embodiment”, “additional embodiment” or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment or maybe found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, in the context of more than one embodiment, or the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

[0045] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.

[0046] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.

[0047] A detailed methodology is explained in the following paragraphs of the disclosure.

[0048] An objective of the present disclosure is to provide innovative solutions to enable dynamic transformation and seamless integration of virtual reflections and shadows within real-world contexts in an XR environment.

[0049] Additionally, an objective of the present disclosure is to enhance real -world user experience by dynamically overlaying virtual shadows and reflections onto ground shadows and reflections of real-world entities in the XR environment.

[0050] The above-mentioned objectives are achieved by providing a methodology for enhancing user experience for a user wearing an extended reality (XR) display deviceby transforming at least one of shadow and a reflection of one or more real-world entities.

[0051] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.

[0052] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the figure number, in which the corresponding component is shown. For example, reference numerals starting with digit “1” are shown at least in Fig. 1. Similarly, reference numerals starting with digit “2” are shown at least in Fig. 2.

[0053] FIG. 2 is a block diagram 200 depicting an exemplary extended reality (XR) device 201, according to embodiments of the present disclosure. The XR device 201 may correspond to a device that combines virtual reality (VR), augmented reality (AR), and mixed reality (MR) providing an immersive and interactive experience to a user by overlaying digital information onto real world entities and also displaying virtual objects in the real world environments.

[0054] In an example, the XR device 201 may be a head mounted device (HMD) to be worn on the head of the user immersing the user in a virtual environment and / or a physical environment simultaneously. In another example, the XR device 201 may be a pair of smart glasses which overlay digital information on the real world entities, allowing the user to see physical and / or virtual elements simultaneously. In yet another example, the XR device 201 may be a handheld AR device such as smartphones or tablets which act as AR device when paired with dedicated accessories.

[0055] According to embodiments of the present disclosure, the XR device 201 may implement a system 203 for enhancing user experience for a user using the XR device 201 by transforming at least one of a shadow and a reflection of one or more real -world entities. The system 203 may include a memory 205, a processor(s) 207, artificial intelligence (Al) model(s) 209, a plurality of sensors 211, at least one display 213, and a plurality of modules 215.

[0056] In an example, the processor(s) 207 may be a single processing unit or a number of units, all of which could include multiple computing units. The processor 207 may beimplemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logical processors, virtual processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 207 is configured to fetch and execute computer-readable instructions and data stored in the memory 205.

[0057] The memory 205 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory or Random Access Memory (RAM), such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0058] At least one of a plurality of operations of the system 203 may be implemented through the Al model(s) 205. A function associated with Al may be performed through the non-volatile memory, the volatile memory, and the processor 207.

[0059] The processor(s) 207 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor such as a neural processing unit (NPU).

[0060] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or Al model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0061] Here, being provided through learning means that, by applying a learning technique to a plurality of learning data, a predefined operating rule or Al model of a desired characteristic is made. The learning may be performed in a device itself in which Al according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0062] The Al model(s) 209 may refer to one or more machine learning models pre-trained using predetermined neural network techniques. The Al model(s) 209 may consist ofa plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through the calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0063] The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to decide or predict. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0064] The Al model(s) 209 may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic Al model with multiple pieces of training data by a training technique. The Al model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[0065] The plurality of sensors 211 may be configured to capture real-time environment data from the field of view (FOV) of the XR device 201. According to embodiments of the present disclosure, the real-time environment data may include visual and spatial information of the real-world environment in the FOV of the XR device 201. The plurality of sensors 211 may include but are not limited to, red green blue (RGB) sensor, infrared depth sensor, one or more accelerometers, one or more gyroscopes, one or more magnetometers, ambient light sensor, and proximity sensor.

[0066] In an embodiment, the depth sensor may be utilized to map the real-time environment in three-dimensional (3D) data for accurately overlaying virtual images onto real- world surfaces and for shadow and reflection transformations.

[0067] In an embodiment, the accelerometer may be utilized to determine when the user wearing the XR device 201 has moved in a way that would affect the reflection andshadow data for calibrating real-world positioning for virtual reflection and shadow overlay. Further, the gyroscope may be utilized to obtain real-time spatial orientation information and angular velocity. The data obtained from the gyroscope may be used for aligning and overlaying virtual reflections and shadows on to real-world reflections and shadows.

[0068] Further, magnetometer may be utilized to obtain data useful for keeping virtual reflection and shadow directionally accurate aiding in rotational adjustments of virtual reflections and shadows. Furthermore, the ambient light sensors may be utilized to measure environmental lightning conditions for adjusting the brightness and contrast of virtual reflections and shadow overlays for a realistic integration with the physical environment. Furthermore, data from the proximity sensor aid in dynamically adjusting virtual reflections and shadows based on user proximity, ensuring immersive and contextually aware experiences.

[0069] In an example, the display 213 of the XR device 201 may be utilized to present visual information to the user. The display 213 may take various forms depending on the type of XR environment. For example, in VR, the display may be a headset with a high- resolution screen that provides an immersive experience. In another example, in AR, the display may be a smartphone or tablet screen with AR-enabled software that superimposes digital information onto the real world. In another example, in MR, the display may be a combination of VR and AR, providing an immersive and interactive experience.

[0070] Further, the plurality of modules 215 may include a program, a subroutine, a portion of a program, a software component, or a hardware component capable of performing a stated task or function. As used herein, the plurality of modules 215 may be implemented on a hardware component such as a server independently of other modules, or a module can exist with other modules on the same server, or within the same program. The the plurality of modules 215 may be implemented on a hardware component, such as processor one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. The plurality of modules 215 when executed by the processor(s) 207 may be configured to perform any of the functionalities discussed herein.

[0071] In an embodiment, the plurality of modules 215 may be implemented using one or more artificial intelligence (Al) modules that may include a plurality of neural network layers. Examples of neural networks include but are not limited to, Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), and Restricted Boltzmann Machine (RBM). Further, ‘learning’ may be referred to in the disclosure as a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the present subject matter’s mechanism through an Al model.

[0072] A function associated with an Al module may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general- purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor, such as a neural processing unit (NPU). One or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0073] The plurality of modules 215 may include a set of instructions that may be executed according to the embodiments of the present disclosure for enhancing user experience for a user using the XR device 201 by transforming at least one of a shadow and a reflection of one or more real -world entities. The plurality of modules 215 are described in the forthcoming paragraphs in detail in conjunction with FIG. 3, FIG. 4, and FIG. 5.

[0074] FIG. 3 is a block diagram 300 depicting the plurality of modules 215 of the system 203 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure. Further, FIG. 4 is a block diagram 400 depicting an overall flow of operations among the plurality of modules215, according to the embodiments of the present disclosure. Further, FIG. 5 is a schematic diagram 500 depicting an exemplary scenario illustrating the implementation of the system 203 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure. FIG. 3 will be described in conjunction with FIG. 4 and FIG. 5 for the sake of brevity and ease of reference.

[0075] The plurality of modules 215 may include input processing module 301, context vector generation module 303, attribute prediction module 305, virtual appearance generation module 307, and image composition module 309. According to embodiments of the present disclosure, some of the plurality of modules may be implemented on the XR device 201 and the rest of the plurality of modules may be implemented on cloud.

[0076] Initially, at operation 401 , an input is received at the XR device 201. The input includes an image of the field of view (FOV) of the XR device captured using the plurality of sensors 211 within the XR device 201. The FOV of the XR device 201 may include one or more real world entities. The one or more real-world entities correspond to at least one of one or more animate subjects and one or more inanimate objects.

[0077] In the exemplary scenario depicted in FIG. 5, a user 501 may be wearing the XR device 201. In the depicted scenario, the FOV of the XR device 201 may include one or more real world entities such as other user 503, and a smart refrigerator 505. According to the embodiments of the present disclosure, the other user 503 may correspond to an animate subject, while the smart refrigerator 505 may correspond to an inanimate object.

[0078] According to the embodiments of the present disclosure, at operation step 403, the input processing module 301 may be configured to process the input and determine semantic information associated with the FOV of the XR device 201. The semantic information includes the one or more real-world entities in the FOV. For example, the other user 503, and the smart refrigerator 505.

[0079] The semantic information further includes first contextual information associated with the one or more real world entities. According to the embodiments of the present disclosure, the first contextual information corresponds to unknown information of theone or more real world entities with respect to the user 501 wearing the XR device 201.

[0080] Further, the input processing module 301 may further be configured to determine second contextual information associated with the with the user 501 wearing the XR device 201. According to the embodiments of the present disclosure, the second contextual information may be obtained from at least one of one or more internet of things (loT) devices coupled with the XR device 201 and one or more social media platforms associated with the user 501 wearing the XR device 201.

[0081] According to embodiments of the present disclosure, the input processing module 301 may caption image data obtained from a plurality of sources associated with the user 501 wearing the XR device 201 and the one or more real world entities, and generate textual semantics to derive contextual information for further processing of context derivation. In an embodiment, the input processing module 301 may use a predefined combination of convolutional neural network (CNN) and long short term memory (LSTM) model to convert image data to textual data which will be utilized in context generation module.

[0082] In an example, the combination of CNN and LSTM may be a wavelet transform based convolutional neural network (WCNN) which uses two level discrete wavelet decomposition for extracting the visual feature maps highlighting the spatial, spectral and semantic details from the image data. Further, the LSTM may be used to achieve a probability of appropriate word prediction for the image data.

[0083] Further, the textual data and the captioned image data are processed using word embedding techniques. According to embodiments of the present disclosure, the word embedding refer to a technique where individual words of a domain or language are represented as real-valued vectors in a lower dimensional space. Said sub-step generates a common vector space representation of entire data set from multiple smart devices, such as the one of one or more internet of things (loT) devices. Each element of the generated vector space embodies specific textual data captured from respective data source inputs. Further, the textual data feeds along with image captions are processed as generate word vectors using pre-trained models such as global vectors for words representation (GloVe).

[0084] The word embedding results in a set of vectors, one for each piece of textual data collected from the multiple smart devices, that collectively form a context vector space. The context vector space is representative of the entirety of collected data and is utilized for subsequent processing using multi-depth clustering and prioritization. Multi-depth clustering technique groups data points in a dataset of the textual data and captioned image data and create a hierarchical structure within the context vector space, identifying and grouping similar context vectors to form clusters and subclusters.

[0085] In a first-tier clustering of the multi-depth clustering, an unsupervised machine learning algorithm, such as K-means, DBSCAN, or hierarchical clustering, is applied to the context vector space and separates the context vector space into different broad clusters based on natural groupings of the data points.

[0086] Further, in the second- tier clustering of the multi-depth clustering breaks down broad clusters are broken into more specific, prioritized contexts. Such level of clustering can capture more nuanced information about preferences or behaviors of the user (user 501 or the other user 503) enabling improved personalization. Upon clustering, the contexts are prioritized to it determine relevance and importance of information to be displayed to the user 501. During context prioritization, a hierarchical ranking is assigned to each of second-tier clusters based on various factors such as density of data points within cluster which is indicative of frequency of a specific behavior or preference, and recency of data points which reflect current relevance of context, or any other relevant factors.

[0087] Upon context prioritization, context labeling is performed. The context labeling step assigns labels to each of the prioritized, second tier clusters. In an embodiment, the labels may be generated using methods like centroid-based keyword extraction or topic modeling, to provide a human-understandable description of each context.

[0088] Further, the input processing module 301 may perform visual scene understanding on the image of the FOV. In an embodiment, the visual scene understanding may be performed using a hierarchical approach that reasons over multiple levels of domain knowledge with increasing granularity in order to generate accurate scene graphs for both corrupted and clean images. In an example, the hierarchical approach may be thehierarchical knowledge enhanced robust scene graph generation (HiKER-SGG) approach.

[0089] According to the embodiments of the present disclosure, when the one or more real world entities correspond to the one or more animate subjects (for example, the other user 503), the input processing module 301 determines the first contextual information by obtaining first information relating to the one or more animate subjects from one or more social media platforms associated with the one or more animate subjects.

[0090] Thereafter, the input processing module 301 determines first mutual information relating to the user wearing the XR device and the one or more animate subjects based on the obtained first information and the second contextual information.

[0091] Thereafter, the input processing module 301 determines the first contextual information corresponding to the unknown information of the one or more animate subjects based on a difference between the first information and the first mutual information. According to embodiments of the present disclosure, the first contextual information is determined by calculating the correlation between each friend of the user 501 via a technique called Friend2Vec. Friend2Vec is based on a multilayer architecture represented using equation (1), such that each layer represents each social platform with different weights depending on relative engagement on each platform. So, each message / posts is shared and response is exchanged.yjfnP, ;*(Sentiment value of j's response) , , , . „ .

[0092] -yW yM p.. > . . (1), wherein P is a social media post, i is Li=o^j=o^ijnumber of social media posts, and j is number of friends.

[0093] According to embodiments of the present disclosure, when the one or more real world entities is an animate subject, for example the another user 505, the input processing module 301 may generate the correlation between the friends and create a two- dimensional (2D) matrix as depicted below in Table 1. Said 2D correlation matrix is dynamically updated with every information update associated with the social media of the user 501.Table 1

[0094] Further, for a j-th post / message Mj, an Z-th friend Ft, has either interacted with the post (encoded as 1) or not (encoded as 0). Mj is a central node and forms a connection with Fi. The interaction is considered to be at least one of viewing, reaction, and comment to the Mj . With said interaction data, a matrix as depicted in table 2 for the user for the interaction of post is created.Table 2

[0095] Using this matrix, a cluster is created for each friend Ft, with the unknown message or post Mj.

[0096] Each message Mj node in the cluster is weighted with a given formula defined using equation (2).

[0097] WM. > is the sentence importance score of themessage or post Mj. In an example, the importance score is calculated by using a modified KL divergence technique. In implementation, the KL divergence technique removes each word individually in the sentence and calculate the semantic value of the sentence with and without the word. If the semantic value remains within a predetermined threshold, the word is considered to be non-important. Using the same technique, one or more important words in the sentence are determined and corresponding weighted average is computed. The importance score (I) is normalized count of important words present in the sentence.

[0098] For example, the other user 503 may have shared a post celebrating his graduation.However, said information is unknown to the user 501. The above-described methodology may be utilized to identify the unknown post (the one celebrating graduation) for a particular friend (i.e., the other user 503) of the user 501.

[0099] According to embodiments of the present disclosure, when the one or more real world entities correspond to the one or more inanimate objects, for example the smart refrigerator 505, the input processing module 301 may determine the first contextual information by determining associated interests of the user 501 wearing the XR device 201 relating to the one or more inanimate objects based on digital fingerprints of the user 501 wearing the XR device. Thereafter, the input processing module 301 may determine the first contextual information based on the determined associated interests of the user wearing the XR device.

[0100] According to embodiments of the present disclosure, the semantic information further includes light source information associated with one or more light sources in the FOV. According to embodiments of the present disclosure, the input processing module 301 may determine the light source using a predetermined method for light source estimation using a deep neural network to learn a functional relationship between the input red green blue depth (RGB-D) image of the scene and a dominant light direction. The deep neural network network may be trained only once on a variety of scenes, and then, could be utilized to be apply in a new scene with an arbitrary geometry.

[0101] According to embodiments of the present disclosure, light source information estimation is done by the neural network independent of camera pose and bycalculating a transformation to world space after estimation. The light sources are estimated in a coordinate space aligned with camera sensor of the XR device 201.

[0102] For light source estimation, a dominant light direction is modeled in terms of relative Euler angles. The Euler angles are being calculated in a camera coordinate space to make the light source estimation independent of a camera pose. In an embodiment, only two Euler angles (cp and 9) are used to define the direction of a light source. In the neural network, the two Euler angles are directly regressed by the network as cp and 9 are being estimated in the camera coordinate space (i.e., they are relative to camera pose). Further, a relation between the relative Euler angles of light (cp and 9) to the absolute Euler angles of the camera (cpc and 9c) and light (cpl and 91) can be defined using the following equations (3) and (4):

[0103] O1 = cpc + cp . (3)

[0104] 01 = 9c + 9 . (4)

[0105] According to embodiments of the present disclosure, light source estimation is performed using residual blocks of convolutional layers to avoid the problem of vanishing or exploding gradients.

[0106] According to embodiments of the present disclosure, the semantic information further includes one or more attributes associated with an orientation and a dimension of the at least one of the shadow and the reflection of the one or more real-world entities. According to embodiments of the present disclosure, reflection, shadows, and background within the image of the FOV is isolated and segmented using a predetermined GAN model.

[0107] Thereafter, the image processing is performed on the segmented reflection, shadows, and background to determine the one or more attributes of rendering reflection and shadow in reference of one or more real-world entities such as rendering orientation, dimensions (length and width including concave and convex properties). In an embodiment, one or more image properties are processed in grey scale to reduce pixel variation to lower scale or binary. From the binary image, one or more inanimate objects may be counted, labeled, isolated, and measured for properties such as area.Further, for the one or more animate objects, body posture and relative parts positioning may be determined using a deep convolutional neural network (DCNN).

[0108] According to embodiments of the present disclosure, an image block reflection and shadow of the one or more real world entities are considered enclosed area in an ellipse, which is further segmented using imaginary major and minor axis within enclosed area. Thereafter, orientation is determined by calculating axis angles in reference to reflection surface plane, with reference plane in vector form.

[0109] According to embodiments of the present disclosure, for determining dimensions (length and width) of the rendered reflection and shadow major axis length (in pixels) of the major axis of the ellipse that has the same normalized second central moments as the region, returned as a scalar, and minor axis length (in pixels) of the minor axis of the ellipse that has the same normalized second central moments as the region, returned as a scalar.

[0110] According to embodiments of the present disclosure, the input processing module 301 may be configured to determine at least one of a shadow prominence index and a reflection prominence index corresponding to each of the at least one of the shadow and reflection of the one or more real world entities to determine whether to transform at least one of a particular shadow and reflection or not.

[0111] According to embodiments of the present disclosure, first the semantic features for the reflection and shadow are obtained and then the cosine similarity of semantic vectors of reflection and shadow with semantic vector of image is calculated. In an embodiment, the key features are identified in the cases of reflection and shadow. In case of reflection, if the details of the reflection on surface is matching the details of the original object is identified. In case of shadows, the projected shadow is estimated and overlap area with the ground truth of shadow is calculated. The above-mentioned information is used to generate reflection and shadow prominence index (PIR, PIS ) respectively.

[0112] According to embodiments of the present disclosure, prior to determining the shadow prominence index, for a plurality of overlapping shadows of the one or more real world entities, the input processing module 301 may add the plurality of overlapping shadows to a priority queue when the one or more real world entities associated with theplurality of overlapping shadows correspond to the one or more animate subjects. In an embodiment, when the at least one of the one or more real world entities associated with the plurality of overlapping shadows correspond to the one or more animate subjects and at least another one of the one or more real world entities correspond to one of the one or more inanimate objects, the input processing module 301 may add the plurality of overlapping shadows to waiting queue.

[0113] Thereafter, the input processing module 301 may select from the priority queue or the waiting queue, the plurality of overlapping shadows, for determining the shadow prominence index, based on a likelihood of natural separation of the plurality of overlapping shadows due to mobility of the one or more inanimate objects associated with the plurality of overlapping shadows.

[0114] According to embodiments of the present disclosure, for each shadow of the plurality of overlapping shadows, the input processing module 301 may determine an estimated shadow based on the light source information and determine an overlap between the estimated shadow and a corresponding image of a ground truth shadow. Further, the input processing module 301 may determine the shadow prominence index when the overlap is greater than a predefined threshold.

[0115] According to embodiments of the present disclosure, for each of the one or more real world entities, the input processing module 301 may determine a corresponding image semantic vector from corresponding image of the each of the one or more real world entities, and determine a corresponding feature semantic vector associated with corresponding reflection of the each of the one or more real world entities. The the corresponding feature semantic vector may be determined by modifying a convolutional operation, i.e., layer feature extraction mathematics as the reflection entities are more prominent near the object and with increasing distance, the reflection fades away.

[0116] In conventional convolutional operation, conventional operator extracts all the features in an image using the simple multiplication, i.e., £( / * A), where I is the image pixel value, and K is the kernel value. This allows all the features to be extracted have same value in real life. However, in the reflective surfaces, the reflection fades away as the pixel is away from the base image. The illuminance of the later pixels is affected dueto this. Hence, the features in that area of the reflection are lesser significant that the pixels in the top of the reflective surface.

[0117] In one or more embodiments, the convolutional operation id modified as< is the image pixel value, and K is the kernel value. Said modification is adaptive with a such that a is closer to 1 for initial pixels, and reduces gradually when moved away from the base surface.

[0118] Further, the input processing module 301 may determine the reflection prominence index based on a modified cosine similarity between the corresponding image semantic vector and the feature semantic vector such that pixels near the reflective surface in the image of the FOV have higher value than the pixels at the far end of the reflective surface. According to embodiments of the present disclosure, the semantic information may be used to generate context vector as described in the forthcoming paragraphs.

[0119] According to embodiments of the present disclosure, at operation step 405, the input processing module 301 may be configured to determine segmentation information associated with the image of the FOV. In an example, the input processing module 301 may determine the segmentation information using the HiKER-SGG method.

[0120] According to embodiments of the present disclosure, at operation step 407, the context vector generation module 303, determines the context vector based on the first contextual information and the second contextual information. For the determination of the context vector, the first contextual information and the second contextual information is sent to through a neural network (weights vector) which provides weightage (importance) to each factor in the first contextual information and the second contextual information. An output from the weights vector is sent through an encoder based on a neural network to provide a combined weighted context vector. The context vector is used by the attribute prediction module 305 for determining a plurality of attributes for generating one or more virtual appearances, as described below.

[0121] According to embodiments of the present disclosure, at operation step 409, the attribute prediction module 305 may be configured to determine a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities based onthe context vector, the semantic information, and segmentation information associated with an image of the FOV.

[0122] The plurality of attributes for generating one or more virtual appearances include at least one relevant object to be virtually presented corresponding to at least one of the one or more real world entities, motion probability of the corresponding at least one of the one or more real world entities, and orientation and dimension probability of the corresponding at least one of the one or more real world entities in a subsequent FOV of the XR display device.

[0123] According to embodiments of the present disclosure, the attribute prediction module 305 may determine the at least one relevant object and a specific virtual representation of the at least one relevant object based on the context vector, and corresponding one or more attributes associated with the orientation and the dimension of the at least one of the shadow and reflection of corresponding one or more real world entities. In particular, the attribute prediction module 305 may predict from the context vector, the most appropriate object relevant to the context along with a way of presenting that object (for example, a person with graduation cap or just graduation cap) for the particular scene.

[0124] With a labeled training set, having inputs as the context vector along with reflection and shadow attributes (i.e. orientation and dimensions), a neural network is trained with p layers having 2Ap neurons in each layer. The final layer of the neural network provides the probabilities of the objects that could be displayed. During inferencing, the output object to be overlaid is matched with various instances of the images for that object from the memory, and a set of top-k images is obtained. Further, occurrences of full-fledged image and derived image are calculated to understand if a full-fledged image or derived image needs to be sent to the Conditional Generative Network (CGN) in next step using the following formula in equation (5) below:

[0125] O0= max(n(full ), n(derived)) ....(5)

[0126] In an example, if the object to be overlaid is determined to be a comic character, then by full-fledged image means the complete appearance of the comic character, and the derived image mean a logo for the comic character.

[0127] Further, the the attribute prediction module 305 may predict, from the segmented image, if the one or more real world entities (i.e., animate and inanimate) will be changing poses and co-ordinates in a subsequent frame along with the next position. For animate subjects like humans, pets, etc., animate subjects states [V1,V2,V3, •••, VT] are observed, DAGs of said subjects states are created and feed into stacked EN- blocks, each of which includes a DA-GCO and a TUO, aiming to update the attributes of bones and joints along with respective temporal dynamics in the observed subjects states. Finally, the encoder outputs the final update result represented as a DAG. The obtained DAG is then fed into a decoder, which is composed of a DA-GRU and a MLP, sequentially predicting future subject states.

[0128] For inanimate objects, a pre-trained neural network trained on previously available context vectors is utilized which include history data for the positioning of the inanimate objects like connected smart devices. Said data helps the pre-trained neural network to predict if an inanimate object in the FOV is going to change its position.

[0129] Further, the attribute prediction module 305 may determine the motion probability associated with a likelihood of movement of the corresponding at least one of the one or more real world entities in the subsequent FOV based on the segmentation information associated with the image of the FOV. In particular, the attribute prediction module 305 predict, based on the light source information, the next orientation of reflection and shadow for the one or more animate subjects and the one or more inanimate objects.

[0130] Further, the attribute prediction module 305 may determine the orientation and dimension probability based on the segmentation information associated with the image of the FOV, the determined motion probability associated with the subsequent FOV, and the light source information. In particular, the final concatenated vector from animate subjects and inanimate objects, along with the segmented image at t = T are sent to the next frames stitching module where the new image at t = T + 1 is generated. The image at t = T + 1 along with light source information is sent to shadow and reflection generator to generate shadows and reflection for the new image based on the one or more attributes associated with an orientation and a dimension of the at least one of the shadow and the reflection of the one or more real-world entities, and the at least one of a shadow prominence index and a reflection prominence indexcorresponding to each of the at least one of the shadow and reflection of the one or more real world entities.

[0131] According to the embodiments of the present disclosure, at operation step 411, the virtual appearance generation module 307 may generate the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV. According to the embodiments of the present disclosure, the virtual appearance generation module 307 may generate the one or more virtual appearances corresponding to the at least one relevant object based on the orientation and segmentation probability in a pre-determined order of the motion probability using a layered conditional generative network (CGN), for example, n-layered CGN where every n-th CGN creates at the least one of the shadow and the reflection overlay for a new object based on the overlays of previous CGNs enabling the system 203 to ensure that the least one of the shadow and the reflection generated for an object are accurate and immersive.

[0132] A first generator neural network of the n-layered CGN focusses on generating the objects having low probability of movement like the smart refrigerator 505, floor backgrounds, soft toys, dolls, etc. Such objects are referred to as static objects. The first generator neural network generates a base layer over which other dynamic objects are overlaid.

[0133] Referring to FIG. 5 since the other user 503 is standing, the probability of movement of the other user 503 at a particular point of time is mid (for example, approximately 40%). Hence, the second (1st of n CGN) conditional generator neural network (CGN) takes the base image from previous CGN and overlays the animate subject’s (i.e., the other user 503) layer.

[0134] Further, at a later moment, a cat may be relaxing but since is a curious creature, a probability of movement of the cat at the particular point of time is high (for example, approximately 60%). Hence, the third (2nd of n CGN) conditional generator neural network (CGN) takes the base image from previous CGN and overlays the animate subject’s (i.e., cat) layer.

[0135] Further, contextually, a standing dog would have highest probability of movement.Hence, the final (n-th CGN) conditional generator neural network (CGN) takes thebase image from previous CGN and overlays the animate subject’s (i.e., dog’s) layer. Hence producing the final image which will be viewed by the user.

[0136] According to the embodiments of the present disclosure, at operation step 413, the image composition module 309 may be configured to overlaying generated the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed as depicted at operation step 415. In an example, as depicted in FIG. 5, shadow of the other user 503 may be transformed to the shadow 507 of a person wearing a graduation cap informing the user 501 wearing the XR device 201 about the graduation news of the other user 503. Further, the shadow 509 of the smart refrigerator 505 may be transformed to a shadow depiction of earth indicating that the appliance is running efficiently with low energy usage which saves earth.

[0137] FIG. 6 is a flow diagram depicting a method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to an embodiment of the present disclosure. The method 600 includes a series of operations 601 through 607 executed by one or more components of the XR device 201, in particular the processor 207.

[0138] At step 601, the processor 207 determines semantic information associated with a field of view (FOV) of the XR display device. The semantic information includes the one or more real -world entities in the FOV, the first contextual information associated with the one or more real world entities, such that the first contextual information corresponds to unknown information of the one or more real world entities with respect to a user wearing the XR device, the light source information associated with one or more light sources in the FOV, and the one or more attributes associated with an orientation and a dimension of the at least one of the shadow and the reflection of the one or more real-world entities.

[0139] According to the embodiments of the present disclosure, the one or more real-world entities correspond to at least one of one or more animate subjects and one or more inanimate objects. According to the embodiments of the present disclosure, when the one or more real world entities correspond to the one or more animate subjects, the processor 207 obtains the first information relating to the one or more animate subjectsfrom one or more social media platforms associated with the one or more animate subjects, determining first mutual information relating to the user wearing the XR device and the one or more animate subjects based on the obtained first information and the second contextual information. Further, the processor 207 determines the first contextual information comprises by determining the first contextual information corresponding to the unknown information of the one or more animate subjects based on a difference between the first information and the first mutual information.

[0140] According to the embodiments of the present disclosure, the second contextual information corresponds to information associated with the user wearing the XR display device, obtained from at least one of one or more internet of things (loT) devices coupled with the XR display device and one or more social media platforms associated with the user wearing the XR device.

[0141] According to the embodiments of the present disclosure, when the one or more real world entities correspond to the one or more inanimate objects, the processor 207 determines associated interests of the user wearing the XR device relating to the one or more inanimate objects based on digital fingerprints of the user wearing the XR device and determines the second contextual information based on the determined associated interests of the user wearing the XR device 201.

[0142] According to the embodiments of the present disclosure, the the processor 207 determines at least one of a shadow prominence index and a reflection prominence index corresponding to each of the at least one of the shadow and reflection of the one or more real world entities to determine whether to transform at least one of a particular shadow and reflection or not. Prior to determining the shadow prominence index, the processor 207, for a plurality of overlapping shadows of the one or more real world entities, when the one or more real world entities associated with the plurality of overlapping shadows correspond to the one or more animate subjects, adding the plurality of overlapping shadows to a priority queue and when the at least one of the one or more real world entities associated with the plurality of overlapping shadows correspond to the one or more animate subjects and at least another one of the one or more real world entities correspond to one of the one or more inanimate objects, adding the plurality of overlapping shadows to waiting queue.

[0143] Further, the processor 207 selects, from the priority queue or the waiting queue, the plurality of overlapping shadows, for determining the shadow prominence index, based on a likelihood of natural separation of the plurality of overlapping shadows due to mobility of the one or more inanimate objects associated with the plurality of overlapping shadows.

[0144] Furthermore, the processor 207 determines the shadow prominence index for each shadow of the plurality of overlapping shadows by determining an estimated shadow based on the light source information, determining an overlap between the estimated shadow and a corresponding image of a ground truth shadow, and determining the shadow prominence index when the overlap is greater than the predefined threshold.

[0145] Furthermore, the processor 207 determines the reflection prominence index for each of the one or more real world entities by determining a corresponding image semantic vector from corresponding image of the each of the one or more real world entities, determining a corresponding feature semantic vector associated with corresponding reflection of the each of the one or more real world entities, and determining the reflection prominence index based on a modified cosine similarity between the corresponding image semantic vector and the feature semantic vector.

[0146] At step 603, the processor 207 determines a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities based on a context vector, the semantic information, and segmentation information associated with an image of the FOV.

[0147] According to embodiments of the present disclosure, the plurality of attributes for generating the one or more virtual appearances include at least one relevant object to be virtually presented corresponding to at least one of the one or more real world entities, motion probability of the corresponding at least one of the one or more real world entities, and orientation and dimension probability of the corresponding at least one of the one or more real world entities in a subsequent FOV of the XR display device.

[0148] According to embodiments of the present disclosure, the processor 207 determines the plurality of attributes for generating the one or more virtual appearances by determining the at least one relevant object and a specific virtual representation of theat least one relevant object based on the context vector, and corresponding one or more attributes associated with the orientation and the dimension of the at least one of the shadow and reflection of corresponding one or more real world entities, determining the motion probability associated with a likelihood of movement of the corresponding at least one of the one or more real world entities in the subsequent FOV based on the segmentation information associated with the image of the FOV, and determining the orientation and segmentation probability based on the segmentation information associated with the image of the FOV, the determined motion probability associated with the subsequent FOV, and the light source information.

[0149] At step 605, the processor 207, generates the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV. According to embodiments of the present disclosure, the processor 207 generates the one or more virtual appearances by generating the one or more virtual appearances corresponding to the at least one relevant object based on the orientation and dimension probability in a pre-determined order of the motion probability using a layered conditional generative network.

[0150] At step 607, the processor 207, overlays the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.

[0151] FIGS. 7A-7P are exemplary use cases implementing the method 600 for transforming at least one of a shadow and a reflection of one or more real-world entities, according to embodiments of the present disclosure.

[0152] Referring to FIG. 7 A, the user 501 is sitting in living room wearing the XR device 201.The user 501 sees reflection 703 of an air purifier 701 as butterflies which gently fly inside air purifier's reflection. Such reflection may indicate to the user 501 of purification of air and a fresher, healthier environment. Thus, the concept of air cleanliness and freshness is conveyed to the user 501 in a visually engaging manner.

[0153] Referring to the scenario depicted in FIG. 7B, Sophia and her friends are doing intense workout in a gym where Sophia is pushing herself hard, unaware that she’ s dehydrated. Trainer Lisa, only by observing her, may not be able to know her condition. However, when Lisa wears the XR device 201 during training session, the hydration / dehydrationlevel of the trainees may be indicated in corresponding reflections. For example, as depicted in the figure, Lisa sees the red reflection of Sophia signifying dehydration and intervenes in real-time if necessary.

[0154] Referring to the scenario depicted in FIG. 7C, Christina is engaged in a relaxation session with Sophia (wearing the XR device 201). The smartwatch wore by Christina indicates low stress factor. Sophia, via the the XR device 201, sees Christina’s reflection as person sitting in a meditative pose with animated bubbles moving outward which symbolizes release of stress and flow of energy. This gives Sophia an immersive glimpse into Christina’s state of mind and meditation practice leading to enhanced wellness session engagement.

[0155] Referring to the scenario depicted in FIG. 7D, Sophia (wearing the XR device 201) sees Christina’s reflection as flames instead of her real reflection indicating high stress levels of Christina. This gives Sophia an immersive glimpse into Christina’s state of mind. This emotional state visualization enhances empathy, support and communication in various settings such as counselling sessions.

[0156] Referring to the scenario depicted in FIG. 7E, a user wearing the XR device 201, sees in field of vision, red reflection of air purifier indicating degradation in air quality level making air quality instantly recognizable to users wearing the XR device 201. This immediate visual cue provides an intuitive and unobtrusive method of alerting them to change in air quality.

[0157] Referring to the scenario depicted in FIG. 7F, Sophia (wearing the XR device 201) sees washing machine’s reflection transform into flowing water in her kitchen while dryer’s reflection transforms to sun, symbolizing the drying process. This unique feature transforms Sophia’s laundry routine into an enjoyable activity with immersive experience. This transformation of reflections signifies the cleaning and drying processes facilitated by the washer and dryer, reflecting its nature of work.

[0158] Referring to a smart home scenario depicted in FIG. 7G, Sophia (wearing the XR device 201) sees refrigerator reflection displaying a green earth indicating that the appliance is running efficiently with low energy usage which saves earth promoting environmental awareness.

[0159] Referring to the scenario depicted in FIG. 7H, at children’s birthday party, all children wearing the XR device 201 focus is on music system which is playing birthday songs. In the sound bar reflection on surface, the children sees the visual representation of the music beats playing through the sound bar. The beats pulsate, enhancing the user's visual experience. This combination of auditory and visual stimuli creates a multisensory immersion, enriching the overall enjoyment of the music and party.

[0160] Referring to the scenario depicted in FIG. 71, Sophia is having an upcoming trip to New York. Alice wearing XR device 201, sees Sophia’s shadow resembling the Statue of Liberty instead of Sophia’s normal shadow on floor. It serves as a playful yet revealing hint about Sophia’s trip to New York turning an ordinary moment into an unforgettable experience.

[0161] Referring to the scenario depicted in FIG. 7J, John may have missed seeing Matthew’ s social media post about obtaining an masters degree recently. However, while wearing the XR device 201, John notices Matthew’s shadow, where Matthew is wearing his graduation dress, indicating that Matthew has indeed completed his MBA. This enables John to see Matthew’s happy moment in graduation attire despite not having seen the related post.

[0162] Referring to the scenario depicted in FIG. 7K, the XR device 201 enables an engaging way of sharing a excting good news in a highly immersive and personal manner. For example, Alex has not yet shared to his friend Tom that he is going to be a father. Tom wearing the XR device 201 sees shadow of Alex as playful shadow of a baby creating a memorable and emotionally engaging experience.

[0163] Referring to the scenario depicted in FIG. 7L, Alice wearing the XR device 201 looks at Sophia’s reflection and sees an array of virtual balloons in the reflection. This indicates to Alice, that Sophia’s birthday is just around the corner providing a joyful reminder. It serves as joyful reminder for Alice.

[0164] Referring to the scenario depicted in FIG. 7M, John may support a global goal programm. When John wears the XR device 201 in Mathew’s presence who recently donated towards the same lobal goal programm, John sees the donation drive by Mathew in his reflection.

[0165] Referring to the scenario depicted in FIG. 7N, Christina has yet to share about her recently bought new mobile phone with her friend David. David wearing the XR device 201, sees Christina’s reflection in form of her new mobile phone. This enhances social interactions and communication, making the sharing of new information more engaging and delightful among friends.

[0166] Referring to the scenario depicted in FIG. 70, Mathew may have recently purchased a lawn tennis racket, a trendy sports equipment. John, known for his interest in Sports and Lawn Tennis, after wearing the XR device 201 sees Racket in Mathew’s shadow which appears as an overlay within John's mixed reality experience, catching his attention. This scenario shows potential application of the XR device 201 for personalized advertisements or showing interests like playing lawn tennis.

[0167] Referring to the scenario depicted in FIG. 7P, a user wearing the XR device 201 gaze at a pool reflecting the Eiffel Tower, and observes movement in the reflection and sees a time-lapse of the Eiffel Tower’s construction. The time-lapse may illustrate a structure accelerating as Eiffel Tower take shape, reflecting the passing years. Time fast-forwards, showing the Eiffel Tower’s completion and its majestic presence. This provides the user an mesmerizing experience providing a profound glimpse into the history of Eiffel Tower.

[0168] At least by virtue of aforesaid, the present subj ect matter at least provides the following advantages:

[0169] The system and method described herein enable blending of virtual reflections and shadows with the physical environment enhances the realism of XR experiences leading to more immersive and believable interactions with users.

[0170] The system and method described herein focus on overlaying virtual reflections and shadows on physical entities revolutionizing the way the reflections and shadows are perceived and interacted with.

[0171] The system and method described herein enable dynamic manipulation of ground surface reflection and shadow allowing for interactive experiences, that respects real- world physics and lighting conditions.

[0172] The system and method described herein creates an XR experience that opens new possibilities in various fields like entertainment, gaming, fashion, health, tourism, advertising, and more.

[0173] In this application, unless specifically stated otherwise, the use of the singular includes the plural, and the use of “or” means “and / or.” Furthermore, the use of the terms “including” or “having” is not limiting. Any range described herein will be understood to include the endpoints and all values between the endpoints. Features of the disclosed embodiments may be combined, rearranged, omitted, etc., within the scope of the invention to produce additional embodiments. Furthermore, certain features may sometimes be used to advantage without a corresponding use of other features.

[0174] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be appreciated that a vast number of variations exist.

Claims

Claims

1. A method (600), at an extended reality (XR) device (201), for transforming at least one of a shadow and a reflection of one or more real-world entities, the method (600) comprising:determining semantic information associated with a field of view (FOV) of the XR device (201);determining a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities based on a context vector, the semantic information, and segmentation information associated with an image of the FOV;generating the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV; andoverlaying the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.

2. The method (600) as claimed in claim 1, wherein the one or more real -world entities correspond to at least one of one or more animate subjects and one or more inanimate objects.

3. The method (600) as claimed in claim 1, wherein the semantic information includes at least one of:the one or more real-world entities in the FOV,first contextual information associated with the one or more real world entities, wherein the first contextual information corresponds to unknown information of the one or more real world entities with respect to a user wearing the XR device (201),light source information associated with one or more light sources in the FOV, andone or more attributes associated with an orientation and a dimension of the at least one of the shadow and the reflection of the one or more real-world entities.

4. The method (600) as claimed in claim 3, wherein the context vector is determined based on the first contextual information and second contextual information associated with the user wearing the XR device (201).

5. The method (600) as claimed in claim 4, wherein:when the one or more real world entities correspond to the one or more animate subjects, determining the first contextual information comprises:obtaining first information relating to the one or more animate subjects from one or more social media platforms associated with the one or more animate subjects,determining first mutual information relating to the user wearing the XR device (201 ) and the one or more animate subjects based on the obtained first information and the second contextual information, anddetermining the first contextual information corresponding to the unknown information of the one or more animate subjects based on a difference between the first information and the first mutual information; andwhen the one or more real world entities correspond to the one or more inanimate objects, determining the first contextual information comprises:determining associated interests of the user wearing the XR device (201) relating to the one or more inanimate objects based on digital fingerprints of the user wearing the XR device (201), anddetermining the first contextual information based on the determined associated interests of the user wearing the XR device (201).

6. The method (600) as claimed in claim 3, wherein the second contextual information corresponds to information associated with the user wearing the XR device (201), obtained from at least one of one or more internet of things (loT) devices coupled with the XR device (201) and one or more social media platforms associated with the user wearing the XR device (201).

7. The method (600) as claimed in claim 2, further comprising determining at least one of a shadow prominence index and a reflection prominence index corresponding to each of the at least one of the shadow and reflection of the one or more real world entities to determine whether to transform at least one of a particular shadow and reflection or not.

8. The method (600) as claimed in claim 7, wherein prior to determining the shadow prominence index, the method (600) comprises:for a plurality of overlapping shadows of the one or more real world entities: when the one or more real world entities associated with the plurality of overlapping shadows correspond to the one or more animate subjects, adding the plurality of overlapping shadows to a priority queue;when the at least one of the one or more real world entities associated with the plurality of overlapping shadows correspond to the one or more animate subjects and at least another one of the one or more real world entities correspond to one of the one or more inanimate objects, adding the plurality of overlapping shadows to waiting queue; and selecting, from the priority queue or the waiting queue, the plurality of overlapping shadows, for determining the shadow prominence index, based on a likelihood of natural separation of the plurality of overlapping shadows due to mobility of the one or more inanimate objects associated with the plurality of overlapping shadows.

9. The method (600) as claimed in claim 8, wherein determining the shadow prominence index comprises:for each shadow of the plurality of overlapping shadows:determining an estimated shadow based on the light source information; determining an overlap between the estimated shadow and a corresponding image of a ground truth shadow; anddetermining the shadow prominence index when the overlap is greater than a predefined threshold.

10. The method (600) as claimed in claim 7, wherein determining the reflection prominence index comprises:for each of the one or more real world entities:determining a corresponding image semantic vector from corresponding image of the each of the one or more real world entities;determining a corresponding feature semantic vector associated with corresponding reflection of the each of the one or more real world entities; anddetermining the reflection prominence index based on a cosine similarity between the corresponding image semantic vector and the feature semantic vector.

11. The method (600) as claimed in claim 3, wherein the plurality of attributes for generating the one or more virtual appearances include:at least one relevant object to be virtually presented corresponding to at least one of the one or more real world entities;motion probability of the corresponding at least one of the one or more real world entities; andorientation and dimension probability of the corresponding at least one of the one or more real world entities in a subsequent FOV of the XR device (201).

12. The method (600) as claimed in claim 11, wherein determining the plurality of attributes for generating the one or more virtual appearances comprises:determining the at least one relevant object and a specific virtual representation of the at least one relevant object based on the context vector, and corresponding one or more attributes associated with the orientation and the dimension of the at least one of the shadow and reflection of corresponding one or more real world entities;determining the motion probability associated with a likelihood of movement of the corresponding at least one of the one or more real world entities in the subsequent FOV based on the segmentation information associated with the image of the FOV; anddetermining the orientation and segmentation probability based on the segmentation information associated with the image of the FOV, the determined motion probability associated with the subsequent FOV, and the light source information.

13. The method (600) as claimed in claim 11, wherein generating the one or more virtual appearances comprises:generating the one or more virtual appearances corresponding to the at least one relevant object based on the orientation and segmentation probability in a pre-determined order of the motion probability using a layered conditional generative network.

14. An extended reality (XR) display device (201), for transforming at least one of a shadow and a reflection of one or more real-world entities, XR display device (201) comprises:a memory (205);a processor (207) coupled with the memory (205), the processor (207) being configured to:determine semantic information associated with a field of view (FOV) of the XR display device;determine a plurality of attributes for generating one or more virtual appearances corresponding to the at least one of the shadow and the reflection of the one or more real world entities based on a context vector, the semantic information, and segmentation information associated with an image of the FOV;generate the one or more virtual appearances based on the determined plurality of attributes, the semantic information, and the image of the FOV; andoverlay the one or more virtual appearances on the image of the FOV such that the at least one of the shadow and the reflection of the one or more real-world entities is transformed.

15. The XR display device (201) as claimed in claim 14, wherein the one or more real-world entities correspond to at least one of one or more animate subjects and one or more inanimate objects.