Video see-through device for performing hand tracking and method for operating the same

The video see-through device enhances AR/VR experiences by using an AI model to dynamically adapt to user environments, addressing the lack of contextual adaptation and personalization in current technologies, resulting in improved hand tracking efficiency and user engagement.

WO2026019121A1PCT designated stage Publication Date: 2026-01-22SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/009466
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-07-02
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Current AR/VR technologies lack contextual adaptation and personalization, failing to dynamically improve or adapt to the user's environment and behavior over time, leading to a static user experience and potential frustration.

Method used

A video see-through device and method utilizing an AI model to capture and process environmental data, enabling 3D/2D scene information retrieval, dynamic object detection, and continuous adaptation through machine learning to enhance hand tracking efficiency.

Benefits of technology

The system provides a more immersive and intuitive AR/VR experience by accurately adapting to user surroundings, improving hand tracking efficiency through real-time object detection and continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009466_22012026_PF_FP_ABST
    Figure KR2025009466_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Providing a video see-through (VST) device for performing hand tracking and a method for operating the same. The VST device initiates the process by capturing detailed environmental data using advanced sensors and then utilizes an artificial intelligence (AI) model to render 3D / 2D scene information. The AI model distinguishes hands of a user from static elements through background removal, object possession time analysis, and 3D connectivity estimation. The VST device performs hand tracking by differentiating hands of the user from other objects using trajectory analysis. Continuous updates to the AI model, based on new data and interactions, further refines the object detection and tracking capabilities. The iterative enhancement allows the VST device to adapt and respond more accurately to user movements, delivering a more immersive and intuitive AR / VR experience.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO SEE-THROUGH DEVICE FOR PERFORMING HAND TRACKING AND METHOD FOR OPERATING THE SAME

[0001] The present invention generally relates to the field of virtual reality / augmented reality (VR / AR) and more particularly provides a video see-through (VST) device for performing hand tracking and a method for operating the same. In particular, the present invention relates to providing a VST device and a method for continuously improving the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding.

[0002] The following description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.

[0003] Virtual Reality (VR) and Augmented Reality (AR) have seen significant advancements over the years, which has transformed numerous industries. Improved hardware offers better resolution, field of view, and tracking, making immersive experiences more accessible and realistic. Software development provide robust tools for creating sophisticated VR / AR applications. These technologies are now widely used in gaming, education, healthcare, and industrial training, enhancing user interaction through precise hand and eye-tracking capabilities. In current times, VR / AR hand-tracking technologies have significantly enhanced user interaction with virtual environments. The current devices incorporate advanced hand-tracking capabilities, allowing users to manipulate virtual objects and navigate interfaces without physical controllers. This technology finds applications in education, providing hands-on training simulations; in healthcare, aiding in surgical training and rehabilitation; and in various industrial settings for remote collaboration and complex task simulations. The precise capture of hand movements makes VR / AR experiences more immersive and intuitive, driving greater adoption across different sectors.

[0004] Being an emerging technology, the AR / VR field face several challenges, including high development costs and the need for powerful hardware to deliver seamless experiences. User comfort and safety are also concerns, with issues like motion sickness and prolonged use discomfort being common. Technical limitations, such as insufficient battery life and limited field of view, hinder widespread adoption. Privacy and data security are critical, as these technologies often collect sensitive user information. Lastly, achieving a broad market acceptance is challenging due to the high cost of devices and the need for compelling, practical applications beyond entertainment.

[0005] However, one of the most significant challenges associated with AR / VR is the lack of contextual adaptation and personalization over time, despite continuous use in the same location. However, current AR / VR technologies often fail to improve or adapt dynamically to the user's environment and behaviour. This issue arises because the devices typically do not effectively remember or adapt to the spatial layout of the environment or the user's habitual interactions. Consequently, the user experience remains static and does not evolve to become more intuitive or tailored to the user's preferences and specific settings. This lack of improvement may lead to user frustration and diminish the perceived value and effectiveness of AR / VR technology. Further, users also expect that as they repeatedly use AR / VR devices in environments like their living room or bedroom, the devices should learn and optimize the user experience based on accumulated data and interactions.

[0006] There is therefore a need for a system and method to overcome the challenges associated with the existing technologies and to provide the techniques for advancements in machine learning, better integration of adaptive algorithms, and more sophisticated environmental mapping and user behaviour tracking thereby continuously improving the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding.

[0007] The present disclosure overcomes one or more shortcomings of the prior art and provides additional advantages. Embodiments and aspects of the disclosure described in detail herein are considered a part of the claimed disclosure.

[0008] According to an aspect of the present disclosure, a method, operated by a video see-through (VST) device, for performing hand tracking comprises: obtaining at least one image of a real-world by capturing a user's surrounding using a camera and a field of view (FOV) of the VST device by using at least one sensor while the user interacts with the VST device; obtaining, by using an artificial intelligence (AI) model, at least one 2D view corresponding to the FOV from a 3D model corresponding to the user's surrounding; identifying at least one background object that is commonly included in the at least one image and the obtained at least one 2D view; removing the identified at least one background object from the at least one image; detecting a hand of the user from the at least one image in which the at least one background object has been removed; and performing hand tracking of the hand of the user by applying image processing analysis on the hand of the user.

[0009] According to an aspect of the present disclosure, a video see-through (VST) device for performing hand tracking, the VST device comprises: a camera; at least one sensor; at least one processor including processing circuitry; and memory storing one or more instructions that, when executed by the at least one processor individually or collectively, cause the VST device to: obtain at least one image of a real-world by capturing a user's surrounding using the camera and a field of view (FOV) of the VST device by using the at least one sensor while the user interacts with the VST device, obtain, by using an artificial intelligence (AI) model, the at least one 2D view corresponding to the FOV from a 3D model corresponding to the user's surrounding, identify at least one background object that is commonly included in the at least one image and the obtained at least one 2D view, remove the identified at least one background object from the at least one image, detect a hand of the user from the at least one image in which the at least one background object has been removed, and perform hand tracking of the hand of the user by applying image processing analysis on the hand of the user.

[0010] The foregoing technical solution is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.

[0011] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings in which like reference characters identify correspondingly throughout. Some embodiments of system and / or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying Figs., in which:

[0012] Fig. 1A depicts an exemplary environment illustrating a common scenario of using the AR / VR technology, in accordance with embodiments of the present disclosure;

[0013] Fig. 1B depicts an exemplary environment illustrating example image frame depicting the application of the existing technologies to track the hands of the user, in accordance with embodiments of the present disclosure;

[0014] Fig. 2 depicts a flow diagram illustrating the process of hand tracking using AR / VR technology, in accordance with embodiments of the present disclosure;

[0015] Fig. 3 depicts an exemplary environment illustrating user interacting with AR / VR technology, in accordance with embodiments of the present disclosure;

[0016] Fig. 4 depicts a flow diagram illustrating an overview of the proposed technique of continuously improving the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding, in accordance with embodiments of the present disclosure;

[0017] Fig. 5 depicts a block diagram illustrating detailed mechanism of retrieving the 3D / 2D scene information from the artificial intelligence (AI) model, in accordance with embodiments of the present disclosure;

[0018] Fig. 6 depicts an exemplary surrounding illustrating the virtual camera placement by continuously updating the AI model, in accordance with embodiments of the present disclosure;

[0019] Fig.7 depicts a block diagram illustrating dynamic object detection in an AR / VR environment, leveraging the AI models to reduce clutter, in accordance with embodiments of the present disclosure;

[0020] Figs. 8A-8C depict different environments illustrating the gradual removal of the clutter from background so as to effectively perform hand tracking, in accordance with the embodiment of the present disclosure;

[0021] Fig. 9 depicts a block diagram illustrating, in accordance with the embodiments of the present disclosure;

[0022] Fig. 10 depicts a block diagram illustrating, in accordance with the embodiments of the present disclosure;

[0023] Fig. 11 depicts a block diagram illustrating, in accordance with the embodiments of the present disclosure;

[0024] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in a computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0025] The foregoing has broadly outlined the features and technical advantages of the present disclosure in order that the detailed description of the disclosure that follows may be better understood. It should be appreciated by those skilled in the art that the conception and specific embodiment disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure.

[0026] The novel features which are believed to be characteristic of the disclosure, both as to its organization and method of operation, together with further objects and advantages will be better understood from the following description when considered in connection with the accompanying figures. It is to be expressly understood, however, that each of the figures is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure.

[0027] In current times, Virtual Reality (VR) and Augmented Reality (AR) technologies have made significant strides, transforming various sectors like gaming, education, healthcare, and industrial training. Modern advancements in hardware have led to better resolution, wider fields of view, and improved tracking capabilities, enhancing the realism and accessibility of immersive experiences. Software developments have provided robust tools for creating sophisticated applications, allowing for precise hand and eye tracking that facilitates intuitive interactions without physical controllers. However, despite these innovations, the AR / VR field faces notable challenges. High development costs and the need for powerful hardware remain significant barriers to widespread adoption. User comfort and safety are also concerns, with issues such as motion sickness and discomfort from prolonged use. Technical limitations, including insufficient battery life and limited field of view, further hinder the seamless experience necessary for broader acceptance. Additionally, privacy and data security are critical issues due to the sensitive information these technologies often collect. Out of all these, one of the most pressing challenges is the lack of contextual adaptation and personalization. In current times, the AR / VR devices typically do not improve or adapt dynamically to the user's environment and behaviour over time, leading to a static user experience. This failure to evolve may result in user frustration and reduce the perceived value and effectiveness of AR / VR technology.

[0028] In order to overcome the above-mentioned challenges, the present disclosure provides a method and system for enhancing hand tracking efficiency using the AR / VR device. The disclosed technique may capture the detailed environmental data using advanced sensors. An artificial intelligence (AI) model may be then utilized to render 3D / 2D scene information from the captured data, enabling accurate identification of surroundings. This AI model, along with advanced machine learning techniques, facilitates the detection of dynamic objects in real time by distinguishing them from static elements through processes like background removal, object possession time analysis, and 3D connectivity estimation. Once dynamic objects are identified, the proposed technique may focus on hand tracking by differentiating hands of a user from other non-rigid and rigid objects using trajectory analysis. The gesture identification unit may therefore play a crucial role in recognizing and interpreting user gestures, ensuring precise hand movement capture. Further, continuous updates to the AI model based on new data and user interactions may further refine the system's object detection and tracking capabilities, iteratively improving hand tracking efficiency. This ongoing enhancement process may therefore allow the system to adapt and respond more accurately to user movements, providing a more immersive and intuitive AR / VR experience.

[0029] In the present disclosure, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.

[0030] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the spirit and the scope of the disclosure.

[0031] The terms “comprise”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device, or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a device or system or apparatus proceeded by “comprises… a” does not, without more constraints, preclude the existence of other elements or additional elements in the device or system or apparatus.

[0032] The terms “an embodiment”, “embodiment”, “embodiments”, “the embodiment”, “the embodiments”, “one or more embodiments”, “some embodiments”, and “one embodiment” mean “one or more (but not all) embodiments of the invention(s)” unless expressly specified otherwise.

[0033] The terms “including”, “comprising”, “having” and variations thereof mean “including but not limited to” unless expressly specified otherwise.

[0034] The terms “surrounding”, “space” and “environment” have been used interchangeably in the present disclosure.

[0035] Figure 1A depicts an exemplary environment 100A illustrating a common scenario of using the AR / VR technology. In this, a user 102 has been depicted using some kind of AR / VR device 104. In one embodiment, the AR / VR device 104 may be a Video see-through (VST) device or a Head-Mounted Display (HMD) device, smart glasses, mixed reality enabling devices, AR / VR Internet of Things (IoT) devices among others. Enabling hand tracking via the AR / VR technology has found application in various sectors such as gaming, education, and healthcare by providing more immersive and interactive experiences. In an exemplary scenario, in medical training, surgeons may practice procedures with precise hand movements, while in education, students may engage with virtual lab experiments. To achieve this, the hands 106 of the user 102 may be tracked, which has been illustrated in Fig. 1B of the present disclosure as explained in the upcoming paragraphs.

[0036] Figure 1B depicts an exemplary environment 100B illustrating example image frame depicting the application of the existing technologies to track the hands 106 of the user 102. Various artificial intelligence (AI) models may be deployed to process the data to create a real-time, accurate representation of the hands 106 within the virtual environment. The system may be configured to recognize various gestures such as pointing 108 or open hand 110 among others enabling users to interact naturally with virtual objects. The existing technology to achieve the hand tracking via deploying the AR / VR has been further explained in the upcoming paragraphs in conjunction with the Fig. 2 of the present disclosure.

[0037] Figure 2 depicts a flow diagram 200 illustrating the process of hand tracking using AR / VR technology. In this, the first step may comprise of image frame acquisition 202 as depicted where the images of the user's hands may be captured using sensors such as cameras or depth sensors mounted on AR / VR devices. These images provide visual data that may be processed to detect and track hand movements. This may be followed by the process of hand segmentation 204 which may involve isolating the user's hands from the background and other objects in the captured images. This step is crucial for tracking hand movements without interference from the surrounding environment. Once the hands are segmented, the next step may comprise that of hand tracking 206 which may involve tracking hands movements in subsequent frames. Tracking mechanism may be deployed which aid in analyzing the motion and position of the hands over time, allowing for real-time updates of their location and orientation. Feature extraction 208 may then involve identifying relevant characteristics or key points of the hand of the user, such as finger joints or palm orientation as depicted in Fig. 1B. These features may provide valuable information about the hand's shape, posture, and movement, which may be subsequently used for gesture recognition. After extracting the features, the process of classification 210 may be initiated where various classification mechanisms may be deployed to analyze the extracted features to recognize specific hand gestures or actions. In one embodiment, different machine learning techniques, such as neural networks, may be commonly used to classify hand movements based on learned patterns from training data. Finally, the recognized hand gestures are translated into corresponding output gestures 212 such as commands or actions within the AR / VR application. This output, in one of the embodiments, may include interactions with virtual objects, navigation through menus, or triggering specific events based on the user's gestures.

[0038] As already explained in the foregoing paragraphs that the techniques of hand tracking via the AR / VR techniques being used in the current scenarios, as discussed in reference to Fig. 2 of the present disclosure face most significant challenge of the lack of contextual adaptation and personalization over time, despite continuous use in the same location. However, current AR / VR technologies often fail to improve or adapt dynamically to the user's environment and behaviour. This issue arises because the devices typically do not effectively remember or adapt to the spatial layout of the environment or the user's habitual interactions. Consequently, the user experience remains static and does not evolve to become more intuitive or tailored to the user's preferences and specific settings. This lack of improvement may lead to user frustration and diminish the perceived value and effectiveness of AR / VR technology. To overcome these challenges, the present disclosure offers a novel approach for continuously improving the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding, as discussed in the upcoming paragraphs in conjunction with the Figs. 3-11 of the present disclosure.

[0039] Fig. 3 illustrates an exemplary environment 300 of a user interacting with AR / VR technology. In this exemplary scenario, the surrounding has been depicted as a workstation but it may include any user surrounding such as bedroom, living room, operation theatre, parks, malls among others. Now in this environment 300, the user 302 wearing the AR / VR device 306 may initiate the interactions to attain effective tracking of the hands 304. In this scenario, various background objects 308 such as flowerpot, helmet, water bottle among others are also present which may be captured by the AR / VR device 306 once the user 302 initiates the said interactions. In this given environment 300, the disclosed technique of hand tracking has been explained from an overview in the upcoming paragraphs in conjunction with Fig. 4.

[0040] Fig. 4 depicts a flow diagram 400 illustrating an overview of the proposed technique of continuously improving the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding. In this, at step 402, a camera has been depicted which may be linked to the AR / VR device. In one embodiment, the camera may or may not constitute as one of the principal components of the AR / VR device in use. The camera may then provide real-world input frames to the hand tracking system which may basically depict the surrounding of the user as visible to them. Once the surrounding of the user interacting with the AR / VR device has been identified by the input camera frames then, at step 404, a three-dimensional (3D) scene information / two-dimensional (2D) scene information corresponding to the identified surrounding may be retrieved from an artificial intelligence (AI) model 404. In one of the scenarios, when the AR / VR device is being used again in the same environment 300 then the retrieved 3D / 2D scene information from the AI model 404 may be associated with the previously captured frames by the AR / VR device when it was last used in the same environment 300. Whereas in another scenario, if the AR / VR device is being used for the first time in the environment 300 i.e., the environment 300 is new for the AR / VR device then the AI model 404 may not have any associated 3D / 2D scene information corresponding to the new environment 300 and thus the input frames acquired by the camera 402 in the first session may be used for training the 3D model and facilitating the corresponding surrounding’s retrieval in all the future sessions of the concerned surrounding.

[0041] The 3D / 2D scene information of the user surrounding (for example, environment 300 of Fig. 3) is learned or stored by the AR / VR device during previous usage of the device in the same surrounding. In an embodiment, the retrieved 3D / 2D scene information from the AI model 404 may or may not constitute all the objects as present in the real-world captured input camera frame. For example, referring back to Fig. 3, if the AI model 404 has an image of the environment 300 stored in its database such that the stored image does not have water bottle, mobile device and sunglasses in it then that would be the partial retrieval of the real-world image. In other words, the retrieved images from the AI model are similar to the identified surrounding only and may or may not constitute all the objects present in its real-world counterpart. Coming back to Fig. 4, after retrieving the 3D / 2D scene information of the identified surrounding, at step 406, detection of dynamic objects in the surroundings is initiated. This detection may be enabled by an elaborative mechanism which have been discussed in detail in later section of this disclosure. Once the dynamic objects i.e., user hands are detected, then at step 408, the user hands are tracked effectively. In one embodiment, the hand tracking may be done by removing all the static objects and only keeping the detected dynamic objects in the rendered frame such that at step 410, tracked and rendered hands may be displayed. In another embodiment, both the static and dynamic objects may comprise of the one or more first objects whereas the one or more second objects may only comprise of the dynamic objects.

[0042] Fig. 5 depicts a block diagram 500 illustrating detailed mechanism of retrieving the 3D / 2D scene information from the AI model, in accordance with the embodiments of the present disclosure. In this, sensors 502 has been depicted which may be used by the AR / VR devices for capturing data regarding the surrounding environment. In one embodiment, these sensors may comprise cameras, LiDAR, infrared sensors, GPS (Global Positioning System), IMU (Inertial Measurement Unit) among others, to capture comprehensive data about the surrounding environment. These sensors 502 may also be used for collecting information as they gather visual and depth information, including the spatial layout, object positions, and ambient conditions, providing a detailed representation of the physical space. This collected information may be then fed to an AI model database 504. In an embodiment, the AI model may constitute of neural radiance field (NeRF) which may facilitate in enabling learning of novel view synthesis, scene geometry, and the reflectance properties of any given scene. In the present disclosure, the "NeRF" is a neural network that may reconstruct complex 3D scenes from a partial set of 2D images. In an embodiment, the NeRF may take 2D images representing a scene as input and interpolates between them to render a complete 3D scene.

[0043] This an AI model database 504 may contain extensive datasets used to train the AI models. These datasets may include labeled images and environmental data that represent various types of surroundings and objects. This data is essential for the initial training of the AI models, ensuring they may recognize and understand different environments accurately. Therefore, when the AR / VR device needs to identify a new environment, it may retrieve 506 a AI model 508 from the AI model database 504. In one embodiment, these models may be trained on large datasets to recognize common environmental features and objects. The AI model 508 may then use the data captured by the sensors to identify and map the environment in real-time. It may recognize patterns, objects, and spatial relationships based on its training. The AI model 508 may generate a detailed map of the surroundings, identifying key elements such as walls, furniture, and other objects. It may also adapt and refine its understanding of the environment over time, incorporating new data to improve accuracy and robustness. In another embodiment, the AI model 508 may be loaded onto the AR / VR device, ready to process incoming data from the sensor 502. Therefore, by combining data captured by the sensor 502 with sophisticated AI models, AR / VR systems may effectively identify and map surroundings, providing users with an immersive and interactive experience that aligns closely with the physical world.

[0044] Along with the process as explained in the above paragraph, data captured by the sensor 502 may also be utilized for determining viewing location and angle 510. In this, the AR / VR device may determine the user's current position (viewing location) and orientation (viewing angle) within the identified physical space. In one embodiment, along with the tracking of the user’s hand, this may also involve tracking the user’s head movements and body position. This information may be therefore crucial for calculating the correct perspective from which the virtual environment should be rendered, ensuring it aligns with the user's actual view. In particular, the proposed AI model may be configured in such a way that it may give the list of objects in the user’s surrounding along with their identified locations from the queried viewpoint i.e., the viewpoint of the user using the AR / VR device.

[0045] Based on the user's viewing location and angle 510, virtual camera 512 may be placed within the AR / VR environment by using the AI model 508. The virtual camera 512 may mimic the user's eyes, capturing the scene from the same viewpoint as the user would see it in the real world. Further, as the user moves, the virtual camera 512 may continuously adjust its position and orientation to match the user's movements, maintaining an accurate and consistent viewpoint. This phenomenon of placing the virtual cameras has been further explained in the later part of this disclosure in reference to Fig. 6. The virtual camera 512 may capture the 3D virtual environment from its viewpoint, generating a 2D image that may represent the user's view. In one embodiment, this may involve complex rendering techniques to ensure realistic lighting, textures, and depth. In another embodiment, in AR, this generated 2D image may be overlaid onto the real-world view seen through the AR / VR device's display, seamlessly integrating virtual objects with the physical surroundings. However, in VR, the 2D image may fully immerse the user in the virtual environment, replacing their view of the real world. This process may be happening in real-time, with the system continuously rendering new 2D images as the user's viewpoint changes, providing a smooth and immersive experience. Therefore, by using sensors 502 to capture environmental data and determining the user's viewing location and viewing angle 510, the AR / VR systems may accurately place a virtual camera 512 to simulate the user's viewpoint. The virtual camera 512 may then render 2D images that are either integrated with the real world in AR or fully immersive in VR, creating a seamless and dynamic user experience.

[0046] Subsequent to this, the 3D / 2D rendered information may be then combined with the inputs from camera 516 of the AR / VR device to accomplish dynamic object detection 518. Once the dynamic objects i.e., user hands are detected, then the hand tracking 520 may be accomplished effectively, after which, the tracked and rendered hands may be displayed.

[0047] In addition to this, the block diagram 500 also depicts the process of enabling continuous improvement in the hand tracking efficiency by facilitating constant adaptation of the VR / AR to any given user surrounding. In this, the rendered 3D / 2D scene information may be combined with the retrieved input frames from the camera 516 for object recognition 522 to continuously update 524 the AI model. In particular, a feedback loop may be established, where new data from the real-world input camera frames continuously refine the AI model's 508 understanding through incremental updates. Adaptive models may integrate the newly acquired information, adjusting recognition parameters to remain accurate and responsive to changes. This continuous learning process may therefore correct errors and enhance predictions, leading to improved accuracy and adaptability. Benefits include better object recognition accuracy, the ability to adapt to dynamic environments, enhanced user experiences, personalized interactions, and maintaining relevance as new objects emerge in a given surrounding. This real-time adaptation therefore ensures a more reliable and immersive AR / VR experience for users.

[0048] As referred in the foregoing paragraphs, Fig. 6 depicts an exemplary surrounding 600 illustrating the placing of the virtual cameras 602 by continuously updating the AI model, which may be NeRF model, in one of the embodiments. In this, the virtual cameras 602 which are generally associated with the AR / VR device may simulate the user's viewpoint within the virtual or augmented environment, capturing and rendering scenes in real-time to match the user’ viewpoint. This ensures that the virtual elements are properly aligned with the physical world in AR or create an immersive experience in VR. To elaborate on this, the disclosed mechanism proposes that each time a user enters and interacts with the room via the AR / VR device, the AR / VR device may capture multiple viewpoints through its sensors. As the user moves around, different viewpoints of the room are recorded, providing a comprehensive set of data points from various angles. Over time, these accumulated data points may create a more detailed and extensive dataset of the room. Each new session initiated by the user may add more viewpoints, effectively increasing the number of virtual cameras 602 or viewpoints from which the environment has been observed and recorded. Now with each additional data point collected, the system may therefore place more virtual cameras within the AI model and are used to update the AI model through offline training. These virtual cameras 602 may therefore represent the different viewpoints from which the room has been observed, allowing the system to render the room from multiple angles with greater detail and accuracy as depicted via Figs. 604-610. As the AI model is trained with more data, it may become better at synthesizing new viewpoints and rendering the environment more accurately. The increase in virtual camera placements (new viewpoints) may therefore allow the AI model to interpolate and extrapolate more effectively, producing high-quality renders from virtually any angle as depicted by Fig 6. This may, in addition, lead to improved AI model, enriched with more data points, which may thereby enhance the system's ability to detect and track dynamic objects and / or hand movements accurately.

[0049] Fig.7, in turn depicts a block diagram 700 illustrating dynamic object detection 702 in an AR / VR environment, leveraging the AI models to reduce clutter. In one embodiment, the clutter may comprise of the background information. Upon receiving the input camera frames and the rendered 3D / 2D scene information from the AI model, the two scenes may be compared. In one embodiment, at step 704, the AI model may be queried for previously encountered objects i.e., those objects which are included in both of the received scene information and their positions from various viewpoints. Once identified, such static objects from the AI model may be removed from the scene to reduce visual clutter, allowing the system to focus on detecting dynamic objects such as hands.

[0050] Referring to Figs. 8A-8C of the present disclosure which depict different environments 800A-800C of gradually removing the clutter from background so as to effectively perform hand tracking, in accordance with the embodiment of the present disclosure. In Fig. 8(a), let us suppose helmet 802 is the identified common or previously encountered static object then after performing the step 704 of Fig. 7, the helmet 802 may be removed as depicted in Fig. 8(b) (shown using dotted lines).

[0051] Referring back to Fig. 7, at step 706, the process of 3D object connectivity estimation may be carried. This estimation may help in determining that how the objects interact with each other in 3D. Typically, hands perform air interactions or interaction with a single objects like mobile device, hence, they may be connected to a lesser number of objects and all the objects with higher connectivity are also removed to further reduce the clutter. In view of Fig. 8(c), the system 804 may be removed as it has higher connectivity than other objects in the identified surrounding.

[0052] Then at step 708, in Fig. 7, the process of object possession estimation may be carried. In this, the objects may be analyzed based on their possession of the Field of View (FoV) or screen time. Hands and body parts typically occupy the FoV for longer durations compared to other objects in atypical AR / VR environment. So, based on the possession time, the objects with less possession time are removed, narrowing down the potential dynamic objects to those likely to be hands or parts of the body. Taking cue from Fig. 8(d), other object having less possession time such as mobile device 806 may be removed. In the last, at step 710, rigid / non-rigid object detection may be carried out such that in this module, the rigid objects may be identified using temporal frames. First of all, the trajectory of each object may be identified where rigid objects have consistent and predictable trajectories and may be therefore removed from the scene. As evident from Fig. 8(f), all objects 802-810 may be removed yielding Fig. 8(e) with only tracked and rendered hands 812.

[0053] Fig. 9 depicts a system diagram 900 illustrating an AR / VR device 902. In an embodiment, the AR / VR device 902 may be a video see-through (VST) device. However, the AR / VR device 902 is not limited thereto, and may be implemented as a head mounted display (HMD) device, smart glasses, mixed reality enabling devices, AR / VR Internet of Things (IoT) devices among others. Referring to Fig. 9, the AR / VR device 902 may comprise a sensor 904, an input / output (I / O) interface 906, a processor 908, a display 910, a communication interface 912, and memory 914. The camera 904, the I / O interface 906, the processor 908, the display 910, the communication interface 912, and the memory 914 may be electrically and / or physically connected to each other. Fig. 9 illustrates only essential components for describing operations of the AR / VR device 902, and the components included in the AR / VR device 902 are not limited to those shown in Fig. 9.

[0054] The sensor 904 may comprise at least one camera which may be configured to obtain at least one image of a real-world corresponding to the user’s surrounding. In an embodiment, sensor 902 may further comprise of LiDAR, and depth sensors, GPS, IMU etc. which continuously capture real-time data from the environment. In another embodiment, the collected data may include images, depth information, and spatial coordinates of objects. Further, basic preprocessing, such as noise reduction and initial object detection, may occur at the sensor level to prepare the data for further analysis.

[0055] The I / O interface 906 may be configured to receive the collected data by the sensor 904 as input frames, such as gestures. In an embodiment, the I / O interface 906 may further comprises a microphone which configured to receive voice commands. This I / O interface 906 may provide an effective user interface for interaction with the AR / VR system.

[0056] The processor 908 may execute a program code or one or more instructions stored in the memory 914. The processor 908 may include one or a plurality of processors. The one or the plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processor such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The processor 908 may include multiple cores and is configured to execute the instructions stored in a memory 914.

[0057] The processor 908 according to one or more embodiments of the disclosure may include various processing circuitry and / or multiple processors. For example, as used in the present disclosure, including the claims, the term “processor” may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when “a processor”, “at least one processor”, and “one or more processors” are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the processor 908 may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner. The processor 908 may execute program instructions stored in the at least one memory 914 to achieve or perform various functions.

[0058] The memory 914 may store one or more instructions to be executed by the processor 908. The memory 914 may include one or more non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard disks, optical disks, floppy disks, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the at memory 914 may, in some examples, be considered a non-transitory storage medium. The term “non-transitory” may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term “non-transitory” should not be interpreted that the memory 914 is non-movable. In certain examples, a non-transitory storage medium may store data that may, over time, change (e.g., in Random Access Memory (RAM) or cache).

[0059] The processor 908 may then integrate sensor data and user inputs and may be then configured to apply the AI models from the memory 914. In an embodiment, the memory 914 may comprise an AI model 916, tracking code 918, and a gesture identification code 920. In the present disclosure, the term ‘code’ included in the memory 914 refers to a unit for processing functions or operations performed by the processor 908, and may be implemented as software such as instructions, algorithms, data structures, or program code.

[0060] In an embodiment, the AI model 916 may be retrieved by the processor 908 from the memory 906. In an embodiment, the AI model 916 may constitute of neural radiance field (NeRF) In the present disclosure, the “NeRF” is a neural network that may reconstruct complex 3D scenes from a partial set of 2D images. In an embodiment, the NeRF may take 2D images representing a scene as input and interpolates between them to render a complete 3D scene. The AI model 916 may render 3D / 2D scene information corresponding to the identified surroundings captured via the input image frames from the sensor 904. In an embodiment, the processor 908 in conjunction with the AI model 916 may be configured to retrieve one or more 2D views corresponding to the FOV of the VST device while the user interacts with the VST device. The ‘one or more 2D views corresponding to the FOV’ may refer to one or more 2D images stored in 3D model corresponding to the identified surroundings by FOV of the VST device rendered using the AI model 916. The processor 908 in conjunction with the AI model 916 may be configured to correlate the at least one image obtained by the sensor 904 with the retrieved one or more 2D views rendered by using the AI model 916 corresponding to the user’s surrounding.

[0061] The processor 908 in conjunction with the AI model 916 may then carry out the process of detecting the dynamic objects in the identified surrounding in real-time. In an embodiment, the processor 908 may identify at least one static object that is included in both of the real-world corresponding to the FOV of the VST device and the retrieved one or more 2D views corresponding to the FOV of the VST device by comparing the image frames from the sensor 904 with the one or more 2D views.

[0062] The processor 908 may be configured to remove the background information from the identified surrounding such as removing previously encountered objects, objects with low possession time or objects with high connectivity. The processor 908 may then carry out rigid / non-rigid object detection to determine the dynamic object in the real-world. In an embodiment, the processor 908 may detect the user’s hand from the image frame corresponding to the identified surrounding in which the background information has been removed. Once the dynamic objects, i.e., user’s hand are detected then the processor 908 may be configured to track objects and user movements across frames by executing the tracking code 918. Subsequently, the processor 908 may be configured to recognize and interpret user gestures based on sensor data and predefined gesture models by executing the gesture identification code 920. In an embodiment, the processor 908 may perform hand tracking of the user’s hand by applying image processing analysis on the user’s hand, by executing the tracking code 918.

[0063] The display 910 may be configured to render the updated AR / VR environment, integrating dynamically detected objects and user interactions. The display 910 may be implemented by, for example, at least one of a liquid-crystal display (LCD), a thin-film-transistor liquid-crystal display (TFT-LCD), an organic light-emitting diode (OLED) display, a flexible display, a three-dimensional (3D) display, or an electrophoretic display.

[0064] However, the present disclosure is not limited thereto. In a case in which the AR / VR device 902 is implemented as augmented reality glasses, the display 910 may be configured as a lens optical system and may include a waveguide and an optical engine. The optical engine may include a projector configured to generate light of a three-dimensional virtual object configured as a virtual image, and project the light to the waveguide. The optical engine may include, for example, an image panel imaging panel, an illumination optical system, a projection optical system, and the like. In an embodiment of the disclosure, the optical engine may be arranged in the frame or temples of the augmented reality glasses. In an embodiment of the disclosure, the optical engine may display the virtual bounding region or the at least one prompt by projecting, to the waveguide, light of the virtual bounding region or the at least one prompt for providing an image to the user, under control of the processor 908.

[0065] The communication interface 912 may ensure smooth data flow between various hardware components of the system such as sensor 904, I / O interface 906, processor 908, memory 914 and display 910. In addition to this, the AI model 916 may be continuously updated based on new sensor data and user inputs, refining object detection and tracking dynamically, where this process of updating the AI model 916 may be iteratively performed until the three-dimensional (3D) scene information is retrieved entirely for the identified surrounding.

[0066] Fig. 10 is a flowchart showing steps of a method 1000 for enhancing hand tracking efficiency in a Video see-through (VST) device, in accordance with an embodiment of the present disclosure. The method 1000 may also be described in the general context of computer executable instructions. Generally, computer executable instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types.

[0067] The order in which the method 1000 is described is not intended to be construed as a limitation, and any number of the described method blocks may be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described.

[0068] At step 1002, the method 1000 may include identifying a surrounding in a real world while a user starts interacting with the VST device. In one embodiment, the sensor 904 (referring to Fig. 9) may be configured to capture the input frames and send it to the processor 908 (referring to Fig. 9), which in turn, in conjunction with the memory 914 (referring to Fig. 9)may be configured to identify the surrounding in which the user is interacting with the help of the AR / VR device.

[0069] At step 1004, the method 1000 may include retrieving a three-dimensional (3D) scene information corresponding to a portion of the identified surrounding using a AI model. In one embodiment, the processing in conjunction with the AI model may be configured to retrieve the 3D scene information corresponding to the identified surrounding. In an embodiment, the AI model may constitute of neural radiance field (NeRF) which may facilitate in enabling learning of novel view synthesis, scene geometry, and the reflectance properties of any given scene.

[0070] At step 1006, the method 1000 may include locating one or more first objects included in both of the real world and the 3D scene information, and one or more second objects included only in the real world. In an embodiment, the one or more first object may be both static and dynamic objects whereas the one or more second object may comprise dynamic objects. In another embodiment, the processor 908 in conjunction with the AI model may be configured to locate the objects and determine their positions which may also be determined by the viewing angle of the user.

[0071] At step 1008, the method 1000 may include performing hand tracking of user’s hand by applying image processing analysis upon the one or more second objects. In one embodiment, the processor 908 in conjunction with the AI model and tracking code 918 (referring to Fig. 9) may be configured to perform hand tracking of the user’s hand only on the detected dynamic objects.

[0072] Fig. 11 is a flowchart showing steps of a method 1100 for enhancing hand tracking efficiency in a video see-through (VST) device, in accordance with an embodiment of the present disclosure. The method 1100 may also be described in the general context of computer executable instructions. Generally, computer executable instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types.

[0073] The order in which the method 1100 is described is not intended to be construed as a limitation, and any number of the described method blocks may be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described.

[0074] At step 1102, the method 1100 may include obtaining one or more images of the real world corresponding to a current user surrounding and a current field of view (FOV) of a user by using at least one sensor while the current user interacts with the VST device. In one embodiment, the processor 908 (referring to Fig. 9) may be configured to obtain the one or more images of the real world corresponding to a current user surrounding by using a camera, and the current field of view (FOV) of the user by using at least one sensor 904 (referring to Fig. 9) during the HMD session.

[0075] At step 1104, the method 1100 may include obtaining, by using the AI model, one or more 2D views corresponding to the current FOV from a 3D model corresponding to the current user surrounding. In an embodiment, the processor 908 in conjunction with the AI model may be configured to retrieve one or more 2D views corresponding to the current FOV of the VST device. In an embodiment, the AI model may constitute of neural radiance field (NeRF) In the present disclosure, the “NeRF” is a neural network that may reconstruct complex 3D scenes from a partial set of 2D images. In an embodiment, the NeRF may take 2D images representing a scene as input and interpolates between them to render a complete 3D scene.

[0076] The processor 908 may further configured to obtain viewing location and viewing angle of the user by using at least one sensor 904, and place a virtual camera in the 3D model using the AI model based on the obtained viewing location and viewing angle of the user. The processor 908 may render the at least one 2D view that represents a viewpoint of the user by using the virtual camera.

[0077] The processor 908 may further configured to detect one or more static objects and location information of the one or more static objects from the one or more 2D view rendered by using the virtual camera.

[0078] At step 1106, the method 1100 may include the identifying one or more background objects that are commonly included in at least one image and the retrieved one or more 2D views corresponding to the current FOV of the VST device. In another embodiment, the processor 908 in conjunction with the AI model may be configured to identify one or more background objects between the current FOV of the VST device and the retrieved one or more 2D views corresponding to the current FOV of the VST device. In another embodiment, the processor 908 in conjunction with the AI model may be configured to correlate the one or more images obtained by capturing the user’s surroundings and the FOV of the VST device with the retrieved one or more 2D views derived using the pre-trained 3D model corresponding to the current user surrounding. In an embodiment, the processor 908 may further configured to detect one or more dynamic objects by comparing the one or more images and the one or more 2D views and by using object information of the one or more static objects and location information of the one or more static objects.

[0079] At step 1108, the method 1100 may include removing the identified one or more background objects from the one or more images corresponding to the current FOV. In an embodiment, the processor 908 may determine possession time of one or more background objects and remove at least one object from among the one or more background objects with the possession time less than a pre-determined threshold. The possession time may constitute screen time occupied by each of the one or more background objects. In an embodiment, the processor 908 may determine connectivity of one or more background objects with each other while analyzing their interactions in the 3D model and remove at least one object with connectivity higher than a pre-defined value from among the one or more background objects. In an embodiment, the processor 908 may identify at least one rigid object by analyzing trajectory of one or more background objects from temporal image frames obtained by using the camera and remove the identified at least one rigid object.

[0080] At step 1110, the method 1100 may include detecting the user’s hand from the one or more images. In an embodiment, the processor 908 may detect the user’s hand from the one or more images in which the one or more background objects have been removed.

[0081] At step 1112, the method 1100 may include performing hand tracking of the user’s hand by applying image processing analysis on the one or more dynamic object i.e., the user’s hand. In another embodiment, the processor 908 in conjunction with the AI model may be configured to perform hand tracking on the one or more images of the real-world in which the one or more background objects have been removed.

[0082] The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.

[0083] Embodiments herein disclose a method operated by a video-see-through (VST) device. In an embodiment, the method may comprise: identifying a surrounding in a real world while a user starts interacting with the VST device; retrieving a three-dimensional (3D) scene information corresponding to a portion of the identified surrounding using a pre-trained learning model; locating one or more first objects present both in the real world and the 3D scene information, and one or more second objects present only in the real world; and performing hand tracking of user’s hand by applying image processing analysis upon the one or more second objects.

[0084] In an embodiment, after performing the hand tracking of the user’s hand, the method may further comprise: updating the 3D scene information along with location of the one or more first objects and location of the one or more second objects; and iteratively performing steps (a)-(e) until the three-dimensional (3D) scene information is retrieved entirely for the identified surrounding.

[0085] In an embodiment, the pre-trained learning model may be trained by: capturing 2D images of the one or more similar surroundings using at least one of sensors and cameras mounted on the VST device; extracting depth information from the captured 2D images; and rendering a continuous 3D representation of the one or more similar surroundings based on the extracted depth information and the captured 2D images.

[0086] In an embodiment, the method may further comprise: generating the 3D scene information of the identified location using the pre-trained learning model at each instance when the user's hand starts interacting with the VST device in one or more surroundings similar to the identified surrounding; and removing background information from the 3D scene information using the pre-trained learning model, wherein the background information corresponds to common objects located in the captured 2D images and the generated 3D scene information.

[0087] In an embodiment, the one or more first objects may be both static and dynamic objects and the one or more second objects may be dynamic objects.

[0088] In an embodiment, the method may further comprise: receiving the captured 2D images of the one or more similar surroundings; receiving the 3D scene information along with location of the one or more first objects and location of the one or more second objects; comparing the captured 2D image and the 3D scene information and removing commonly occurring the one or more first objects and the one or more second objects; determining possession time of remaining objects, wherein the possession time constitutes screen time occupied by each of the remaining objects; removing one or more objects of the remaining objects with the possession time less than a pre-determined threshold; determining connectivity of objects with each other while analysing their interactions in the received 3D scene information; removing the objects with connectivity higher than a pre-defined value; and identifying and removing the static objects in the received 3D scene information to entirely remove the background information from the 3D scene information.

[0089] Embodiments herein disclose a method operated by a video-see-through (VST) device. In an embodiment, the method may comprise: obtaining at least one image of a real-world by capturing a user’s surrounding using a camera and a field of view (FOV) of the VST device by using at least one sensor while the user interacts with the VST device; obtaining, by using an artificial intelligence (AI) model, at least one 2D view corresponding to the FOV of the VST device from a 3D model corresponding to the user’s surrounding; identifying at least one background object that is commonly included in the at least one image and the obtained at least one 2D view; removing the identified at least one background object from the at least one image; detecting a hand of the user from the at least one image in which the at least one background object has been removed; and performing hand tracking of the hand of the user by applying image processing analysis on the hand of the user.

[0090] In an embodiment, the retrieving of the at least one 2D view corresponding to the FOV may comprise: correlating the obtained at least one image with the retrieved at least one 2D view rendered by using the AI model corresponding to the user’s surrounding.

[0091] In an embodiment, the retrieving of the at least one 2D view corresponding to the FOV may comprise: obtaining viewing location and viewing angle of the user by using the at least one sensor, wherein the at least one sensor comprises at least one of global positioning system (GPS), inertial measurement unit (IMU), light wave detection and ranging (LIDAR), or infrared sensor; placing a virtual camera in the 3D model using the AI model based on the obtained viewing location and viewing angle of the user; and rendering the at least one 2D view that represents a viewpoint of the user by using the virtual camera.

[0092] In an embodiment, the removing of the at least one background object from the at least one image may further comprise removing at least one static object from the at least one image of the real-world.

[0093] In an embodiment, the removing of the at least one background object from the at least one image may further comprise: determining possession time of remaining objects, wherein the possession time constitutes screen time occupied by each of the remaining objects; and removing at least one object from among the remaining objects with the possession time less than a pre-determined threshold.

[0094] In an embodiment, the removing of the at least one background object from the at least one image may further comprise: determining connectivity of remaining objects with each other while analyzing their interactions in the 3D model; and removing at least one object with connectivity higher than a pre-defined value.

[0095] In an embodiment, the removing of the at least one background object from the at least one image may further comprise: identifying at least one rigid object by analysing trajectory of remaining objects from temporal image frames obtained by using the camera; and removing the identified at least one rigid object.

[0096] In an embodiment, after performing the hand tracking of the user’s hand, the method may further comprise: updating the 3D model along with location of the at least one background object and location of the at least one dynamic object; and iteratively performing steps of obtaining of the at least one image of a real-world and the FOV of the user, retrieving of the at least one 2D view corresponding to the FOV of the user, identifying of the at least one background object and the at least one dynamic object, removing of the at least one background object from the at least one image of the real-world, performing of the hand tracking, and the updating of the 3D model, until the 3D model is retrieved entirely for the user’s surrounding.

[0097] Embodiments herein disclose a video see-through (VST) device for performing hand tracking. In an embodiment, the VST device may comprise: at least one processor; memory, wherein the at least one processor in conjunction with the memory may be configured to: identify a surrounding in a real world while a user starts interacting with the VST device, retrieve, by using a pre-trained learning model, a three-dimensional (3D) scene information corresponding to a portion of the identified surrounding, locate one or more first objects present both in the real world and the 3D scene information, and one or more second objects present only in the real world, and perform hand tracking of user’s hand by image processing analysis upon the one or more second objects.

[0098] In an embodiment, the at least one processor in conjunction with the memory may be further configured to: update the 3D scene information along with location of the one or more first objects and location of the one or more second objects, and iteratively perform steps (a)-(e) until the three-dimensional (3D) scene information is retrieved entirely for the identified surrounding.

[0099] In an embodiment, the at least one processor in conjunction with the memory may be configured to: capture 2D images of the one or more similar surroundings using at least one of sensors and cameras mounted on the VST device, extract depth information from the captured 2D images, and render a continuous 3D representation of the one or more similar surroundings based on the extracted depth information and the captured 2D images.

[0100] In an embodiment, the at least one processor in conjunction with the memory may be configured to: generate the 3D scene information of the identified location using the pre-trained learning model at each instance when the user's hand starts interacting with the VST device in one or more surroundings similar to the identified surrounding, and remove background information from the 3D scene information using the pre-trained learning model. In an embodiment, the background information may correspond to common objects located in the captured 2D images and the generated 3D scene information.

[0101] In an embodiment, the one or more first objects may be both static and dynamic objects and the one or more second objects may be dynamic objects.

[0102] In an embodiment, the at least one processor in conjunction with the memory may be configured to: receive the captured 2D images of the one or more similar surroundings, receive the 3D scene information along with location of the one or more first objects and location of the one or more second objects, compare the captured 2D image and the 3D scene information and removing commonly occurring the one or more first objects and the one or more second objects, determine possession time of remaining objects, wherein the possession time constitutes screen time occupied by each of the remaining objects, remove one or more objects of the remaining objects with the possession time less than a pre-determined threshold, determine connectivity of objects with each other while analysing their interactions in the received 3D scene information, remove the objects with connectivity higher than a pre-defined value, and identify and remove the static objects in the received 3D scene information to entirely remove the background information from the 3D scene information.

[0103] Embodiments herein disclose a video see-through (VST) device for performing hand tracking. In an embodiment, the VST device may comprise: a camera; at least one sensor; at least one processor including processing circuitry; and memory storing one or more instructions that, when executed by the at least one processor individually or collectively, cause the VST device to: obtain at least one image of a real-world by capturing a user’s surrounding using the camera and a field of view (FOV) of the VST device by using the at least one sensor while the user interacts with the VST device, obtain, by using an artificial intelligence (AI) model, the at least one 2D view corresponding to the FOV of the VST device from a 3D model corresponding to the user’s surrounding, identify at least one background object that is commonly included in the at least one image and the obtained at least one 2D view, removing the identified at least one background object from the at least one image, detect a hand of the user from the at least one image in which the at least one background object has been removed, and perform hand tracking of the hand of the user by applying image processing analysis on the hand of the user.

[0104] In an embodiment, the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to: correlate the obtained at least one image with the retrieved at least one 2D view rendered by using the AI model corresponding to the user’s surrounding.

[0105] In an embodiment, the at least one sensor including at least one of global positioning system (GPS), inertial measurement unit (IMU), light wave detection and ranging (LIDAR), or infrared sensor, and configured to obtain viewing location and viewing angle of the user. In an embodiment, the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to: place a virtual camera in the 3D model using the AI model based on the obtained viewing location and viewing angle of the user, and render the at least one 2D view that represents a viewpoint of the user by using the virtual camera.

[0106] In an embodiment, the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to: remove at least one static object from the at least one image.

[0107] In an embodiment, the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to: determine possession time of at least one background object, wherein the possession time constitutes screen time occupied by each of the at least one background object; and remove at least one object from among the at least one background object with the possession time less than a pre-determined threshold.

[0108] In an embodiment, the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to: determine connectivity of at least one background object with each other while analyzing interactions between each of the at least one background object in the 3D model; and remove at least one object with connectivity higher than a pre-defined value.

[0109] Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Further, any skilled person in the art would appreciate that the reconstruction error mentioned in the foregoing paragraphs may be considered as a value that overshoots the determined threshold value and must not be construed as an error as such.

[0110] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer- readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., are non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0111] Suitable processors include, by way of example, a general-purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a graphic processing unit (GPU), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) circuits, any other type of integrated circuit (IC), and / or a state machine.

[0112] In an embodiment, the present disclosure provides techniques for continuously improving hand tracking in AR / VR thereby enhancing user interaction by making movements more precise and responsive. This may also lead to a more immersive and intuitive experience, reducing errors, and enabling more complex and natural interactions, thereby increasing the overall effectiveness and appeal of AR / VR applications.

[0113] In an embodiment, the present disclosure provides techniques for obtaining known background objects’ information using the AI model of identified / familiar locations in AR / VR may enhance scene understanding and contextual accuracy. It may further allow the system to differentiate between static and dynamic elements, reducing computational load and clutter, thereby improving object detection, interaction accuracy, and overall user experience.

[0114] In an embodiment, the present disclosure provides technique for performing image analysis only on dynamic objects (one or more second objects) in AR / VR may increase the efficiency by reducing computational load and focusing resources on relevant elements. This approach may further enhance real-time performance, minimize latency, and improve the accuracy of object interactions, leading to a smoother and more responsive user experience.

[0115] In an embodiment, the present disclosure provides techniques for distinguishing when a new object may be introduced in the surrounding, alongside differentiating hands or dynamic objects from new static objects, thereby enhancing interaction precision and scene awareness. This may further ensure that the system may accurately track and respond to dynamic changes, improving user experience and enabling more complex and realistic interactions in AR / VR environments.

Claims

1.A method, operated by a video see-through (VST) device, for performing hand tracking, the method comprising:obtaining at least one image of a real-world by capturing a user’s surrounding using a camera and a field of view (FOV) of the VST device, by using at least one sensor while the user interacts with the VST device;obtaining, by using an artificial intelligence (AI) model, at least one 2D view corresponding to the FOV from a 3D model corresponding to the user’s surrounding;identifying at least one background object that is commonly included in the at least one image and the obtained at least one 2D view;removing the identified at least one background object from the at least one image;detecting a hand of the user from the at least one image in which the at least one background object has been removed; andperforming hand tracking of the hand of the user by applying image processing analysis on the hand of the user.2.The method of claim 1, wherein the obtaining of the at least one 2D view corresponding to the FOV comprises:correlating the obtained at least one image with the obtained at least one 2D view rendered by using the AI model corresponding to the user’s surrounding.3.The method of claim 1, wherein the obtaining of the at least one 2D view corresponding to the FOV comprises:obtaining viewing location and viewing angle of the user by using the at least one sensor, wherein the at least one sensor comprises at least one of global positioning system (GPS), inertial measurement unit (IMU), light wave detection and ranging (LIDAR), or infrared sensor;placing a virtual camera in the 3D model using the AI model based on the obtained viewing location and viewing angle of the user; andrendering the at least one 2D view that represents a viewpoint of the user by using the virtual camera.4.The method of claim 1, wherein the removing of the at least one background object from the at least one image further comprises removing at least one static object from the at least one image.5.The method of claim 1, wherein the removing of the at least one background object from the at least one image further comprises:determining possession time of at least one background object, wherein the possession time constitutes screen time occupied by each of the at least one background object; andremoving at least one object from among the at least one background object with the possession time less than a pre-determined threshold.6.The method of claim 1, wherein the removing of the at least one background object from the at least one image further comprises:determining connectivity of at least one background object with each other while analyzing interactions of the user between each of the at least one background object in the 3D model; andremoving at least one object with connectivity higher than a pre-defined value.7.The method of claim 1, wherein the removing of the at least one background object from the at least one image further comprises:identifying at least one rigid object by analyzing trajectory of the at least one background object from temporal image frames obtained by using the camera; andremoving the identified at least one rigid object.8.A video see-through (VST) device for performing hand tracking, the VST device comprising:a camera;at least one sensor;at least one processor including processing circuitry; andmemory storing one or more instructions that, when executed by the at least one processor individually or collectively, cause the VST device to:obtain at least one image of a real-world by capturing a user’s surrounding using the camera and a field of view (FOV) of the VST device, by using the at least one sensor while the user interacts with the VST device,obtain, by using an artificial intelligence (AI) model, the at least one 2D view corresponding to the FOV from a 3D model corresponding to the user’s surrounding,identify at least one background object that is commonly included in the at least one image and the obtained at least one 2D view,remove the identified at least one background object from the at least one image,detect a hand of the user from the at least one image in which the at least one background object has been removed, andperform hand tracking of the hand of the user by applying image processing analysis on the hand of the user.9.The VST device of claim 8, wherein the one or more instructions are further configured to, when executed by the at least one processor individually or collectively, cause the VST device to:correlate the obtained at least one image with the obtained at least one 2D view rendered by using the AI model corresponding to the user’s surrounding.10.The VST device of claim 8, wherein the at least one sensor includes at least one of global positioning system (GPS), inertial measurement unit (IMU), light wave detection and ranging (LIDAR), or infrared sensor, and configured to obtain viewing location and viewing angle of the user, andwherein the one or more instructions are is further configured to, when executed by the at least one processor individually or collectively, cause the VST device to:place a virtual camera in the 3D model using the AI model based on the obtained viewing location and viewing angle of the user, andrender the at least one 2D view that represents a viewpoint of the user by using the virtual camera.11.The VST device of claim 8, wherein the one or more instructions are further configured to, when executed by the at least one processor individually or collectively, cause the VST device to:remove at least one static object from the at least one image of the real-world.12.The VST device of claim 8, wherein the one or more instructions are further configured to, when executed by the at least one processor individually or collectively, cause the VST device to:determine possession time of at least one background object, wherein the possession time constitutes screen time occupied by each of the at least one background object; andremove at least one object from among the at least one background object with the possession time less than a pre-determined threshold.13.The VST device of claim 8, wherein the one or more instructions are further configured to, when executed by the at least one processor individually or collectively, cause the VST device to:determine connectivity of at least one background object with each other while analyzing interactions between each of the at least one background object in the 3D model; andremove at least one object with connectivity higher than a pre-defined value.

Citation Information

Patent Citations

  • 2d image data generation system using of 3D model, and thereof method

    KR1020170074413A

  • Gaze-based object placement within a virtual reality environment

    US20190286231A1

  • In-car cloud VR device and method

    US20210366194A1

  • Techniques to set focus in camera in a mixed-reality environment with hand gesture interaction

    US20220141379A1

  • Visual AURA around field of view

    US20220244776A1