Method and apparatus for determining pose of user wearing head mounted display device
By employing a multi-frame pre-filtering technique and motion compensation, the method and device enhance pose estimation accuracy in HMD systems by distinguishing static from dynamic objects, addressing inaccuracies caused by dynamic objects in existing methods.
Patent Information
- Application Number
- PCT/KR2025/001975
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-11
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods for determining the pose of a user wearing a head-mounted display (HMD) device are hindered by the presence of dynamic objects, leading to inaccuracies in feature detection and matching, which result in erroneous depth calculations and reduced accuracy in pose estimation.
A method and device that differentiate between static and dynamic objects using a multi-frame pre-filtering technique, coupled with velocity and depth-based processing, to enhance pose estimation accuracy by incorporating motion compensation and leveraging a higher number of feature points.
The method and device effectively distinguish static from dynamic objects, ensuring accurate pose estimation and maintaining consistency in localization and mapping, even in unfamiliar scenes with dynamic elements.
Smart Images

Figure KR2025001975_28082025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR DETERMINING POSE OF USER WEARING HEAD MOUNTED DISPLAY DEVICE
[0001] The present disclosure relates to smart devices and Metaverse, and more specifically related to determining pose of a user wearing a head mounted display (HMD) device.
[0002] In broad terms, Head-Mounted Display (HMD) devices can fall into one of three categories: Extended Reality (XR), Augmented Reality (AR), or Virtual Reality (VR) devices. These displays are capable of accurately tracking user movements within a scene, utilizing both localization and mappingmodels. Localization models analyze key features within images, while mapping models use visual cues to estimate user pose.
[0003] Traditional approaches to localization and mapping rely on feature detection and matching methods. However, these methods are often hindered by the quality of image features and can suffer from degraded performance when objects in image frames or videos displayed in the HMD devices are in motion relative to the cameras of the HMD devices. Feature detection is a method used to detect features in an image frame, while matching isused to track features between two or more image frames.. In feature matching, featuresat the current timestamp are detected using landmarks in previous or historical image frames. One of the techniques employed is motion-based method, which estimates the movement of previously detected landmarks from the historical image frames to the image frames at the current timestamp, and the estimated movements are then processed to estimate the final pose.
[0004] However, in dynamic scenes, the movement of dynamic objects differs from their static counterparts. This discrepancy leads to a loss of features and a reduction in accuracy, as the motion calculation fails to match the movement of dynamic objects.
[0005] The descriptor based matching method involves detecting the features of image frames in the current timestamp and comparing them with those of historical image frames. Descriptors are then calculated for all features of the image frames, which are used to match the features between the historical and current frames. The matched features are subsequently utilized for further processing. However, in situations where dynamic objects are present in the scene, the descriptors may match these objects in both the historical and current frames. This is due to the movement of both the camera and the object, resulting in erroneous depth calculations for the features of the frames. As a result, this introduces errors in the pose calculation.
[0006] Thus, it is desired to address the above-mentioned disadvantages or other shortcomings or at least provide a useful alternative.
[0007] Embodiments disclosed herein provide a method of a head mounted display (HMD) device for determining pose of a user wearing the HMD device, within anExtended Reality (XR) scene. The method may comprise determining plurality of objects in the XR scene captured by at least one camera associated of the HMD device. The method may comprise determining at least one static objects and at least one dynamic objects available in the plurality of objects of the XR scene. The method may comprise detecting a motion of the HMD device. The method may comprise determining change in a position of the at least one static objects and a change in the position of the at least one dynamic object based on the motion of the HMD devices and a plurality of frames of the XR scene. The method may comprise determining the pose of the user wearing HMD devices based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object.
[0008] Determining the pose of the HMD device based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object may comprise determining a level of change in the position of the at least one static object. The level of change in position may be determined based on change in the position of the at least one static object. The determining the pose of the HMD device based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object may comprise determining a level of change in the position of the at least one dynamic object. The level of change in position may be determined based on change in the position of the at least one dynamic object. The determining the pose of the HMD device based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object may comprise determining the pose of the HMD device based on the level of the change in the position of the at least one static object and the level of the change in the position of the at least one dynamic object.
[0009] The method may comprise detecting the motion of the HMD device includes receiving sensor data from motion sensors. The method may comprise determining the motion of the HMD device based on the sensor data. The method may comprise receiving a feature points from the plurality of objects in the XR scene captured by the at least one camera of the HMD device. The method may comprise determining by matching each of the feature points of the plurality of feature points in the plurality of frames for depth estimation. The method may comprise determining a refined pose of the user based on the motion of the HMD device and each of the feature points in the plurality of feature points based on the estimated depth.
[0010] Determining the at least one static object and the at least one dynamic object available in the plurality of objects of the XR scene may comprise receiving the plurality of frames of the XR scene. Determining the at least one static object and the at least one dynamic object available in the plurality of objects of the XR scene may comprise determining a plurality of feature points in the plurality of frames. Determining the at least one static object and the at least one dynamic object available in the plurality of objects of the XR scene may comprise segmenting the plurality of feature points into the at least one probable static object and the at least one probabledynamic object based on a deep learning model in the frames. Determining the at least one static object and the at least one dynamic object available in the plurality of objects of the XR scene may comprise marking the plurality of segmented feature points as the at least one probable static object and the at least one probable dynamic object.
[0011] Determining the level of change in the position of the at least one object may comprise determining the plurality of objects in the plurality of frames of the XR scene. Determining the level of change in the position of the at least one object may comprise determining a velocity integration for each object. Determining the level of change in the position of the at least one object may comprise measuring a relative velocity for each object by comparing at least two frames of plurality of frames. Determining the level of change in the position of the at least one object may comprise determining the at least one dynamic object based on the relative velocity. Determining the level of change in the position of the at least one object may comprise determining an inverse of the velocity integration for the dynamic object. Determining the level of change in the position of the at least one object may comprise determining the level of change in the position of the at least one dynamic object based on the inverse of the velocity integration. Determining the level of change in the position of the at least one object may comprise updating the position of the at least one dynamic object based on the determined level of change in the position.
[0012] Determining the at least one static object and at least one dynamic object available in the plurality of objects of the XR scene may comprise determining at least one feature count in the plurality of frames. The feature count may be a number of trackable features in the plurality of frames. The more the number of features in a scene, better will be the determined pose and the pose is the overall translational and rotational movement of the HMD. The transitional pose may be a linear movement and the rotational pose may be an orientation information with respect to coordinate axis. The features relative velocity is compared across the frames and compared with the relative velocity of probable static objects. In case the relative velocity of a feature is much higher as compared to average static feature relative velocity, the feature may be classified as confident dynamic object.
[0013] Determining the pose of the user wearing the HMD based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object relative to the HMD may comprise receiving the motion information of the HMD using a motion sensor in the HMD and the plurality of frames from the HMD and bundling the motion information of the HMD and the frames from the HMD for each image frame. An initial pose of the user wearing the HMD may be determined based on the bundled motion of the HMD and the plurality of frames. Determining the pose of the user wearing the HMD based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object relative to the HMD may comprise performing a refinement of the pose based on the initial pose of the user wearing the HMD and the plurality of frames from the HMD. The refined pose may be determined by minimizing re-projection error of three-dimensional landmarks.
[0014] The three-dimensional landmarks may be the plurality of feature points from the plurality of objects in the XR scene captured by at least one camera.
[0015] Detecting the motion of the HMD may comprise determining an average velocity of the plurality of objects and an average velocity of the at least one camera in the plurality of frames. Detecting the motion of the HMD may comprise determining a difference in a location of the of objects based on the determined average velocity of the plurality of objects and the average velocity of the camera in the frames. Detecting the motion of the HMD may comprise determining a motion compensated frame based on the determined difference in the location of the objects based on the determined average velocity of the objects and the average velocity of the camera in the frames.
[0016] The method may comprise merging the motion compensated frames with the static object in the frames. The merged frames may be used for the pose estimation.
[0017] Merging the motion compensated frames with the static object in the frames may comprise capturing plurality of camera rays projected from a center of the camera passing through feature in the motion compensated frames and the corresponding features in the previous frames. Merging the motion compensated frames with the static object in the frames may comprise determining the depth / location of the feature based on the intersection of camera rays, and the depth / location of the feature is determined the motion compensated frames.
[0018] Embodiments disclosed herein provide a HMD device for determining pose of the user from the XR scene, The HMD device may comprises: at least one camera, memory storing instructions, and at processor operably coupled to the camera, the plurality of sensors, and the memory. The instructions, when executed by the at least one processor, may cause the HMD device to determine multiple objects in the XR scene captured by the camera associated with the HMD device. The instructions, when executed by the at least one processor, may cause the HMD device to determine one or more static objects and one or more dynamic objects available in the objects of the XR scene. The instructions, when executed by the at least one processor, may cause the HMD device to detect a motion of the HMD device and determine a change in a position of the one or more static objects and a change in position of the one or more dynamic objects based on the motion of the HMD device and multiple frames. The instructions, when executed by the at least one processor, may cause the HMD device to determine the pose of the user wearing the HMD device based on change in the position of the one or more static objects and the one or more change in the position of the dynamic objects.
[0019] Embodiments disclosed herein provide a non-transitory computer readable storage medium storing insturctions. The instructions, when executed by at least one processor of a head mounted display (HMD) device, may cause the HMD device to determine multiple objects in the XR scene captured by the camera associated with the HMD device. The instructions, when executed by the at least one processor, may cause the HMD device to determine one or more static objects and one or more dynamic objects available in the objects of the XR scene. The instructions, when executed by the at least one processor, may cause the HMD device to detect a motion of the HMD device and determine a change in a position of the one or more static objects and a change in position of the one or more dynamic objects based on the motion of the HMD device and multiple frames. The instructions, when executed by the at least one processor, may cause the HMD device to determine the pose of the user wearing the HMD device based on change in the position of the one or more static objects and the one or more change in the position of the dynamic objects.
[0020] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating preferred embodiments and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
[0021] The objective of the proposed disclosure is to provide a method and HMD device for estimation of pose of a user wearing the HMD device from a Scene comprising of static and dynamic objects.
[0022] Another object of the embodiments herein is to provide solution for localizing and mapping frames within an unfamiliar scene. This is achieved through a multi-frame pre-filtering technique that effectively distinguishes static objects from dynamic ones, coupled with velocity and depth-based processing to expedite accuracy by employing more number of landmarks from both static and dynamic objects in the unfamiliar scene, all while ensuring that accuracy remains uncompromised.
[0023] Yet another object of the embodiments herein is to provide globally consistent map that keeps pose information to enable loop closure and relocalization incurring minimal errors.
[0024] Yet another object of the embodiments herein is to utilize the extent of motion of the dynamic objects, as measured, as a compensatory parameter in order to determine the ultimate pose of the HMD device from its approximate pose.
[0025] This invention is illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the drawings, in which:
[0026] FIG. 1 illustrates an exemplary depiction of stationary and moving elements within a given setting, according to the embodiments disclosed herein;
[0027] FIG. 2A depicts the direction of camera movement in the event of precise head tracking by a user wearing a HMD device, according to the embodiments disclosed herein;
[0028] FIG. 2B is a representation illustrating the direction of the camera movement in case of inaccurate head tracking of the user wearing the HMD device, according to the embodiments disclosed herein;
[0029] FIG. 3 is a block diagram illustrates the implementation of a view alteration, based on the head tracking movements received from the HMD device, according to the embodiments disclosed herein;
[0030] FIG. 4 is a graph that illustrates the impact of dynamic features in a view displayed on the HMD device's screen and the corresponding head movements of the user wearing the device, according to the embodiments disclosed herein;
[0031] FIG. 5A shows a Six Degrees of Freedom (6DOF) tracking of a user wearing the HMD device, according to the embodiments disclosed herein;
[0032] FIG. 5B depicts the localization and mapping of an object as perceived by the user donning the HMD device, according to the embodiments disclosed herein;
[0033] FIG. 6 depicts a refined block diagram that showcases the Simultaneous Localization and Mapping (SLAM) approach, which enables the determination of the user's pose while donning the HMD device, according to the embodiments disclosed herein;
[0034] FIG. 7 is a visual representation manifestation of the feature tracking process within a given scene and the corresponding estimation of the camera pose of the user donning the HMD device, according to the embodiments disclosed herein;
[0035] FIG. 8 is a block diagram illustrating accurate localization and mapping in unknown scenes for determining the pose of the user wearing the HMD device, according to the embodiments disclosed herein;
[0036] FIG. 9 is a visual depiction of the comparison between the original features of frames and their motion-compensated counterparts, according to the embodiments disclosed herein;
[0037] FIG. 10A depicts a visual representation that elegantly illustrates the alterations in motion flow between dynamic and static objects, according to the embodiments disclosed herein;
[0038] FIG. 10B depicts a visual representation marking the dynamic objects and the static objects respectively, according to the embodiments disclosed herein;
[0039] FIG. 11 is a flow diagram illustrating the segmentation of objects or features in the scene, which aids in segmenting the static and dynamic objects while wearing the HMD device, according to the embodiments disclosed herein;
[0040] FIG. 12A is a block diagram illustrates feature processing based on a type of landmark for determining pose of the user wearing the HMD device, according to the embodiments disclosed herein;
[0041] FIG. 12B is a block diagram illustrating velocity calculation and motion estimator based correction, according to the embodiments disclosed herein;
[0042] FIG. 13A is a schematic that illustrates detection of stationary objects within the scene, which is instrumental in determining the pose of the user donning the HMD device, according to the embodiments disclosed herein;
[0043] FIG. 13B schematically depicts the detection of dynamic objects within the scene for the purpose of determining the pose of the user wearing the HMD device, according to the embodiments disclosed herein;
[0044] FIG. 14 depicts the movements of the camera and objects over time, which are utilized to ascertain the user's pose while wearing the HMD device, according to the embodiments disclosed herein;
[0045] FIG. 15 depicts a visually appealing representation of the motion compensation mechanism employed in both past and present frames to determine the posture of the user donning the HMD device, according to the embodiments disclosed herein;
[0046] FIG. 16 depicts a block diagram that elegantly illustrates the system utilized to ascertain the user's pose while donning the HMD device, according to the embodiments disclosed herein; and
[0047] FIG. 17 is a flow diagram that illustrates a method for determining pose of the user wearing the HMD device, according to the embodiments disclosed herein.
[0048] It may be noted that to the extent possible, like reference numerals have been used to represent like elements in the drawing. Further, those of ordinary skill in the art will appreciate that elements in the drawing are illustrated for simplicity and may not have been necessarily drawn to scale. For example, the dimension of some of the elements in the drawing may be exaggerated relative to other elements to help to improve the understanding of aspects of the invention. Furthermore, the elements may have been represented in the drawing by conventional symbols, and the drawings may show only those specific details that are pertinent to the understanding the embodiments of the invention so as not to obscure the drawing with details that will be readily apparent to those of ordinary skill in the art having benefit of the description herein.
[0049] Various embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. In the following description, specific details such as detailed configuration and components are merely provided to assist the overall understanding of these embodiments of the present disclosure. Therefore, it should be apparent to those skilled in the art that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0050] Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.
[0051] Herein, the term "or" as used herein, refers to a non-exclusive or, unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0052] Throughout the specification, the terms "XR scene" and "scene" are used interchangeably.
[0053] Throughout the specification, the terms "features" and "objects" are interchangeably.
[0054] Embodiments disclosed herein provide a method and a HMD device for determining a pose of a user wearing the HMD device from an XR scene comprises determining multiple objects in the XR scene captured by one or more camera associated with the HMD device, determining one or more static object and one or more dynamic object available in the multiple objects of the XR scene, detecting a motion of the HMD device and determining a change in a position of the one or more static object and a change in a position of the one or more dynamic object based on the motion of the HMD device and a multiple frames. The method further includes determining the pose of the user wearing the HMD device based on the change in the position of the one or more static object and the change in the position of the one or more dynamic object.
[0055] When confronted with both static and dynamic objects in the XR scene, the HMD device is specifically engineered to precisely determine the user's pose while wearing the device. The current disclosure presents a motion compensation method that transforms dynamic objects to exhibit static behavior. Additionally, it involves dynamically thresholding probable dynamic objects and static objects into confident dynamic objects and static objects.
[0056] In conventional techniques, upon receiving frames at the current timestamp and detecting landmarks in preceding frames, moation based detection is utilized to determine the displacement of previously detected landmarks from the preceding frames to the current frames. The identified landmarks are then subjected to further processing to ultimately estimate the user's final pose while wearing the HMD device. However, the movement of dynamic objects diverges from that of static objects, resulting in misalignment during motioncomputation. This leads to a loss of features and a subsequent decrease in accuracy.
[0057] Certain techniques involve detecting objects in both current and preceding frames, with feature descriptor methods used to calculate descriptors for all objects. These descriptors are then utilized to match objects between frames, with the matched objects serving as a basis for further processing. However, when dynamic objects are present in the scene, the descriptors may match for these objects in both preceding and current frames. This can lead to erroneous depth calculations for the features, ultimately resulting in errors in the user's pose calculation due to movement of both the camera and objects.
[0058] In current methods, LiDAR SLAM devices that are feature-aware and capable of detecting dynamic objects rely on the automatic generation of training data. These dynamic objects may include moving cars, humans, and the like, posing a significant challenge for the device as it primarily assumes static objects. In contrast, the present disclosure introduces a novel approach that leverages a sequence of frames to distinguish between dynamic and static objects in the scene, rather than relying solely on LiDAR data. Additionally, a motion compensation technique is employed to transform dynamic objects into static ones, as opposed to filtering them out like existing methods.
[0059] The current techniques tend to discard dynamic features owing to a significant deviation from the underlying optical flow. Similarly, the descriptor-based approach often leads to incorrect matching of dynamic features with those in preceding frames, resulting in less precise pose estimation. Moreover, the use of incorrectly matched feature points in the image can cause inconsistent non-linear optimization. In contrast, the present disclosure introduces a motion compensation technique that enables the existing methods to accommodate dynamic objects even after translation. By utilizing a higher number of matched feature points (both static and dynamic), the present disclosure achieves more accurate pose estimation.
[0060] FIG. 1 illustrates an exemplary depiction of stationary and moving elements within a given setting, according to the embodiments disclosed herein;
[0061] FIG 1 depicts a scene comprising both static objects (102) and dynamic objects (101). The scene is composed of multiple objects, including traffic lights, moving vehicles, trees, buildings, and more. The objects in the image are classified based on their respective types: static and dynamic. Static objects remain constant within the scene, while dynamic objects can change position irrespectove of user's movements while wearing the HMD device. The static objects serve as stable reference points for a SLAM device to localize and map the surroundings. Examples of static objects include walls, floors, buildings, and other permanent structures. The localization process accurately navigates and comprehends the environment of the scene. Once localized, the SLAM device can use this information to update its map, plan paths, and perform tasks based on its knowledge of its position relative to its environment.
[0062] FIG 2A depicts the camera movement direction in case of precise head tracking of the user wearing the HMD device (1700), as per the disclosed embodiments. If the user moves towards the right while looking at the 3D object, keeping it at the center of their eyesight, the camera's perspective in the digital environment is determined by the user's viewpoint. The virtual camera updates accordingly when the user moves physically, such as walking or using a handheld controller to simulate movement. If the user moves or turns right in the real-world space where they are using the HMD device (1700), the sensors in the HMD device (1700) track the movement. To provide a comfortable and immersive experience, the user's attention is focused on a specific object within their field of vision. As the user moves, the virtual camera adjusts its position and orientation to ensure that the objects remain at the center of the user's vision.
[0063] In an embodiment, a cube (a) is positioned at the center of the camera and subsequently rendered at the center of the frame, as depicted in the FIG 2A and FIG 2B. As the camera moves from time T-1 (201a) to T (201b) and then to T+1 (201c), the 3D object within the scene is precisely rendered. This results in an accurate depiction of the cube, providing users with a seamless and captivating experience when utilizing any HMD device (1700).
[0064] FIG. 2B is a representation illustrating the direction of the camera movement in case of inaccurate head tracking of the user wearing the HMD device (1700), according to the embodiments disclosed herein.
[0065] With reference to FIG2B, the positioning of a cube (a) in the camera is elucidated in the event of imprecise head tracking. When the user direction is precisely computed, as seen in timestamps T-1 (201d) and T (201e), objects are rendered at the frame's center, in tandem with the user's actual movement. However, at timestamp T+1 (201f), owing to a pose calculation error, head movement becomes inaccurate, leading to the object being rendered marginally to the right of the image's center, in a divergent orientation.
[0066] FIG. 3 is a block diagram illustrates illustrates the implementation of a view alteration, based on the head tracking movements received from the HMD device (1700), according to the embodiments disclosed herein.
[0067] The third figure encompasses an Extended Reality (XR) camera (301), a head movements tracker (302), a relative movement determination unit (303), and a view changer (304). At the S1, the XR scene is projected onto the Head-Mounted Display (HMD device) (1700). The XR scenes can comprise a digital environment or space that users engage with using XR cameras. These scenes may be a virtual world in Virtual Reality (VR), an augmented view of the real world with digital overlays in Augmented Reality (AR), or a mixed environment where virtual objects seamlessly interact with the user's physical surroundings.
[0068] At S2, the head movements are meticulously monitored (302) through the use of a head movement tracker. By analyzing the tracked head movements, the relative movements are accurately determined. The user's head movements, while wearing the HMD device (1700), are gauged through a variety of sensors including gyroscopes, accelerometers, and sometimes magnetometers. These sensors are capable of detecting even the slightest changes in the HMD device's (1700) orientation and movement, as they continuously measure the headset's orientation in the 3-dimensional space. This includes pitch (up and down movement), yaw (side-to-side rotation), and roll (tilting from side to side). Furthermore, some advanced VR devices possess sensors that can track the headset's position within a defined physical space, allowing for both rotational and translational movements (forward / backward, left / right, up / down). The sensor data collected from the various sensors is then transmitted to the VR device, which promptly updates the virtual environment in real-time to match the user's head movements.
[0069] At S3, as change in view displayed on a user display affected by the movement and in turn depends on the accurate head tracking. The user display is in configuration with the HMD device (1700). The inaccuracies in head tracking affects smoothness of the rendered view creating a jittery experience for the user. The relative movement determination unit (303) determines relative movements for reference frames (N-1) (303).
[0070] At S4, the transformations and view modifications are executed sequentially to reveal the updated view through a view changer. This process is then repeated for subsequent frames. The determination of the changes in the rendered view involves the calculation of the pose that is utilized in object rendering. As the user's movement directly impacts the changes in the rendered view on their display, the accuracy of the head tracking is critical to ensure a seamless and smooth experience. Any discrepancies in the head tracking may result in an unsteady and jittery view for the user.
[0071] In an embodiment, the accuracy of the system is significantly impacted by the presence of dynamic objects within the scene. This is due to the low object count in the scene and poor matching in existing methods. The present disclosure addresses the issue of low feature count by including missed dynamic features, and in turn, provides an accurate view change with user movement. A feature count refers to the number of trackable features, such as edges, lines, corners, and sharp intensity changes, within the frames using which the camera's pose is being calculated. The pose of a camera object comprises two components: the translational component, which provides linear movement of the object from the origin, and the rotational component, which provides information about the object's orientation with respect to the coordinate axis.
[0072] FIG. 4 is a graph that illustrates the impact of dynamic features in a view displayed on the HMD device's (1700) screen and the corresponding head movements of the user wearing the device, according to the embodiments disclosed herein;
[0073] With reference to FIG. 4, the graph aptly depicts the user's precise path, serving as the ground truth for their trajectory. The identification of dynamic objects in the scene enables unrestricted movement, while the absence of such elements confines the user's motion to a specific area. The X and Y axes signify the distance covered in the respective directions, measured in meters. The map is generated by projecting the actual 3-dimensional trajectory onto the X-Y plane, rendering it effortlessly comprehensible. In the graph of the FIG. 4, line 'a' denotes local motion, 'b' represents no local motion, and 'c' represents the actual map. The graph delineates the impact of dynamic features in view, as well as the tracking of head movement.
[0074] FIG. 5A shows a Six Degrees of Freedom (6DOF) tracking of a user (501) wearing the HMD device (1700), according to the embodiments disclosed herein. FIG. 5B depicts the localization and mapping of an object as perceived by the user donning the HMD device (1700), according to the embodiments disclosed herein. The 6DOF tracking of the user (501) wearing the HMD device (1700) track the movement of the user in 3D space using parameters such as (pitch, yaw, roll, forward or backward, left or right, up or down). The tracking is specifically applied to the user (501) wearing the HMD device (1700), which is a device used in the XR devices to provide a visual and auditory immersive experience. The HMD device (1700) can accurately capture and record the user's (501) movement in three-dimensional space along all axes. The HMD device (1700) allows for a high level of precision in tracking both rotational movements (like tilting or turning the head) and translational movements (like moving forward, backward, left, right, up, or down).
[0075] The 6DOF tracking in the context of an HMD device (1700) enables a highly immersive VR experience. It allows the user (501) to move their head freely in all directions and accurately reflects the user (501) movements within the virtual environment. This creates a strong sense of presence and realism, as the virtual world responds to the user's (501) movements in a natural and intuitive way.
[0076] The user (501) is equipped with the HMD device (1700) on the head that typically covers the user's eyes and sometimes ears. The HMD device (1700) provides the visual and auditory elements of the virtual environment.
[0077] The localization involves determining the position and the orientation of the HMD device (1700) or a robot within a given environment, whether known or unknown. On the other hand, mapping involves creating a comprehensive representation or model of the environment being navigated by the HMD device (1700) or robot. This representation can be either 2D, such as floor plans, or 3D, such as complete models of spaces. To effectively navigate, the HMD device (1700) or robot requires an accurate map of the environment and an understanding of its own position within that map. As the HMD device (1700) moves, localization continuously updates the position of objects, while mapping updates the environment representation. The applications of localization and mapping are numerous, including autonomous robots, self-driving cars, drones, augmented reality systems, and other related fields.
[0078] FIG. 6 depicts a refined block diagram that showcases the Simultaneous Localization and Mapping (SLAM) approach, which enables the determination of the user's pose while donning the HMD device (1700), according to the embodiments disclosed herein.
[0079] Referring to the FIG. 6, the SLAM camera includes feature extraction and matching component (603), depth estimation component (604), sensor fusion and bundle adjustment unit (605) and local and global bundle adjustment unit (607).
[0080] At S1, a sequence of frames (602) are given as input to feature extraction and matching component (603). The sequence of frames are referred as multiple frames or frames or media frames. The media frames are retrieved from the XR scenes.
[0081] At S2, the feature extraction and matching component (603) extracts and matches the feature received from the sequence of frames. The sequence of frames are retrieved from the XR scenes and determine the multiple objects from the XR scenes. The XR scenes are captured using the camera associated with the HMD device (1700). The camera can be SLAM camera. The static objects and dynamic objects are determined from the multiple objects. The multiple objects are received from the sequenc of frames.
[0082] At S3, the depth estimation component (604) uses the matched features for depth estimation. The depth estimation component (604) estimate the depth in the sequence of frames. The depth estimation component (604) determines the distances to objects in the scene. The depth estimation creates accurate 3D maps and understand the spatial layout of the environment. The depth estimation uses two or more cameras with a known baseline (distance between them) to capture images from slightly different perspectives. By analyzing disparity (difference in pixel positions of corresponding points) between the sequence of frames, the HMD device (1700) can determine depth information. The Time of Flight (ToF) sensors emit a light signal and measure the time it takes for the signal to bounce back. This provides direct depth information for each pixel in the camera's field of view.
[0083] The LiDAR systems emit laser pulses and measure the time it takes for the pulses to return after bouncing off objects. This creates a point cloud, which can be used for 3D mapping and localization.
[0084] At S4, the sensor fusion and Bundle Adjustment (BA) component (605) collects the outputs from the different modules such as feature extraction and matching component (603), depth estimation component (604) and the like to combine data received from the different modules to obtain more accurate and comprehensive understanding of the environment than would be possible with individual sensor alone. Each sensor provides distinct types of information. For example, cameras can be used for visual information, IMU (Inertial Measurement Unit) for acceleration and gyroscope data, and LiDAR for depth information. The sensor fusion method integrates data from different sensors to enhance accuracy in tasks like localization and mapping. The BA employs a mathematical optimization technique to refine the estimated 3D structure of the scene and the camera poses within the scene. It simultaneously enhances the positions of the 3D points in the environment and the positions and orientations of the cameras observing these points. The refinement is executed by minimizing the difference between observed feature locations in images and the predicted locations based on the current 3D structure and camera poses.
[0085] The S5 unit employs a local and global bundle adjustment approach (606) to refine the pose of the user. This method involves identifying multiple landmarks, such as landmark 1, landmark 2, and landmark 3, based on the user's pose and new poses. As illustrated in FIG 6, the distances d1, d2, and d3 represent the distance from the user's pose to landmark 1, landmark 2, and landmark 3, respectively. Similarly, d1, d2, and d3 represent the distance from the robot's pose to landmark 1, landmark 2, and landmark 3, respectively. The Δx in 607 denotes the distance between the user's current position and the updated or current position and orientation of the user within the environment.
[0086] ...(1)
[0087] ...(2)
[0088] ...(3)
[0089] ...(4)
[0090] Pred - IMU Predicted Pose
[0091] refTracker - Refined Tracker Pose
[0092] refMapper - Refined Mapper Pose
[0093] ...(5)
[0094] ...(6)
[0095] ...(7)
[0096] ...(8)
[0097] In an embodiment, the user poses and new user poses are utilized to illustrate the landmark 1, landmark 2, and landmark 3. Additionally, distances such as d1, d2, and d3 refer to the distance between each landmark and the user pose. The Local bundle adjustment and global bundle adjustment process refines the structures and camera parameters in a 3D reconstruction pipeline. The structures include 3D geometric information such as spatial relationships and positions of points in space. These structures can be 3D points or other spatial relationships. The local bundle adjustment refines the parameters of a limited subset of images in the reconstruction, typically a small neighborhood around a specific area of interest. Meanwhile, the global bundle adjustment optimizes the parameters of the entire reconstruction, including camera poses and 3D points.
[0098] FIG. 7 is a visual representation manifestation of the feature tracking process within a given scene and the corresponding estimation of the camera pose of the user donning the HMD device (1700), according to the embodiments disclosed herein.
[0099] The SLAM camera is demonstrated in the scene. The marked points in the FIG. 7 (701a, 701b, 701c, ..., 701n.) are feature points. The feature points or the objects in the 3D space are referred as landmarks. The objects can be the static objects and the dynamic objects. A tracks (702) as shown in the FIG. 7 is drawn in static environment and the camera is moving in the environment. The static environment can be the scene or the video. In an embodiment, the tracks (702) can get diverted from actual motion path of the HMD device (1700) in presence of the moving objects in the video or the scene. Referring to the FIG. 7, the marked point represents the identified static objects in the scene. The visual representation shows the historical data of the motion of the user wearing the HMD device (1700). The static objects represented in the FIG. 7, can be more than the showed static objects and the image is only for the illustration purpose.
[0100] FIG. 8 is a block diagram illustrating accurate localization and mapping in unknown scenes for determining the pose of the user wearing the HMD device (1700), according to the embodiments disclosed herein.
[0101] At S1, the sequence of frames (602) are given as input to feature extraction and matching component (603). The sequence of frames are referred as multiple frames or frames or media frames. The media frames are retrieved from the XR scenes.
[0102] AT S2, the feature extraction and matching component (603) extracts and matches the feature received from the sequence of frames. The sequence of frames are retrieved from the XR scenes and determine the multiple objects from the XR scenes. The XR scenes are captured using the camera associated with the HMD device (1700). The camera can be SLAM camera. The static objects and dynamic objects are determined from the multiple objects. The multiple objects are received from the sequenc of frames.
[0103] The feature extraction and matching component (603) includes an object segmentation component (802), a motion estimation component (803), a motion compensation component (804), a landmark merging component (805).
[0104] The object segmentation component (802) receives the input of image segments having dynamic objects. The object segmentation component (802) classifies as probable objects and probable dynamic objects. The object segmentation component (802) includes a pre-filtering method that filters out the objects as static and dynamic objects.
[0105] The motion estimation component (803) determines motion flow based on the pre-filtering. The motion estimation determines the movement or change in position of objects or user viewpoint between the consecutive frames of the sequence frames. The motion estimation helps tracking the user's head movements, hand movements, or even the movement of objects within the environment. The HMD device (1700) adjusts the perpective of the virtual environment based on the user head movements.
[0106] The motion compensation component (804) adjusts the virtual content in response to the detected motion of the HMD device (1700) in order to maintain stable and coherent visual experience. The motion compensation compensated the relative movement of dynamic objects in the scene with respect to the user such that the dynamic object behaves as if they are static.Based on the probable dynamic objects, confident dynamic objects are captured. The motion estimation component (803) filters out the objects and obtain the confident dynamic objects.
[0107] The static objects and the corrected dynamic objects are merged to match the landmarks and identify the variation in the motion. The motion flow of different objects in case of static objects is smooth when compared to dynamic scenes which sometimes can be impulsive. When comparing the motion flow of same object in case of Dynamic vs Static, it can be seen that the flow is completely different for the same object depending upon if it is the dynamic obejcts or static objects. When such objects are used for tracking, they will introduce more errors.
[0108] In an embodiments, static and dynamic features are determined using the equations given below:cc
[0109] The features on each probable dynamic objects in the sequence of frames is matched first with the same feature of the same dynamic objects on consequent frames. The trail of the matched features provides the motion of the object in the scene. Given a series of such motion changes for the feature, an average velocity can be calculated of an object using all the features present on the object using equation 10.
[0110] At S3, the depth estimation component (604) uses the matched features for depth estimation. The depth estimation component (604) estimate the depth in the sequence of frames. The depth estimation component (604) determines the distances to objects in the scene. The depth estimation creates accurate 3D maps and understand the spatial layout of the environment. The depth estimation uses two or more cameras with a known baseline (distance between them) to capture images from slightly different perspectives. By analyzing disparity (difference in pixel positions of corresponding points) between the sequence of frames, the HMD device (1700) can determine depth information. The Time of Flight (ToF) sensors emit a light signal and measure the time it takes for the signal to bounce back. This provides direct depth information for each pixel in the camera's field of view.
[0111] The LiDAR systems emit laser pulses and measure the time it takes for the pulses to return after bouncing off objects. This creates a point cloud, which can be used for 3D mapping and localization.
[0112] At S4, the sensor fusion and Bundle Adjustment (BA) component (605) collects the outputs from the different modules such as feature extraction and matching component (603), depth estimation component (604) and the like to combine data received from the different modules to obtain more accurate and comprehensive understanding of the environment than would not be possible with individual sensor alone. Different sensors provide different types of information. For instance, the SLAM device can use cameras for visual information, IMU (Inertial Measurement Unit) for acceleration and gyroscope data, and LiDAR for depth information. Sensor fusion method integrate data from the different sensors to improve accuracy in tasks like localization and mapping. The BA is a mathematical optimization technique used to refine the estimated 3D structure of the scene and the camera poses within the scene. The BA simultaneously refines the positions of the 3D points in the environment and the positions and orientations of the cameras observing these points. The refining is performed by minimizing the difference between observed feature locations in images and the predicted locations based on the current 3D structure and camera poses.
[0113] At S5, local and global bundle adjustment component (606) is used for pose refinement. The pose refinement method includes determine multiple landmarks such as landmark 1, landmark 2, and landmark 3 and the like are shown based on the pose of the user and new poses of the user. Referring to FIG. 6, d1, d2, d3 and the like refers to the distance from the landmark 1, landmark 2, and landmark 3 to the user pose, respectively, and d1, d2, d3 refers to the distance from the landmark 1, landmark 2, and landmark 3 to the robot pose, respectively. The Δx in 607 refers to the distance between the user pose and the new user pose. The new user pose refers to updated or current position and orientation of the user within the environment.
[0114] In an embodiment, using a feature detector and a matching technique to match the features from precceding frames and wait for feature matching of dynamic objects.Using the motion estimator results to address and negate the effect of object motion such that the objects are considered as if they are static and applying the processing same as static features. Finally merging both types of features for further processing.
[0115] In intertial measurement component (IMU) pre-integration, processes the gyro meter and accelerometer sensor data and integrates the features to estimate relative rotation, velocity and position. Uses these estimates to estimate an initial camera pose to be refined further in next steps.
[0116] Mapping method maintains a consistent global map for accurate loop closing and relocalization in cases the tracking is lost. Thereafter, the mapper parallelly keeps updating the map with new information from tracker and refines the pose. The mapping method uses only the confident static features to avoid inaccuracies arising from dynamic features.
[0117] In the depth estimation, the projection of landmarks in the multiple frames are used and performs a non-linear optimization to estimate depth in order to minimize the overall reprojection error. A loop closing happens when the user visits the same location after some time. Moreover, the mapping method closes the loop if this happens and updates the tracker with new refined poses after running the non-linear optimization. Relocalization is used when tracking is lost.
[0118] In the sensor fusion and pose estimation (805), initial pose is calculated from the IMU, the landmarks and projections of the landmarks. The depth estimator performs a refinement iteration in order to minimize the reprojection error of all the landmarks in the multiple frames in which it is visible.
[0119] FIG. 9 is a visual depiction of the comparison between the original features of frames and their motion-compensated counterparts, according to the embodiments disclosed herein;
[0120] Referring to the FIG. 9, the visual image (image 1, image 2) includes one or more persons (person 1, person 2). The images can be the multiple frames or the sequence of frames. The images and the number of features in the images can be more than two and not restricted. The person 1 can represent the dynamic features, on which motion compensation is applied, in order to make them behave as static features and used along with other static features resulting in increase of feature count and ultimately the accuracy. The person 1 can be identified as a static object. The image can include other features such as chairs, tables and the like.
[0121] The motion estimator technique determines average motion of static objects and uses the average motion to classify the objects as either static or dynamic. The objects are classified as static or dynamic based on the pixel fraction, and the threshold is dynamically computed using confidence values from previous frames. In the existing methods, all the matched features are used for pose calculation, which will affect accuracy and accumulate drifts over time. The present disclosure handles matched dynamic features separately and compensates the motion so that these features can be used as if they are static. The matching marked (a) represents the dynamic features, on which motion compensation is applied, in order to make them behave as static features and can be used along with other static features resulting in increase of feature count and ultimately the accuracy. The matching marked as (b) is identified as the static objects.
[0122] FIG. 10A depicts a visual representation that elegantly illustrates the alterations in motion flow between dynamic and static objects, according to the embodiments disclosed herein.
[0123] FIG. 10B depicts a visual representation marking the dynamic objects and the static objects respectively, according to the embodiments disclosed herein.
[0124] Referring to 10A, the multiple frame sequences are compared. The images includes the multiple objects. The person 1 can represent the dynamic features, on which motion compensation is applied, in order to make them behave as static features and used along with other static features resulting in increase of feature count and ultimately the accuracy. The person 2 can represent the static object. The image 1 and image 2 are compared to determine that there is a change in the camera pose and dynamic objects movement. Image 3 represents the difference in the motion of the multiple objects. The multiple objects can be both static objects and dynamic objects.
[0125] The motion flow of the multiple objects in case of the static objects is smooth when compared to dynamic objects, which are sometimes impulsive. The images shows the difference between the two images and helps visualize the motion of different objects in two consecutive frames (image 3). When comparing the motion flow of the same object in case of dynamic objects versus static objects, it is seen that the flow is completely different for the same objects depending on whether the objects are dynamic or static. When such objects are used for tracking, the objects introduce more errors. Hence, the proposed method of compensates the motion details helps make tracking more stable. Particualrly, the more the number of features in a scene, better will be the determined pose and the pose is the overall translational and rotational movement of the HMD. The transitional pose is a linear movement and the rotational pose is an orientation information with respect to coordinate axis. Finally, the features relative velocity is compared across the frames and compared with the relative velocity of probable static objects. In case the relative velocity of a feature is much higher as compared to average static feature relative velocity, the feature is classified as confident dynamic object.
[0126] In an embodiment, a method for accurate localization and mapping in unknown scenes is disclosed. The object segmentation (Pre-Filtering) estimates the confidence of all the objects in the scene as probable static and dynamic objects. The method uses a segmentation engine with attention based on object contours to provide accurate segmentation around object edges. To decrease the computational complexity, the segmentation is applied at lower resolution and while upscaling using the original contours for correct object wise segmentation. The output received can be segmentation in current frame.
[0127] FIG. 11 is a flow diagram illustrating the segmentation of objects or features in the scene, which aids in segmenting the static and dynamic objects while wearing the HMD device (1700), according to the embodiments disclosed herein.
[0128] Referring to 1201 in the FIG. 11 includes the sequence of frames (602). The multiple frames are given as input to different convoluition blocks (1107, 1108, 1109, 1110, ...). The frames given as inputs to the different convolution block can be N (1102), N-1 (1103), N-2 (1104), N-3 (1105), and N-4 (1106), respectively.
[0129] In an embodiment, in the object segmentation, the network consists of a limited set of convolutional blocks that are used iteratively in subsequent iterations. All the intermediate feature vectors from the previous frame is stashed in memory to be used for next frame.
[0130] Upon Arrival of Next Frame (i.e., Current Frame):
[0131] The Intermediate Feature Vector of frame N-5 is discarded.
[0132] Intermediate Feature Vector of frame N-4 is convolved with conv Block4.
[0133] The Intermediate Feature Vector of frame N-3 is convolved with conv Block3.
[0134] The Intermediate Feature Vector of frame N-2 is convolved with conv Block2.
[0135] The Intermediate Feature Vector of frame N-1 is convolved with conv Block1.
[0136] The Intermediate feature vectors is stashed for next frame and these features are passed thorough conv Block [6..9], followed by deconvBlock to generate the final probable candidate map for the current frame.
[0137] Conv1 : H X W -> H / 2 x W / 2
[0138] Conv2 : H / 2 x W / 2 -> H / 4 x W / 4
[0139] Conv3 : H / 4 x W / 4 -> H / 8 x W / 8
[0140] Conv4 : H / 8 x W / 8 -> H / 16 x W / 16
[0141] Conv5 : H / 16 x W / 16 - > H / 32 x W / 32
[0142] 1113 is a feature encoding of conv block 1 (1107), conv block 2 (1109), conv block 3 (1110), conv block 4 (1111), and conv block 5 (1112). Further, the conv blocks provide the corresponding output as conv block 6 (1114), conv block 7 (1115), conv block 8 (1116), and conv block 9 (1117). Finally, the output received as deconvolution blocks (1118) that provides a probable candidate map (1119).
[0143] FIG. 12A is a block diagram illustrates feature merging based on a type of landmark for determining pose of the user wearing the HMD device (1700), according to the embodiments disclosed herein.
[0144] At S1, the dynamic objects (1201) are given as inputs to the feature detection model (1202). The dynamic objects refers to elements in the environment that are subjected to change or movement over time. The feature detection model (1202) includes identifying distinctive points or regions in the images that can be used for further analyses, such as matching, tracking, or 3D reconstruction.
[0145] At S2 and S3, the velocity is determined from the detected features. The velocity (1203) is determined using the equations 9, 10 and 11. Different velocities are determined for the motion estimation and correction (1204). The calculations include determining relative velocity and average velocity.
[0146] Static objects (1206) may be determined in the scene. At least one object among a plurality of objects may be categorized as confident static objects. The static objetcs are static (i.e. non moving) objects. The static objects may be essential in SLAM to accurately calculate the user's head pose.
[0147] At S5, the static objects (1206) are used for the feature detection (1207). Similar to the feature detection model (1202), the features (i.e. distinctive points or regions in the images that can be used for further analyses, such as matching, tracking, or 3D reconstruction) can be detected from the static objects (1206).
[0148] At S4 and S6, the static features and the motion corrected dynamic features are used for feature merging (1205). The static features and the motion corrected dynamic features are merged to obtain overall features in the scene. The merged features can be used for further processing.
[0149] FIG. 12B is a block diagram illustrating velocity calculation and motion estimator based correction, according to the embodiments disclosed herein.
[0150] Referring to the FIG. 12B, a velocity calculation and motion estimator based correction (1208) is performed by a process in which, at 1209, prior velocity information from the last k frames of each feature point is considered.
[0151] At 1210, the velocity integration for each feature point (1210) is performed.
[0152] At 1211, estimating the position of dynamic feature points using the inverse of integrated velocity (1211) based on the relative velocity calculation from previous to current for each feature point using a motion estimator (1212). At 1213, dynamic feature points in static context are provided as output.
[0153] The features in the each frame are matched first with the same feature of the same car on subsequent frames. The trail of the matched features provides the motion of the object in the scene. Given the series of such motion changes for the feature, the average velocity of an object is calculated using all the features present on the object.
[0154] ...(9)
[0155] ...(10)
[0156] ...(11)
[0157] FIG. 13A is a schematic that illustrates detection of stationary objects within the scene, which is instrumental in determining the pose of the user donning the HMD device (1700), according to the embodiments disclosed herein. The points e, f and g represents the camera. The FIG. 13A shows the images and ray coming from the camera and going through the projection of point on the image plane is the camera ray. The intersection of the camera rays is used to calculate the depth of the objects in real world. The points indicated in the FIG. 13A are static (a, b, c). When the object is static the camera rays intersect to a single 3D point in space, but when the object becomes dynamic as can be seen in FIG. 13B, they fail to triangulate in 3D space hence converging to find wrong 3D point.
[0158] In an embodiment, the objects detected are static objects and the 3D landmark (d) is the point where the user is gazing at. The P1, P2 and P3 represents the projections of the 3D landmark in the different frames.
[0159] FIG. 13B schematically depicts the detection of dynamic objects within the scene for the purpose of determining the pose of the user wearing the HMD device (1700), according to the embodiments disclosed herein. Referring to FIG. 13B, the point P1 is static at time T-1 and is not moving with respect to surroundings. All the features at the timestamp T-1 include one motion factor i.e. camera motion. The point P2 is the projection of same 3D point in space as P1 at time T-1, and is moving with respect to the surroundings at time T. The other surrounding features have one motion factor i.e. camera motion and the point P2 have two motion factors camera motion and its own motion. The point P3 is the projection of same 3D point in space as P2 at time T and is still moving relative to the surroundings at time T + 1. The other surrounding features have one motion factor i.e. camera motion , but the point P3 have two motion factors camera motion and its own motion. Motion Compensation can move the feature from point 1 to point 2 and correct depth is calculated as shown in FIG. 13B.
[0160] FIG. 14 depicts the movements of the camera and objects over time, which are utilized to ascertain the user's pose while wearing the HMD device (1700), according to the embodiments disclosed herein. The visual representation shows a depiction of the objects moving closer to the camera and the size of the image increases. The camera shows different viewing angles of object and object movement with time i.e., T-3, T-2, T-1, T, T+1, T+2, T+3.
[0161] The objects are represented as (O1, O2, ..., O7) and is moving in front of the camera. As the object comes closer to the camera, the size of the object increase with time and as the object starts moving away from the camera the size of the object starts decreasing. The count of features on dynamic objects increase with time and if used with other static features, the robustness of tracking is maintained. If the dynamic features are not used, the feature count goes down, resulting in decrease in pose accuracy.
[0162] FIG. 15 depicts a visually appealing representation of the motion compensation mechanism employed in both past and present frames to determine the posture of the user donning the HMD device (1700), according to the embodiments disclosed herein.
[0163] Referring to the FIG. 15, depiction of motion compensation is shown with respect to the previous frames and the current frame (n, n+1) respectively. The various parameters of both frames are as follows:
[0164]
[0165] In an embodiment, motion compensation is implemented to decrease the number of detected features. This compensation accounts for errors arising from multiple features obtained from dynamic objects, which, in turn, reduces depth accuracy due to additional errors. Moreover, the accuracy of pose decreases as the count of data points reduces. However, in motion estimator-based correction, the number of detected features in the frame increases by utilizing both static (1601a) and corrected dynamic features (1601b). This increase in landmarks leads to a higher number of projections in various images, resulting in a more precise solver minimization. As a result, the accuracy of depth also improves with the increase in projection.
[0166] In an embodiment, the merging of landmarks involves the integration of motion compensated landmarks with static features (1601a) that are already present in the frame. The projection of these features is calculated in previous frames to determine depth. The merged features are then forwarded to a tightly coupled solver for precise pose calculation. These previous frames are interchangeably referred to as preceding frames.
[0167] FIG. 16 depicts a block diagram that elegantly illustrates the system utilized to ascertain the user's pose while donning the HMD device (1700), according to the embodiments disclosed herein.
[0168] Referring to the FIG. 16, in the HMD device (1700), a camera (1701) is connected to the multiple sensors (1702). The HMD device (1700) determines the multiple objects in the XR scene captured by the camera (1701) associated with the HMD device (1700). The pose estimation controller (1705) in configuration with the camera (1701), the multiple sensors, the memory (1703) and the processor (1704) to determine the pose of the user wearing the HMD device (1700).
[0169] The memory (1703) is configured to store instructions to be executed by the processor (1704). The memory (1703) can include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory (1603) may, in some examples, be considered a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted that the memory (1703) is non-movable. In some examples, the memory (1703) is configured to store larger amounts of information. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in Random Access Memory (RAM) or cache).
[0170] The processor (1704) may include one or a plurality of processors. The one or the plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The processor (1604) may include multiple cores and is configured to execute the instructions stored in the memory (1703).
[0171] The multiple sensors (1702) can include accelerometers, gyroscopes, magnetometers, intertial measurement units, laser measurement units, laser based positional tracking, ultrasonic sensors, facial expression sensors and the like. The various sensors used in the HMD device (1700) are used to track the motion, orientation, and position of the head of the user, and the like.
[0172] The pose estimation controller (1705) determines the pose of the user wearing the HMD device (1700). The multiple objects are determined from the XR scene captured by the camera associated with the HMD device (1700). The multiple objects are classified as static objects and dynamic obejcts in the XR scene. The motion of the HMD device (1700) is determined based on the change in position of the static object and the change in the position of the dynamic objects. The pose of the user is determined by measuring a level of change in the position of the static objects and the dynamic objects. The sensor data is collected from the multiple sensors to determine the refined pose of the user. The feature points and matching with the multiple features points in the different frames to determine the depth of the feature points in the XR scene.
[0173] In an embodiment, the pose estimation controller (1705) use a sequence of preceding frames and the current frames to classify the objects in the scene called pre-filtering. The classified objects are marked as probable static obejcts and the dynamic objects. The probable objects are then passed through visual motion classified to estimate the average motion incurred by the static obejcts and dynamically calculate the thrshod for confidently classifying the probalble dynamic objects into the confident static and dynamic objects. The confident static and dynamic objects are processed separately, using the velocity correction based method followed by feature merging. Also, only the static features are used for mapping the user in the scene, in order to avoid error accumulation in mapper.
[0174] The pose estimation controller (1705) may be implemented using at least one processor. The pose estimation controller (1705) and the processor (1704) may be integrally referred to as at least one processor.
[0175] In an embodiment, the virtual environment allows for seamless multi-user interaction by enabling the user to interact with different objects from various angles. Additionally, head movement of the user can be determined to enhance the virtual experience. A globally consistent map is utilized to maintain pose information, which facilitates loop closure and relocalization with minimal error. The present disclosure achieves reduced memory usage through selective utilization of static features, all while maintaining accuracy. Furthermore, computation time is significantly reduced as the CPU processes a smaller number of landmarks for mapping, without compromising accuracy.
[0176] FIG. 17 is a method flow diagram of determining pose of the user wearing the HMD device (1700), according to the embodiments disclosed herein.
[0177] Referring to the FIG. 17, the method for determining pose of the user wearing the HMD device (1700). At 1801, the multiple objects are determined in the XR scene captured by the camera associated with the HMD device (1700). The multiple objects can be static obejcts and the dynamic objects. The dynamic objects refers to elements in the environment that are subjected to change or movement over time. The feature detection model (1202) includes identifying distinctive points or regions in the images that can be used for further analyses, such as matching, tracking, or 3D reconstruction.
[0178] At 1802, the static objects and the dynamic objects available in the objects of the XR scene. The XR scenes can be a digital environment or space that the users interact with using XR cameras. The XR scenes can be the virtual environment in the VR, the augmented view of the real world with digital overlays in the AR, or the mixed environment where virtual objects interact with the user's physical surroundings.
[0179] At 1803, the motion of the HMD device (1700) are detected. The motion of the HMD device (1700) is determined by receiving the data from the motion sensors and the feature points from the XR scene captured by the camera. The feature points in the feature points for depth estimation and determining the refined pose of the user based on the motion of the HMD device (1700) and the feature points based on the estimated depth.
[0180] At 1804, the change in the position of the static obejcts and the change in the position of the dynamic objects are determined based on the motion of the HMD device (1700) and the multiple frames.
[0181] At 1805, the pose of the user wearing the HMD device (1700) is determined based on the change on the position of the static obejcts and the change in the position of the dynamic objects. The level of change in the position of the dynamic objects include determining the objects and based on the objects the velocity integration is determined. The relative velocity is determined based on the dynamic obejcts and inverse of the velocity integration is determined for the dynamic obejcts to identify the level of change in the position of the dynamic objects and updating the position of the dynamic objects based on the change in the position.
[0182] The various actions, acts, blocks, steps, or the like in the method may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some of the actions, acts, blocks, steps, or the like may be omitted, added, modified, skipped, or the like without departing from the scope of the invention.
[0183] An essential aspect of interacting with virtual objects or other users in a virtual or augmented scenario is the precise estimation of head movement and user motion within the scene. However, when a user is wearing a HMD and interacting with various dynamic objects such as moving users, pets, or thrown balls, the existing method may fall short due to inadequate feature matching, leading to inaccurate pose estimation. The proposed solution enables the HMD device (1700) to treat these objects differently, thereby facilitating the estimation of pose using the overall scene. The invention's technical value lies in its ability to seamlessly facilitate interaction within the metaverse or with other virtual or augmented objects. This is achieved through a sturdy and reliable method for estimating pose in a dynamic scene.
[0184] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope of the embodiments as described herein.
Claims
1.A method of a head mounted display, HMD, device (1700) for determining pose of a user wearing the HMD device (1700) within an extended reality, XR, scene, the method comprising:determining (1801) a plurality of objects in the XR scene captured by at least one camera of the HMD device (1700);determining (1802) at least one static object and at least one dynamic object available in the plurality of objects of the XR scene;detecting (1803) a motion of the HMD device (1700);determining (1804) change in a position of the at least one static object and change in a position of the at least one dynamic object based on the motion of the HMD device (1700) and a plurality of frames of the XR scene; anddetermining (1805) the pose of the user wearing the HMD device (1700) based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object.2.The method of claim 1, wherein determining the pose of the HMD device (1700) based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object comprises:determining a level of change in the position of the at least one static object, wherein the level of change in position is determined based on change in the position of the at least one static object;determining a level of change in the position of the at least one dynamic object, wherein the level of change in position is determined based on change in the position of the at least one dynamic object; anddetermining, the pose of the HMD device (1700) based on the level of the change in the position of the at least one static object and the level of the change in the position of the at least one dynamic object.3.The method of claim 1, comprises:receiving sensor data from motion sensors of the HMD device (1700);determining the motion of the HMD device (1700) based on the sensor data;receiving a plurality of feature points from the plurality of objects in the XR scene captured by the at least one camera of the HMD device;determining by matching each of feature point of the plurality of feature points in the plurality of frames for depth estimation; anddetermining a refined pose of the user based on the motion of the HMD device (1700) and each of feature point in the plurality of feature points based on the estimated depth.4.The method of claim 1, wherein determining the at least one static object and the at least one dynamic object available in the plurality of objects of the XR scene comprises:receiving the plurality of frames of the XR scene;determining a plurality of feature points in the plurality of frames;segmenting the plurality of feature points into the at least one probable static object and the at least one probable dynamic object based on a deep learning model in the plurality of frames; andmarking the plurality of segmented feature points as the at least one probable static object and the at least one probable dynamic object.5.The method of claim 2, wherein determining the level of change in the position of the at least one dynamic object comprises:determining the plurality of objects in the plurality of frames of the XR scene;determining a velocity integration for each object of the plurality of objects;measuring a relative velocity for each object by comparing at least two frames of the plurality of frames;determining the at least one dynamic object based on the relative velocity;determining an inverse of the velocity integration for the at least one dynamic object;determining the level of change in the position of the at least one dynamic object based on the inverse of the velocity integration; andupdating the position of the at least one dynamic object based on the determined level of change in the position.6.The method of claim 1, wherein determining at least one static object and at least one dynamic object available in the plurality of objects of the XR scene, comprises;determining at least one feature count in the plurality of frames, wherein the at least one feature count is a number of trackable features in the plurality of frames, andclassifying the at least one feature as the at least one static object and the at least one dynamic object based on relative velocity of the at least one feature, relative velocity of the at least one feature across frames and relative velocity of the probable static objects.7.The method of claim 1, wherein determining the pose of the of the user wearing HMD based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object relative to HMD, comprises;receiving the motion information of the HMD using a motion sensor in the HMD and the plurality of frames from the HMD;bundling the motion information of the HMD and the plurality of frames from the HMD for each image frame;determining an initial pose of the user wearing the HMD based on the bundled motion of the HMD and the plurality of frames; andperforming a refinement of the pose based on the initial pose of the user wearing the HMD and the plurality of frames from the HMD, wherein the refined pose is determined by minimizing re-projection error of three-dimensional landmarks.8.The method of claim 7, wherein the three-dimensional landmarks are the plurality of feature points from the plurality of objects in the XR scene captured by at least one camera.9.The method of claim 1, wherein detecting the motion of the HMD, comprises;determining an average velocity of the plurality of objects and an average velocity of the at least one camera in the plurality of frames;determining a difference in a location of the plurality of objects based on the determined average velocity of the plurality of the objects and the average velocity of the at least one camera in the plurality of frames; anddetermining a motion compensated frame based on the determined difference in the location of the plurality of objects based on the determined average velocity of the plurality of the objects and the average velocity of the at least one camera in the plurality of frames.10.The method of claim 9, comprises:merging the motion compensated frames with the at least one static object in the plurality of frames, wherein the merged frames are used for the pose estimation.11.The method of claim 10, wherein merging the motion compensated frames with the at least one static object in the plurality of frames, comprises;capturing a plurality of camera rays projected from a center of the camera passing through at least one feature in the plurality of motion compensated frames and corresponding features in previous frames,; anddetermining the at least one depth or location of the feature based on the intersection of plurality of camera raysand at least one depth or location of the feature determined in the motion compsensated frames.12.A head mounted display, HMD, device (1700) for determining pose of a user within an Extended Reality (XR) scene comprises:at least one camera (1701);memory (1703) storing instructions;at least one processor (1704) operably coupled to the camera, the plurality of sensors (1702), and the memory (1703):wherein the instructions, when executed by the at least one processor (1704), cause the HMD device (1700) to:determine a plurality of objects in the XR scene captured by the at least one camera (601) of the HMD device (1700);determine at least one static object and at least one dynamic object available in the plurality of objects of the XR scene;detect a motion of the HMD device (1700);determine change in a position of the at least one static object and change in a position of the at least one dynamic object based on the motion of the HMD device (1700) and a plurality of frames of the XR scene; anddetermine the pose of the of the user wearing the HMD device (1700) based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object.13.The HMD device of claim 12, wherein the instructions, when executed by the at least one processor (1704), cause the HMD device (1700) further to be operated according to a method in one of claims 2 to 11.14.A non-transitory computer readable storage medium storing insturctions which, when executed by at least one processor (1704) of a head mounted display, HMD, device (1600), cause the HMD device (1700) to:determine a plurality of objects in an extended reality, XR, scene captured by at least one camera (1701) of the HMD device (1700);determine at least one static object and at least one dynamic object available in the plurality of objects of the XR scene;detect a motion of the HMD device (1700);determine change in a position of the at least one static object and change in a position of the at least one dynamic object based on the motion of the HMD device (1700) and a plurality of frames of the XR scene; anddetermine the pose of the of the user wearing the HMD device (1700) based on the change in the position of the at least one static object and the change in the position of the at least one dynamic object.15.The non-transitory computer readable storage medium of claim 14, wherein the instructions, when executed by the at least one processor (1704), cause the HMD device (1700) further to be operated according to a method in one of claims 2 to 11.
Citation Information
Patent Citations
Organometallic compound, organic light emitting device including the same and electronic apparatus comprising organic light emitting device
KR1020250043941A
Composition for inhibiting halitosis and enhancing oral hygiene
KR102221423B1
object tracking method for CCTV video by use of Deep Learning object detector
KR102253989B1
Real time multi-object tracking device and method by using global motion
KR102434397B1
Object-Tracking System
US20200053292A1