User indication of location information in shared extended reality environments
By employing radar and eye-tracking technologies to identify user-specified locations in shared XR environments, the solution addresses privacy and regulatory concerns, enabling accurate and seamless integration of digital content within real-world settings.
Patent Information
- Application Number
- PCT/EP2023/085123
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
Existing XR technologies face challenges in accurately identifying a user-specified location in a shared extended reality environment without relying on outward-facing cameras, due to privacy concerns and regulatory restrictions.
The use of radar to detect distances to real-world objects combined with eye-tracking to determine the user's gaze direction, allowing for the selection of a known real-world object and determination of its position as the user-specified location, without the need for cameras.
This approach enables accurate and privacy-compliant placement of digital objects in XR environments, ensuring that the digital content fits seamlessly into the real-world environment, and allows for sharing of this information among multiple users.
Smart Images

Figure EP2023085123_19062025_PF_FP_ABST
Abstract
Description
[0001] USER INDICATION OF LOCATION INFORMATION IN SHARED EXTENDED REALITY ENVIRONMENTS
[0002] BACKGROUND
[0003] The present invention relates to locations in extended reality (XR) environments, more particularly to locations in shared XR environments, and still more particularly to user designation of XR environment location information that is common to multiple users of a shared XR environment.
[0004] Some or all of the following abbreviations are used in this specification:
[0005] Abbreviation Explanation
[0006] Ao A Angle of Arrival
[0007] AR Augmented Reality
[0008] DoA Direction of Arrival
[0009] IMU Inertial Measurement Unit
[0010] MR Mixed Reality
[0011] RADAR Radio Detection and Ranging
[0012] SLAM Simultaneous Localization and Mapping
[0013] SNR Signal to Noise Ratio
[0014] VR Virtual Reality
[0015] XR Extended Reality
[0016] The term “extended reality” (XR) is a generic term that encompasses augmented reality (AR), mixed reality (MR), and virtual reality (VR) technologies. In each of these, a user is able to experience immersion in a computer generated environment. The term VR generally refers to environments that are entirely computer generated, whereas in AR and MR, the user experiences a mixture of a real world environment and one that is computer generated. As an example, a user wearing an AR device while sitting at their desk may be able to see not only real world objects such as their telephone, notepad, and pen but also computer generated (i.e., virtual) digital objects such as a virtual vase with flowers and a virtual computer screen. In short, when using AR / XR devices (e.g., AR / XR headset devices), the user is able to see the surrounding real world environment while the technology adds user-perceivable additional content on top of the real environment.
[0017] Such devices are typically constructed using one or more outward directed cameras to analyze the environment around the user (e.g., for object detection). The devices may also include gaze tracking capabilities to detect where in the environment the user has their attention and eye focus (e.g., by generating so-called gaze heatmaps, which statistically show how much time the user spends looking at each part of the visual environment).
[0018] The above functionalities allow additional artificial objects and / or text to be added to the end user views by, for example, superimposing an image of the object and / or text that is configured to create the illusion that it is located at a certain viewing location and viewing distance based on environmental knowledge that the device has acquired from images captured by the camera. These images are analyzed to determine how to adjust the image so that, when overlay ed on top of the user’s view of the real world, it will appear to exist at the particular location and viewing distance in that real world (e.g., on top of a flat surface).
[0019] Digital objects encompass a wide range of media elements, including 3D models, images, sounds, videos, and interactive assets, which can represent anything from simple static geometric shapes to complex, lifelike moving characters or objects. Digital objects play a crucial role in enriching user experiences, as they provide the building blocks for immersive environments and interactive scenarios. They can be created, manipulated, and shared by users, enabling collaborative and engaging experiences that extend beyond the constraints of the physical world. In the context of the metaverse and AR, digital objects serve as the foundation for various applications, such as gaming, education, commerce, and communication, transforming the way we interact with and perceive the digital realm.
[0020] A user may want to place a digital object within the virtual environment within proximity of themself. Commonly a user would want to place a digital object in a specific location, e.g., for inspecting it or to show the object to other people sharing the virtual space. The location in this context refers to a coordinate in the frame of reference created or used by the user. For AR- glasses this would be a coordinate in the SLAM-built map.
[0021] There are two steps involved:
[0022] 1. Interpreting the intent of the user to place a digital object.
[0023] 2. Determining the coordinate of the intended placement. For some XR-devices, an outwards-facing camera and a microphone can be used to perform these steps, for example, using voice commands for step 1, and pointing with a finger for performing step 2.
[0024] The inventors of the herein-described technology have recognized that conventional technology, such as that which is described above, has limitations and problems. One aspect relates to the fact that there may be reasons why future XR headsets will not be equipped with outward facing cameras. Such reasons may arise from privacy concerns, where persons within proximity of XR devices will not accept being surrounded by constantly active cameras from such devices. There are recorded instances of users wearing XR devices being denied entry to some commercial establishments because the devices represented a form of ubiquitous recording. Similarly, and by analogy, it is believed that there are instances in which photos taken by drones are not allowed to be taken at many places without explicit consent or agreement of the persons photographed, or of the persons who own the animals or property being photographed.
[0025] Furthermore, the European Commission has published a document called, “Proposal for a Regulation Of The European Parliament And Of The Council Laying Down Harmonised Rules On Artificial Intelligence (Artificial Intelligence Act) And Amending Certain Union Legislative Acts.” In the proposal, Al systems are divided into different categories related to their societal risk, with different categories facing different levels of regulations. One category is “Unacceptable Risk” which will be severely regulated or even forbidden. That category includes real-time and remote biometric identification systems, which might include forms of facial recognition systems in public places. Examples of “High Risk” which will be strictly regulated include Surveillance systems (e.g., biometric monitoring for law enforcement, facial recognition systems). One implication of this are proposals for bans of facial recognition in public places. Overall, such regulations might put restrictions on the usage of AR glasses / devices with certain functions in public places.
[0026] The above and other examples illustrate a growing concern for privacy.
[0027] To reach market attraction of XR devices, other solutions not based on outward facing cameras are therefore needed. From a technological perspective, one of the problems to be solved is how to project augmented reality material within XR devices so that the end user experiences the augmentation fitting into the real world environment, and doing so without relying on camera and image recognition / analysis for the projection to fit within the surroundings. This could be challenging when, for example, added text or objects need to fit well at suitable viewing distances and on surfaces. Eye-tracking, using cameras that face towards the eyes, has been used for some time. This technology can determine gaze direction with an accuracy as good as 0.7 degrees. If the device has stereo-camera monitoring of both eyes, the device can perform a distance measurement based on where the gaze directions of the two eyes intersect. However, given the close distance between the eyes, the distance estimate is only fairly good up to about a meter, and at longer distances this gaze-direction accuracy is not good enough for a proper distance estimation. For less expensive systems having only a single camera, where only one eye is monitored, eye tracking fails as a method for accurately estimating distance.
[0028] In view of the foregoing, there is a need for technology that addresses the abovedescribed and related problems, including allowing an XR device to accurately identify a coordinate in a shared XR environment without reliance on an outward-facing camera system in order to, for example, place a digital object at a user-desired location and provide useful information about the user-desired location to other devices that are sharing the XR environment.
[0029] SUMMARY
[0030] It should be emphasized that the terms “comprises” and “comprising”, when used in this specification, are taken to specify the presence of stated features, integers, steps or components; but the use of these terms does not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof.
[0031] Moreover, reference letters may be provided in some instances (e.g., in the claims and summary) to facilitate identification of various steps and / or elements. However, the use of reference letters is not intended to impute or suggest that the so-referenced steps and / or elements are to be performed or operated in any particular order.
[0032] In accordance with one aspect of the present invention, the foregoing and other objects are achieved in technology (e.g., methods, apparatuses, nontransitory computer readable storage media, program means) for identifying a user-specified location in an extended reality environment, wherein the extended reality environment is displayed in a viewing area of an extended reality device. Identifying the user-specified location involves obtaining an ego position, wherein the ego position is a location in a shared set of one or more positions of respective known real world objects in the extended reality environment. A user gaze direction toward the viewing rea is detected. Radar is used to detect respective distances to one or more sensed real world objects (107) in the real world environment. The ego position of the extended reality device, the detected user gaze direction, and the respective distances of the one or more respective sensed real world objects are used to select one of the one or more known real world objects. The shared set of one or more positions of respective known real world objects in the extended reality environment is used to ascertain a selected position of the selected known real world object, and the selected position is used as the user-specified location in the extended reality environment.
[0033] In an aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include using the shared set of one or more positions of respective known real world objects in the extended reality environment to ascertain a pose of the selected known real world object, and using the pose of the selected known real world object in the extended reality environment.
[0034] In another aspect of some but not necessarily all embodiments consistent with the invention, the ego position includes a pose of the extended reality device.
[0035] In yet another aspect of some but not necessarily all embodiments consistent with the invention, the selected known real world object has a corresponding one of the one or more sensed real world objects, and actions involved in identifying the user-specified location in the extended reality environment also include using the distance of the corresponding one of the one or more sensed real world objects to ascertain a pose of the extended reality device with reference to the shared set of one or more positions of respective known real world objects in the extended reality environment. In some but not necessarily all such embodiments, actions involved in identifying the user-specified location in the extended reality environment also include using the distance of the corresponding one of the one or more sensed real world objects to ascertain an updated ego position of the extended reality device with reference to the shared set of one or more positions of respective known real world objects in the extended reality environment.
[0036] In still another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include presenting, in the viewing area of the extended reality device, an indication of the selected position. In some but not necessarily all alternatives, such embodiments include receiving user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting the selected position based on further user input. In some but not necessarily all other alternatives, such embodiments include receiving user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, repeating location-identifying steps. In another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include positioning a virtual object at the user-specified location in the extended reality environment. In some but not necessarily all alternatives of such embodiments, the virtual object has a front portion, and posing the virtual object at the user-specified location in the extended reality environment such that a pose of the virtual object includes the front side of the virtual object facing the extended reality device in the extended reality environment. In some but not necessarily all other alternatives of such embodiments, posing the virtual object at the user-specified location in the extended reality environment such that a pose of the virtual object includes receiving user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting one or more of a pose and a size of the virtual object based on further user input.
[0037] In another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include publishing pose of the virtual object to a second extended reality device that uses the shared set of one or more positions of respective known real world objects in the extended reality environment.
[0038] In yet another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include publishing information about the virtual object to at least one other extended reality device.
[0039] In still another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include receiving a position, a pose, a gaze direction, and an estimated distance to a real world object in the real world environment from each of one or more other extended reality devices. Further, the ego position of the extended reality device, the detected user gaze direction, the respective distances of the one or more respective sensed real world objects, the received positions, the received poses, the received gaze directions, and the received estimated distances are used to identify a common point occupied by the selected one of the one or more known real world objects. In some but not necessarily all alternative embodiments, actions involved in identifying the user-specified location in the extended reality environment also include publishing the common point occupied by the selected one of the one or more known real world objects to the one or more other extended reality devices. And in some but not necessarily all other alternative embodiments, actions involved in identifying the user-specified location in the extended reality environment also include publishing the common point occupied by the selected one of the one or more known real world objects to one or more additional extended reality devices that do not include the one or more other extended reality devices.
[0040] In still another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include initiating performance of location-identifying steps in response to detection of a user command, wherein the user command is one or more of: a voice command; a button or switch activation; a reference hand gesture; a reference sound; a reference movement of one or more eyelids; and a steady user gaze in a reference direction for at least a reference period of time.
[0041] In still another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include receiving a label; and associating the label with the user-specified location. In some but not necessarily all alternative embodiments, receiving the label comprises sensing one or more spoken words and using the one or more spoken words as the label.
[0042] In still another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include repeating location-identifying steps for each of a plurality of locations in the extended reality environment; and using the plurality of locations in the extended reality environment to define corners of a projection area in the extended reality environment.
[0043] In another aspect of some but not necessarily all embodiments consistent with the invention, actions involved in identifying the user-specified location in the extended reality environment also include accessing a 3D occupancy grid map of the extended reality environment; and using the user-specified location in the extended reality environment to adjust a location in the 3D occupancy grid map to prevent the location from being inside an occupied voxel.
[0044] In still another aspect of some but not necessarily all embodiments consistent with the invention, detecting the user gaze direction toward the viewing area comprises monitoring user gaze direction over time, and detecting when the monitored user gaze has been stable for at least a predetermined amount of time.
[0045] In still another aspect of some but not necessarily all embodiments consistent with the invention, the shared set of one or more positions of respective known real world objects in the extended reality environment is a shared map of the one or more positions of the respective known real world objects in the extended reality environment.
[0046] BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The objects and advantages of the invention will be understood by reading the following detailed description in conjunction with the drawings in which:
[0048] Figure 1 illustrates a user who is wearing an XR device configured in the form of a headset, through which the user is able to view portions of an XR environment.
[0049] Figure 2 illustrates an XR device having a viewing area with stereoscopic capability for seeing at least a portion of an XR environment.
[0050] Figure 3 is an illustration of a user’s gaze over left and right portions of a viewing area of an extended reality device, and a corresponding gaze heat map indicating portions of the viewing area that were viewed more than others.
[0051] Figure 4 is a block diagram of a nonlimiting exemplary XR device configured to carry out actions in accordance with the invention.
[0052] Figure 5 is, in one respect, a flowchart of actions performed by a device and / or system for specifying a point in a 3D XR environment in accordance with aspects of some exemplary inventive embodiments.
[0053] Figure 6 is, in one respect, a flowchart of actions performed by a device and / or system for specifying a point in a 3D XR environment in accordance with aspects of some exemplary alternative inventive embodiments.
[0054] Figure 7 is, in one respect, a flowchart of actions performed by an XR device in accordance with some but not necessarily all inventive embodiments
[0055] Figure 8 shows an exemplary controller that may be included in an XR device to cause any and / or all of the herein-described and illustrated actions associated with that device to be performed. DETAILED DESCRIPTION
[0056] The various features of the invention will now be described with reference to the figures, in which like parts are identified with the same reference characters.
[0057] The various aspects of the invention will now be described in greater detail in connection with a number of exemplary embodiments. To facilitate an understanding of the invention, many aspects of the invention are described in terms of sequences of actions to be performed by elements of a computer system or other hardware capable of executing programmed instructions. It will be recognized that in each of the embodiments, the various actions could be performed by specialized circuits (e.g., analog and / or discrete logic gates interconnected to perform a specialized function), by one or more processors programmed with a suitable set of instructions, or by a combination of both. The term “circuitry configured to” perform one or more described actions is used herein to refer to any such embodiment (i.e., one or more specialized circuits alone, one or more programmed processors, or any combination of these). Moreover, the invention can additionally be considered to be embodied entirely within any form of non- transitory computer readable carrier, such as solid-state memory, magnetic disk, or optical disk containing an appropriate set of computer instructions that would cause a processor to carry out the techniques described herein. Thus, the various aspects of the invention may be embodied in many different forms, and all such forms are contemplated to be within the scope of the invention. For each of the various aspects of the invention, any such form of embodiments as described above may be referred to herein as “logic configured to” perform a described action, or alternatively as “logic that” performs a described action.
[0058] Embodiments that include aspects of the invention variously relate to the specification of 3-D points in an area of an XR environment that is viewable by means of an XR device (e.g., without limitation, AR glasses), the placement of one or more virtual objects at the specified locations, and the sharing of those points with others who are using respective other XR devices, so that multiple users virtually located in the same area can see virtual objects in consistent positions. In some but not necessarily all embodiments, information about orientations of respective virtual objects placed by one user are shared with other users so those users experience the virtual objects being aligned within the area as would be expected by those users based on their respective positions and orientations in the shared area.
[0059] In another aspect of at least some embodiments, the above is accomplished with as low overhead required from the user as possible (e.g., to avoid the need for a user to perform complex procedures to learn a 3D point in the surrounding space). To give an overview of technical aspects that are found in various inventive embodiments, the inventors of the herein-described technology have observed that wireless communication devices are becoming more and more advanced. The evolution in radio access protocol functionalities over time has enabled wireless communication using large bandwidths, and the wireless devices typically support several different radio frequency bands - even sometimes operating some of them simultaneously. As a new emerging feature, wireless communication devices may also support radar sensing functionalities utilizing its high-end capabilities of radio signal transmission, reception and signal processing. For example, International Application No. PCT / EP2021 / 08358 describes technology in which a mobile communication device, having for example a radar-enhanced 5G modem, is able to determine its own positional coordinates with support from radar-obtained information about the device’s surroundings, and the ability to do so with very good positioning accuracy and without the need to rely on cameras.
[0060] Inventive embodiments described herein are based on a combined usage of communication connectivity between devices (e.g., and without limitation, 5G cellular connectivity) with radar functionality, and eye-tracking sensors in an XR device (e.g., and without limitation, AR eyewear). In some but not necessarily all embodiments, information from an IMU is also used.
[0061] In known technology that enables mobile devices to determine their own positioning information, this is done by correlating radar images of the device’s surroundings with previously obtained radar images of the same area (similar to how SLAM algorithms are used with visual based sensors) and determining how the correlated objects are related to a detailed stored map of objects in that environment. (If a detailed map does not already exist, one can be created by this process as well.) This also enables the device to determine in which direction it is facing.
[0062] Movement of the device during the positioning process can lead to inaccuracy. To counter this effect, the intermittent movement of the user’s position and head pose can be estimated by dead reckoning, since such glasses typically contain an IMU or accelerometer, and at certain times the position can be re-estimated based on the same method just described, thereby improving accuracy or reducing power with less need for sensing overall.
[0063] In aspects of inventive embodiments, after obtaining a position of him / herself in 3D, the user can gaze towards an object and trigger a “specify point” command. The system then responds by estimating what point in space the user is looking at and then letting that position be the specified point for any subsequent actions.
[0064] In a further aspect of some but not necessarily all inventive embodiments, the device has access to a 3D map of the surrounding environment, and the specified point and related direction of the gaze from the first user are made known to other users in that area who also have access to the 3D map. In this way, the user’s selected location can be shared with others, and the virtual environments of those other users can be adapted accordingly so that all users agree upon the contents of the virtual environment and the locations of those contents.
[0065] In yet another aspect of some but not necessarily all inventive embodiments, in the event that a virtual object is placed at a user-selected point, it is assumed that the orientation of the virtual object is also shared with any other users or stored in a scene descriptor for the area.
[0066] In yet another aspect of some but not necessarily all inventive embodiments, in case the map is incomplete, or the user is looking at a dynamic object that is not present in the map or that has moved from its position in the map, a radar measurement can be used to determine the distance to the object. The distance, combined with the detected gaze direction, provides a relative position to the user. If the user has a known position and orientation in the map, the object’s position can be shared with other users.
[0067] Elaborating further on the above discussion, in embodiments consistent with the invention, the technology enables a user of an XR device to specify points in 3D space using radar and eye-tracking. The eye tracking is used to estimate the gaze direction at the time of a trigger command. The estimated gaze direction is used to specify the point (and any environmental meta data relevant at the time of triggering). Radar scanning is used to detect the presence of physical objects in the XR environment. The device or a related system then estimates the exact point based on the user’s position and the estimation of gaze direction relative to a 3D-map of the area (which may be stored in the device or obtained from an external source), as confirmed and aligned with the radar signals used to sense the environment. In some but not necessarily all embodiments, the user gets confirmation of the identified point by some form of indication in the AR glasses. The user can respond by accepting, rejecting, or modifying the location. This allows the user to improve calibration and accuracy of the technology.
[0068] In addition to the 3D point, the pose / orientation of a virtual object placed at the specified location can be set by default so that, for example, the user will see the front-side of a virtual object projected into the XR environment. This can be specified by, for example, adding the 3D position of the user (at the time of specifying) to the information about the specified point, or by adding a vector to the 3D point, pointing towards the 3D position of the user (at the time of specifying).
[0069] In a further aspect of some but not necessarily all inventive embodiments, if triggered by the user, the calibration procedure can further leverage the radar sensing to interpret hand gestures to refine the 3D point and orientation. In a non-limiting example, this can be done by interpreting an extension of the user’s hand to mean increasing the distance. Similarly, a change in pitch and yaw of the arm, or rotation of the hand can be interpreted as a command to change the orientation of a virtual object.
[0070] A user confirmed specified point (and gaze direction, and orientation) in 3D space can be shared with multiple users in the same area that may then use that information to project the same information accurately oriented and made to mix into the environment in the most suitable way for that application. What constitutes “suitable” is application-dependent. Accordingly, a complete description of this is beyond the scope of the invention.
[0071] For multiple users, a combined and coordinated specification of the 3D point can be enhanced by triangulation of the gaze directions of multiple users in relation to the shared map and by radar-based distance measurements by each device to potential objects in the gaze direction.
[0072] Any device of the other users can decide to use the specified point, gaze direction, and orientation of a shared object by displaying a virtual object in the display (e.g., AR glasses). The user might be able to adjust the specified point by an additional triggering command in combination with a new gaze at the time of that command, thus making the scene become truly interactive and updated in real time since a virtual object can be placed and moved around by any one of the users sharing the XR environment.
[0073] In yet another aspect of some but not necessarily all inventive embodiments, the specified position can be determined more accurately by having multiple users look at the same point at the time that they give the triggering command. Their respective gazes can be used to triangulate a location in relation to the common map of the area.
[0074] In still another aspect of some but not necessarily all inventive embodiments, to support a stable function of the command-based gaze tracking, a timer-based trigger criteria is applied to a “specify point” command. In some examples this can be used to limit activation of new points to only those times when the user’s gaze is stable and non-varying during a longer time period than a defined time value. The defined time value in this case is application-dependent.
[0075] These and further aspects of inventive embodiments are described in the following. Figure 1 illustrates a user 101 who is wearing an XR device 103 configured in the form of a headset, through which the user is able to view portions (depending on direction and pose of the XR device 103) of an XR environment 105. The XR environment 105 includes a real -world (i.e., physical) object 107 having any number of surfaces such as the surface 109. To illustrate aspects of inventive embodiments, suppose the user 101 desires to place a digital (i.e., virtual) object 111 on top of the real -world object 107. In order to create a realistic rendering of the digital object 111 that object’s rendering should be consistent with the user’s expectations with respect to location and size of the digital object 111. In some cases, the pose of the digital object 111 (e.g., direction and tilt angle of one or more of the object’s surfaces) may also be important. In order to configure a correct rendering as just described, it is necessary to know the location and distance of the digital object 111 relative to the XR device 103. Since it is desired to render the digital object 111 on top of the real -world object 107, the location of, and distance 113 to the real -world object 107 can be used as proxies for those characteristics of the digital object 111.
[0076] In one aspect of inventive embodiments, location and distance information can both be obtained by scanning the XR environment 105 with radar signals 115 and receiving radar reflections 117. The direction of arrival (DoA) of the radar reflection 117 will indicate the direction of the real -world object 107, and the time delay between transmission of the radar signal 115 and receipt of the radar reflection 117 corresponds to the round-trip distance (i.e., twice the distance) between the XR device 103 and whichever surface of the real -world object 107 the radar signal 115 reflected off of.
[0077] It is recognized that the XR environment may include more than one real -world object 107, and some number of these may all be sensed by radar. This creates an ambiguity with respect to which real -world object the user wants the digital object 111 to be placed on. To address this technological challenge, the user’s gaze direction is also sensed. In particular, when the user instructs the headset device to create the digital object 111 (e.g., by voice command or by pressing one or more buttons on an input device) the user 101 looks at the location where the digital object 111 is to be rendered. The user’s gaze direction at that moment is sensed and recorded. Then, when multiple real -world objects are sensed by the radar scanning, knowledge of the user’s gaze direction is combined with the detected directions of radar-sensed real -world objects. The real world object having the best matching direction is then selected. Since the direction and distance information of the digital object’s location within the XR environment 105 are now known, the image of digital object can then be adjusted so that its apparent location and apparent size at that location create a realistic user experience that the digital object 111 is actually present in the XR environment 105. To compensate for subsequent movement of the XR device 103, the XR device 103 senses the movement (e.g. by means of gyroscopic or Inertial Measurement Unit (IMU) technology) and continues to adjust the rendering of the digital object 111 so that its perceived presence at the location appears to be fixed.
[0078] Further aspects that are relevant to inventive embodiments are now discussed with reference to Figure 2. In this nonlimiting example, an XR device 203 has a viewing area 221 with stereoscopic capability, so the viewing area 221 has two portions, each dedicated for viewing by a respective one of a pair of user’s eyes 201. By peering at or through the viewing area 221 the user 101 is able to see at least a portion of the XR environment 205.
[0079] In this example, the XR environment 205 includes a real -world surface 209 of an object (e.g., the surface could be a tabletop, desktop, countertop, top surface of a shelf, etc.). Further in this example, the user’s gaze 223 is directed towards a real -world location 225 on top of the surface 209. In this stereoscopic example, gaze monitoring of each eye would detect that the user’s left eye is looking at or through a display location xieft,yieft of the portion of the viewing area 221 allocated for use by the left eye, and that the user’s right eye is looking at or through a display location xright, yright of the portion of the viewing area 221 allocated for use by the right eye. In general, the coordinate pair xieft,yieft is not equal to the coordinate pair xright,yright. The gaze direction 229 in this instance can be determined by finding the intersection of the left and right gaze directions 223. In the example of Figure 2, this is shown as the real world location 225 having x,y,z coordinates (the third coordinate, z, in this case represents depth
[0080] Since digital (virtual) objects do not actually exist in the real world, they do not have real -world locations; they can be seen only in the viewing area 221 of the XR device 203. However, by creating left and right images at respective display locations in the viewing area 221 that line up with the user’s respective left and right gaze directions 223, the user 101 will perceive that the digital object 211 is located at an apparent location 227 in the XR environment 205. For convenience, the convention adopted in the herein-described examples will assign the same reference system for real-world locations and apparent locations, so that the x,y,z coordinates will be the same regardless of whichever one (“real-world” or “apparent”) is mentioned. However, this is by no means an essential aspect of all embodiments. To the contrary, different coordinate systems could be used for real-world and apparent locations. It is also noted that, to further enhance the effect that the digital object 211 will be perceived as being located at the apparent location 227 in the XR environment 205, the light of the digital object 211 should appear to come from the distance corresponding to the apparent location 227, so that the eye will focus that light on the retina when adjusting its lens to that distance, and for that reason an adjustable projection system with lenses can be used in the glasses.
[0081] It is further noted that various aspects found in inventive embodiments are not limited for use only with stereoscopic XR devices. To the contrary, inventive aspects are also present in monoscopic embodiments, as well as in embodiments employing a stereoscopic viewing area 221 but gaze monitoring of only one of the user’s eyes 201. In these latter embodiments, gaze direction 229 is based only on the gaze direction of the one monitored eye.
[0082] In a further aspect, and as mentioned earlier, radar scanning of the XR environment 205 is combined with user gaze monitoring to identify a particular real -world location 225. So, for example as shown in Figure 2, when the user is gazing at the real world location 225, a radar reflection from that point within the environment will indicate to the XR device 203 (via its direction of arrival) a radar reflection direction 231 that will be substantially the same as the detected gaze direction 229. The substantial match between those two measures confirms that the distance that is ascertainable from that same radar reflection is the distance to the real world location 225.
[0083] In another aspect of some but not necessarily all embodiments, the XR environment 205 is shared with one or more other XR devices, such as the other XR device 241 illustrated in Figure 2. The XR environment is generally a 3-dimensional space, so for example the XR device 203 and the other XR device 241 are both able to provide views of, and interactions with elements (e.g., the surface 209 of a real world object or a virtual object 211) in the XR device. However, each of the two XR devices 203, 241 presents an experience of the XR environment 205 from a respectively different perspective, based on the location of the XR device within the XR environment 205. Given this aspect, it is advantageous for locations within the XR environment 205 to be specified with reference to a common frame of reference such as a shared map or more generally a shared set of one or more positions. Since locations characterized in this manner will refer to a same location within the XR environment, they can be used by any of the XR devices that share the XR environment. So, for example, if a user of the XR device 203 specifies a location (real -world location 225 or apparent location 227) and places a virtual object 211 there, information about the location and virtual object can be communicated to the other XR device 241 (e.g., by publishing the information to a server that is accessible to the other XR device 241), and the other XR device 241 can then create the same virtual object 211 at the same location 225, 227 within the XR environment 205, thereby enabling a user of the other XR device 241 to experience and in some embodiments also interact with the virtual object 211. While the location of such a virtual object 211 will be the same for all users, the pose of the virtual object is application dependent. In some instances, it may be desirable for a pose to be constant with respect to the XR environment, so for example, a user of the XR device 203 may see a front side of the virtual object 211 while at the same time, a user of the other XR device 241 may be looking at the virtual object 211 from the side or back. But in other embodiments, it may be desirable for the pose of the shared virtual object 211 to be constant with respect to each user, so that if a user of the XR device 203 is looking at a front side of the virtual object 211, a user of the other XR device 241 also sees the front side of the virtual object. This arrangement can be useful when, for example, the virtual object 211 is a display screen or monitor device.
[0084] Various technologies for monitoring the user’s gaze are known, including technologies that detect particular points of interest within a user’s field of vision, and any such technologies may be employed in embodiments consistent with the invention. By way of example and with reference to Figure 3, the positions of a user’s eyes 301 are monitored over time as they look at respective left and right portions of a viewing area 321 of an extended reality device 303. Through the viewing area 321 the user is able to see a surface 309 upon which are placed three objects 311-1, 311-2, 311-3. The user’s various gaze directions (e.g., user gaze 323-1, 323-2, 323-3, 323-4) are detected and recorded. When a sufficient number of samples are recorded, a so-called “gaze heat map” can be created that shows, for any given point within a field of view, how often the user’s gaze was directed towards that point. Areas of low intensity, such as the point 335, are indicated in some way, and are distinguishable from areas of higher intensity 337 in the gaze heat map, which indicate what the viewer spent the most time looking at. For purposes of example only, in Figure 3 three areas of higher intensity 337 correspond to the locations of the three objects 311-1, 311-2, 311-3. The above-mentioned radar object detections in combination with gaze heat map information (or equivalent) can be used to determine candidate positions for augmented reality text and object overlays.
[0085] Inventive embodiments are characterized by having XR capability, as exemplified without limitation by AR glasses. The various aspects of inventive embodiments do not rely on information obtained from outward facing cameras. Therefore, it is not essential that such devices have outward facing cameras, although their presence is not precluded if they are needed for other purposes.
[0086] Another characteristic of inventive embodiments is gaze tracking capability. For example, the technology for gaze tracking can be camera-based, as is known in the art. Still another characteristic of inventive embodiments is radar functionality. This can be in the form of, for example, specialized radar circuitry. In one of a number of possible alternatives, an inventive device can be equipped with a modem (e.g., a modem configured to support cellular communications) configured to provide the radar functionality, as is known in the art. The modem can also enable communications between the XR device and other devices and / or network services.
[0087] In yet another characteristic of inventive embodiments, the device has the capability of determining its own position in, for example, a frame of reference that is shared with other devices. In this way, position information obtained by one device is meaningful when communicated to other devices in a shared environment. As used herein, the term “ego positioning” refers to the capability of determining one’s own position.
[0088] And still further, inventive embodiments are characterized by having some mechanism by which digital objects (i.e., virtual objects) can be displayed to a user of the device. AR glasses are one type of device having such capability.
[0089] Another feature that is present in some but not necessarily all embodiments is a mechanism by which a user can provide triggers and / or commands to the device to, for example, cause a particular process to be performed. For example, and without limitation, a user might be able to control a device using any one or more of: voice commands and eye commands (e.g., by means of particular blinking patterns).
[0090] To illustrate these aspects further, Figure 4 is a block diagram of a nonlimiting exemplary XR device 401 configured to carry out actions in accordance with the invention. The exemplary XR device 401 includes: an optical unit 403 that includes a viewing area 405 through which a user is able to see a portion of an XR environment. The optical unit is also able to superimpose computer generated digital objects within the viewing area 405.
[0091] A gaze tracker 407 that monitors the gaze directions of one or both of a user’s eyes.
[0092] - Radar circuitry 409 or equivalent radar functionality. As one example of the latter, the XR device 401 may include a modem 411 for wireless communication with, for example, a wireless communication network. Such modems 411 typically operate in frequencies that are suitable for radar operation. Accordingly, the modem 411 can be configured to operate as a radar device that is suitable for use in inventive embodiments. With a suitable radar signal bandwidth (e.g., on the order of 1GHz to give a radar range that permits resolution of two objects located 15cm distance from one another), multiple radar transceivers with a few decimeters distance in-between each other (or a radar transceiver with antenna array 419 as illustrated in Figure 4) can be used to achieve centimeter-level accuracy of distance estimation to objects surrounding the XR device 401, and at mm- wave frequencies, a few centimeters would be enough. The actual distance is frequency / wavelength dependent.
[0093] An inertial measurement unit (IMU) 413 for tracking any movement of the XR device 401. This movement information can be used to stabilize images presented in the view area 405, and can also be used as a basis for determining adjustments to a rendered image of a digital object so that the digital object will appear to have a stable location as a user who is wearing the XR device 401 moves around.
[0094] A microphone array 415 that can serve a number of purposes including but not limited to receiving voiced commands and information from a user.
[0095] A controller 417 for controlling the above-described and other components of the XR device 401. The controller 417 may be configured from hardwired circuitry, programmable software controlled processors / elements, or combination of both.
[0096] - Ego positioning functionality 421, which can be embodied as separate circuitry, or as additional functioning performed by the controller 417.
[0097] The following nonlimiting example illustrates the capabilities of an embodiment in accordance with aspects of the invention:
[0098] Assumptions for determination of link budget for beamforming transmit and angle of arrival receive radar:
[0099] • Frequency of operation: the 60GHz ISM band, wavelength (k)=5mm
[0100] • TX array: 4x4 antenna elements, support 2D beamforming in lcm2
[0101] • RX array: 2x2 antenna elements, support 2D AoA in 5mm x 5mm
[0102] • Antenna element gain: G=3dBi
[0103] • Noise figure: NF=8dB
[0104] • Transmit power: PTx=lmW per antenna
[0105] • Integration time: Tint=lms
[0106] • SNRmin = 20dB
[0107] • Radar cross section: RCS=0.1m2 Calculated results:
[0108] RX sensitivity: PRx=-174dBm-10*log(Tint)+NF+SNRmin=-174+30+8+30dBm=-l 16dBm Maximum range: Rmax=(16*PTx*16*G*G*k2*RCS / (PRx*(4*7t)3))°-25=27m
[0109] It can be seen that, even without beamforming on the receive side, a range of 27m is obtainable. For targets in this range (closer than 27m) the TX beam can be swept while monitoring the angle of arrival of the returning echoes. The direction of small objects can then be determined with high precision. For large objects, however, only a midportion of the illuminated part is detected. It may then be difficult to determine if parts of the object are in the gaze direction of the user. But by performing a transmit beam sweep, it is possible to observe the range of angles of arrival of an object when illuminating different parts, especially the edges of an object. Such observations then make it possible to more accurately determine the angular extension of a large object.
[0110] Further aspects of at least some inventive embodiments will now be described with reference to Figure 5 which, in one respect, is a flowchart of actions performed by a device and / or system for specifying a point in a 3D XR environment. In other respects, the blocks depicted in Figure 5 can also be considered to represent means 500 (e.g., hardwired or programmable circuitry or other processing means) for carrying out the described actions.
[0111] In the exemplary embodiment of Figure 5, a first device specifies a point in 3D space in response to a triggering command issued by the user of the first device. This point can thereafter be shared with other devices having access to a 3D map of the environment which is coherent to that of the first device.
[0112] This point in 3D space can be used as a point of placement of a virtual object in the shared XR environment, and even if the user moves or turns, the object shall be virtually stationary at that place. One nonlimiting example is placement of a virtual hologram of a person on a physical chair in the room, or the placement of a virtual screen at a specific location on a wall.
[0113] As shown in Figure 5, initially, the first device obtains its position (e.g., with respect to a shared 3D map of its XR environment) (step 501). This can be achieved using any form of egopositioning algorithm as are known in the art. The device also has access to a shared 3D map of the environment (which need not be very detailed) (step 503). It should be noted that some embodiments work without access to such a 3D map although the accuracy might be impacted depending on the accuracy of the positioning algorithm, radar beam, and gaze tracking. It is further noted that position in this case typically also includes direction / pose / orientation of the device. In some alternative embodiments, step 501 is performed between steps 507 and 509, which are described below.
[0114] A first user looks at a certain place in 3D space (e.g., at a chair) (step 505). This is the point that the first user intends to specify, and optionally in some embodiments where a virtual object will be positioned if that is called for in the application.
[0115] The first user activates the point specification process, herein called “Specify Point” (step 507). This activation can be by means of any of a number of techniques such as, and without limitation, issuance of a voice command, an eye command such as a predefined pattern of blinking, activation of a button associated with the XR device, and the like. It is advantageous for the first user to continue looking at the point during issuance of this command (and in some embodiments for some minor duration after that, sufficient for the system to act on the command which is typically milliseconds). Here, a timer value can be set to ensure only intentional points are determined.
[0116] When the command is recognized / understood by the XR device / system, it estimates the location of the point specified by the user (step 509) by: determining the user’s gaze direction in relation to the position of the user; and
[0117] - using the radar to determine distance to the likely physical object located in the direction of gaze.
[0118] More particularly, the direction and likely distance to candidate object(s) based on the position and pose of the user are correlated with a 3D map of known objects in the room, and the likely object is identified (i.e., an object having correct / matching direction, correct / matching distance). During this step, the position and pose of the first device can optionally be updated depending on how likely it is that the position and pose from step 501 is not up-to-date.
[0119] In an optional action that need not be included in all embodiments, the first device highlights the estimated point produced in step 509 to the first user in the display of the XR device (step 511). As a response to this optional step, the first user can accept or reject this specified point (step 513, also optional), where a rejection either means that the device is to update the position by some form of delta finetuning (e.g., using hand gestures), or is to reinitiate the complete procedure beginning at least from step 505.
[0120] In another optional action that need not be included in all embodiments, and depending upon whether the first device has any need to do so, the first device shares the identified location with another device (e.g., so that other users can have a 3D point in the room that is coherent with that of the first user, so when they look at the “same” virtual object (e.g. a hologram of another user) they look at the same physical place) (step 515).
[0121] In another optional feature, the data of the specified point that is shared with one or more other devices can include more data than just 3D position, such and without limitation: the position of the user at the time of specifying the point, and / or direction towards the user when specifying, so that projection in a proper orientation can be performed (e.g., with the front side of an object facing towards where the user was located at the time of specifying the location).
[0122] Further aspects of at least some alternative inventive embodiments will now be described with reference to Figure 6 which, in one respect, is a flowchart of actions performed by plural XR devices (e.g., AR glasses) and / or systems for specifying a point in a shared 3D XR environment. In other respects, the blocks depicted in Figure 6 can also be considered to represent means 600 (e.g., hardwired or programmable circuitry or other processing means) for carrying out the described actions.
[0123] In the exemplary embodiment of Figure 6, each one of a number of devices that share a 3D XR environment specifies a point in a shared 3D space that is determined collaboratively by a plurality of XR devices that share that space, in response to a triggering command issued by the user of the first device. This point in 3D space can be used as a point of placement of a virtual object in the shared XR environment, and even if the user moves or turns, the object shall be virtually stationary at that place. One nonlimiting example is placement of a virtual hologram of a person on a physical chair in the room, or the placement of a virtual screen at a specific location on a wall, so that respective users of the devices each see the placed virtual object at the same location in the shared 3D XR environment.
[0124] As shown in Figure 6, each one of a plural number of XR devices obtains its position (e.g., with respect to a shared 3D map of its XR environment) (step 601). This can be achieved using any form of ego-positioning algorithm as are known in the art. Each of the devices also has access to a shared 3D map of the environment (which need not be very detailed) (step 603). It should be noted that some embodiments work without access to such a 3D map although the accuracy might be impacted depending on the accuracy of the positioning algorithm, radar beam, and gaze tracking. It is further noted that position in this case typically also includes direction / pose / orientation of the device.
[0125] The multiple users agree on a certain point in the room and all look at that point (e.g., a chair) (step 605) and give the triggering command (e.g., a voice command “Specify Point”) (step 607). Since there are now multiple devices, each having a position / pose, a gaze direction and an estimated distance to likely object(s), these devices send this data to one device or server which calculates a common point (step 609). In case the collection of data does not lead to a sufficiently coherent point, the users can be requested to reperform the procedure starting with (step 607). In case there are more than two devices, and there is a sufficiently coherent common point for some devices but there are devices whose data contribution does not comply with that point, only those devices / users might be requested to reperform the procedure. Alternatively, the coherent point derived from the other devices is simply shared with outlying devices in a later action 617.
[0126] The estimated location is shared with each of the device users who, after seeing the estimated point highlighted in a viewing area of the device (step 613) can then individually accept or reject that point (step 615). Steps 613 and 615 are optional, and they might depend on whether the estimated point was fully coherent between multiple devices or not (e.g., for how many devices the point was sufficiently coherent, and the degree of coherence between these).
[0127] The users can then share the point with other devices not already having access to it (step 617). This step is also optional.
[0128] Additional aspects of inventive embodiments are now described in the following.
[0129] In one aspect, the “specify point” command can be given in different forms. It can be a voice command by a certain word, or it can be by pressing a button on a handheld control device, or by making a hand gesture, by making a sound (e.g., clapping hands), blinking eyelids in a certain pattern, and the like. Similar command methods can also be used for fine-adjustment of the point or its projection orientation, possibly together with gaze (e.g., using movements with, or levers on, handheld control devices; making hand gestures with empty hands; making voice commands to adjust distance; and the like).
[0130] In another class of embodiments the device is configured to provide for the user to name the specified point so that multiple such points can be maintained and referred to in an unambiguous way. Specification can be done by any number of ways such as, without limitation, a voice command (e.g., “Specify Point - Chair” or “Specify Point - Jim's Avatar”). Then the system can differentiate between several such points when placing virtual objects or sharing the points with other users (e.g., that Jim's Avatar shall be placed at a certain defined point, which might be a certain chair).
[0131] In another class of some but not necessarily all embodiments, more than one point can be specified using any of the above-described techniques, and these multiple points are used in combination to define an area in the XR environment. For instance, a projection area like a board can be defined by specifying the corners. Projections to that board will then fit more closely to the edges of the frame, improving the user quality when watching, for example, a presentation or a movie. A special command can be given (e.g., “specify corners”). The XR device can show the user a polygon representing the area, updated as new corners are specified, until done. For regular rectangular boards mounted horizontally, two diagonal comers can be specified. There can be a special command for that, for example, “specify diagonal corners”. Alternatively, one corner and the size of the board can be interpreted via hand gestures or other user input.
[0132] As the radar can be used to create a 3D occupancy grid map (i.e., a map that represents the probability of a voxel (3D version of pixel, i.e., a small volume grid point) being occupied). In many embodiments, the device is configured to correct the user-specified point in 3D space by, for example, ensuring that any placement of a digital object will not be inside an occupied voxel. In a nonlimiting example, such a correction can involve placing the digital object on top of an occupied voxel via vertical projection. This reduction of the degrees of freedom of the estimation of the 3D point will increase the accuracy as well as conform better to the expected behavior of the invention, thus improving the user experience. Similarly for the projection area, the occupancy grid can be used to correct the placement of the virtual screen to remove or minimize the area that is obstructed (e.g., not placing the lowest part of a virtual screen below the table or behind a pillar or other obstruction in the room).
[0133] Still further aspects of some but not necessarily all inventive embodiments will now be described with reference to Figure 7, which in one respect is a flowchart of actions performed by an XR device in accordance with some but not necessarily all inventive embodiments. In other respects, the blocks depicted in Figure 7 can also be considered to represent means 700 (e.g., hardwired or programmable circuitry or other processing means) for carrying out the described actions.
[0134] An XR device in this exemplary class of embodiments is configured to identify a user- specified location in an extended reality environment, wherein the extended reality environment is displayed in a viewing area of an extended reality device. Operation of the device includes obtaining an ego position (step 701), wherein the ego position is a location in a shared set of one or more positions of respective known real world objects in the extended reality environment.
[0135] A user gaze direction toward the viewing area (step 703) is detected.
[0136] Before, after, or concurrent with step 703, radar is used to detect respective distances to one or more sensed real world objects in the real world environment (step 705). The ego position of the extended reality device, the detected user gaze direction, and the respective distances of the one or more respective sensed real world objects are used to select one of the one or more known real world objects (step 707).
[0137] The shared set of one or more positions of respective known real world objects in the extended reality environment are used to ascertain a selected position of the selected known real world object, and using the selected position as the user-specified location in the extended reality environment (step 709).
[0138] 6 Further aspects of embodiments consistent with the invention will now be described with reference to Figure 8, which shows an exemplary controller 801 that may be included in an XR device or system to cause any and / or all of the herein-described and illustrated actions associated with that device or system to be performed. In particular, the controller 801 includes circuitry configured to carry out any one or any combination of the various functions described herein. Such circuitry could, for example, be entirely hard-wired circuitry (e.g., one or more Application Specific Integrated Circuits - “ASICs”). Depicted in the exemplary embodiment of Figure 8, however, is programmable circuitry, comprising a processor 803 coupled to one or more memory devices 805 (e.g., Random Access Memory, Magnetic Disc Drives, Optical Disk Drives, Read Only Memory, etc.) and to an interface 807 that enables bidirectional communication with other elements of a device as described above. A complete list of possible other elements is beyond the scope of this description.
[0139] The memory device(s) 805 store program means 809 (e.g., a set of processor instructions) configured to cause the processor 803 to control other device elements so as to carry out any of the aspects described herein. The memory device(s) 805 may also store data (not shown) representing various constant and variable parameters as may be needed by the processor 803 and / or as may be generated when carrying out its functions such as those specified by the program means 809.
[0140] Various embodiments that are consistent with the invention provide a number of benefits and advantages over conventional technology. Some of these advantages include provision of a simple and efficient technique for specifying a point in a shared 3D XR environment based on detection of a user’s gaze direction.
[0141] In some embodiments a further advantage is the ability to specify an orientation / pose of virtual objects placed in 3D space.
[0142] Yet another benefit and advantage is the ability to share that information among multiple users. Yet another benefit and advantage is the ability to share that information among multiple users without the need to rely on externally facing cameras.
[0143] The invention has been described with reference to particular embodiments. However, it will be readily apparent to those skilled in the art that it is possible to embody the invention in specific forms other than those of the embodiment described above. Thus, the described embodiments are merely illustrative and should not be considered restrictive in any way. The scope of the invention is further illustrated by the appended claims, rather than only by the preceding description, and all variations and equivalents which fall within the range of the claims are intended to be embraced therein.
Claims
CLAIMS:
1. A method of identifying a user-specified location (225) in an extended reality environment (105, 205), wherein the extended reality environment (105, 205) is displayed in a viewing area (221) of an extended reality device (103, 203, 401), the method comprising: a) obtaining (501, 601, 701) an ego position of the extended reality device (103, 203, 401), wherein the ego position is a location in a shared set of one or more positions of respective known real world objects in the extended reality environment (105, 205); b) detecting (505, 703) a user gaze direction (231) toward the viewing area (221); c) using radar (703) to detect respective distances (113) to one or more sensed real world objects (107) in the real world environment; d) using (707) the ego position of the extended reality device (103, 203, 401), the detected user gaze direction (231), and the respective distances (113) of the one or more respective sensed real world objects (107) to select one of the one or more known real world objects (107); and e) using (709) the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) to ascertain a selected position (225) of the selected known real world object (107), and using (709) the selected position as the user-specified location (225) in the extended reality environment (105, 205).
2. The method of claim 1, comprising: using the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) to ascertain a pose of the selected known real world object (107), and using the pose of the selected known real world object (107) in the extended reality environment (105, 205).
3. The method of any one of the previous claims, wherein the ego position includes a pose of the extended reality device (103, 203, 401).
4. The method of either one of claims 1 and 2, wherein the selected known real world object (107) has a corresponding one of the one or more sensed real world objects (107), and wherein the method comprises: using the distance (113) of the corresponding one of the one or more sensed real world objects (107) to ascertain a pose of the extended reality device (103, 203, 401) with reference tothe shared set of one or more positions of respective known real world objects (107) in the extended reality environment.
5. The method of claim 4, comprising: using the distance (113) of the corresponding one of the one or more sensed real world objects (107) to ascertain an updated ego position of the extended reality device (103, 203, 401) with reference to the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205).
6. The method of any one of the previous claims, comprising: presenting, in the viewing area (221) of the extended reality device (103, 203, 401), an indication of the selected position.
7. The method of claim 6, comprising: receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting (513, 615) the selected position based on further user input.
8. The method of claim 6, comprising: receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, repeating steps b) through e).
9. The method of any one of the previous claims, comprising: positioning a virtual object (111, 211) at the user-specified location (225) in the extended reality environment (105, 205).
10. The method of claim 9, wherein the virtual object (111, 211) has a front portion, and wherein the method comprises: posing the virtual object (111, 211) at the user-specified location (225) in the extended reality environment (105, 205) such that a pose of the virtual object (111, 211) includes the front side of the virtual object (111, 211) facing the extended reality device (103, 203, 401) in the extended reality environment (105, 205).
11. The method of claim 9, comprising: receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting one or more of a pose and a size of the virtual object (111, 211) based on further user input.
12. The method of claim 10 or claim 11, comprising: publishing (515, 617) the pose of the virtual object (111, 211) to a second extended reality device (241) that uses the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205).
13. The method of any one of claims 9 through 12, comprising: publishing (515, 617) information about the virtual object (111, 211) to at least one other extended reality device (241).
14. The method of any one of the previous claims, comprising: receiving a position, a pose, a gaze direction (231), and an estimated distance (113) to a real world object (107) in the real world environment from each of one or more other extended reality devices (241); and wherein step d) comprises using (707) the ego position of the extended reality device (103, 203, 401), the detected user gaze direction (231), the respective distances (113) of the one or more respective sensed real world objects (107), the received positions, the received poses, the received gaze directions (231), and the received estimated distances (113) to identify a common point occupied by the selected one of the one or more known real world objects (107).
15. The method of claim 14, comprising: publishing (515, 617) the common point occupied by the selected one of the one or more known real world objects (107) to the one or more other extended reality devices (241).
16. The method of claim 15, comprising: publishing (515, 617) the common point occupied by the selected one of the one or more known real world objects (107) to one or more additional extended reality devices that do not include the one or more other extended reality devices (241).
17. The method of any one of the previous claims comprising: initiating (507, 607) performance of steps a) through e) in response to detection of a user command, wherein the user command is one or more of: a voice command; a button or switch activation; a reference hand gesture; a reference sound; a reference movement of one or more eyelids; and a steady user gaze in a reference direction for at least a reference period of time.
18. The method of any one of the previous claims, comprising: receiving a label; and associating the label with the user-specified location (225).
19. The method of claim 18, wherein receiving the label comprises sensing one or more spoken words and using the one or more spoken words as the label.
20. The method of any one of the previous claims, comprising: repeating steps a) through e) for each of a plurality of locations in the extended reality environment (105, 205); and using the plurality of locations in the extended reality environment (105, 205) to define corners of a projection area in the extended reality environment (105, 205).
21. The method of any one of the previous claims, comprising: accessing a 3D occupancy grid map of the extended reality environment (105, 205); and using the user-specified location in the extended reality environment (105, 205) to adjust a location in the 3D occupancy grid map to prevent the location from being inside an occupied voxel.
22. The method of any one of the previous claims, wherein detecting (505, 703) the user gaze direction (231) toward the viewing area (221) comprises: monitoring user gaze direction (231) over time, and detecting when the monitored user gaze has been stable for at least a reference amount of time.
23. The method of any one of the previous claims, wherein the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) is a shared map of the one or more positions of the respective known real world objects (107) in the extended reality environment (105, 205).
24. A computer program (809) comprising instructions that, when executed by at least one processor (803), causes the at least one processor (803) to carry out the method according to any one of claims 1 through 23.
25. A carrier comprising the computer program (809) of claim 24, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a non-transitory computer readable storage medium (805).
26. An apparatus for identifying a user-specified location (225) in an extended reality environment (105, 205), wherein the extended reality environment (105, 205) is displayed in a viewing area (221) of an extended reality device (103, 203, 401), wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: a) obtaining (501, 601, 701) an ego position of the extended reality device (103, 203, 401), wherein the ego position is a location in a shared set of one or more positions of respective known real world objects in the extended reality environment (105, 205); b) detecting (505, 703) a user gaze direction (231) toward the viewing area (221); c) using radar (703) to detect respective distances (113) to one or more sensed real world objects (107) in the real world environment; d) using (707) the ego position of the extended reality device (103, 203, 401), the detected user gaze direction (231), and the respective distances (113) of the one or more respective sensed real world objects (107) to select one of the one or more known real world objects (107); and e) using (709) the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) to ascertain a selected position (225) of the selected known real world object (107), and using (709) the selected position as the user-specified location (225) in the extended reality environment (105, 205).
27. The apparatus of claim 26, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: using the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) to ascertain a pose of the selected known real world object (107), and using the pose of the selected known real world object (107) in the extended reality environment (105, 205).
28. The apparatus of any one of claims 26 through 27, wherein the ego position includes a pose of the extended reality device (103, 203, 401).
29. The apparatus of either one of claims 26 and 27, wherein the selected known real world object (107) has a corresponding one of the one or more sensed real world objects (107), and wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: using the distance (113) of the corresponding one of the one or more sensed real world objects (107) to ascertain a pose of the extended reality device (103, 203, 401) with reference to the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205).
30. The apparatus of claim 29, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: using the distance (113) of the corresponding one of the one or more sensed real world objects (107) to ascertain an updated ego position of the extended reality device (103, 203, 401) with reference to the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205).
31. The apparatus of any one of claims 26 through 30, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: presenting, in the viewing area (221) of the extended reality device (103, 203, 401), an indication of the selected position.
32. The apparatus of claim 31, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform:receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting (513, 615) the selected position based on further user input.
33. The apparatus of claim 31, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, repeating steps b) through e).
34. The apparatus of any one of claims 26 through 33, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: positioning a virtual object (111, 211) at the user-specified location (225) in the extended reality environment (105, 205).
35. The apparatus of claim 34, wherein the virtual object (111, 211) has a front portion, and wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: posing the virtual object (111, 211) at the user-specified location (225) in the extended reality environment (105, 205) such that a pose of the virtual object (111, 211) includes the front portion of the virtual object (111, 211) facing the extended reality device (103, 203, 401) in the extended reality environment (105, 205).
36. The apparatus of claim 34, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: receiving (513, 615) user input that indicates a rejection of the selected position; and in response to the rejection of the selected position, adjusting one or more of a pose and a size of the virtual object (111, 211) based on further user input.
37. The apparatus of claim 35 or claim 36, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: publishing (515, 617) the pose of the virtual object (111, 211) to a second extended reality device (241) that uses the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205).
38. The apparatus of any one of claims 34 through 37, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: publishing (515, 617) information about the virtual object (111, 211) to at least one other extended reality device (241).
39. The apparatus of any one of claims 26 through 38, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: receiving a position, a pose, a gaze direction (231), and an estimated distance (113) to a real world object (107) in the real world environment from each of one or more other extended reality devices (241); and wherein step d) comprises using (707) the ego position of the extended reality device (103, 203, 401), the detected user gaze direction (231), the respective distances (113) of the one or more respective sensed real world objects (107), the received positions, the received poses, the received gaze directions (231), and the received estimated distances (113) to identify a common point occupied by the selected one of the one or more known real world objects (107).
40. The apparatus of claim 39, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: publishing (515, 617) the common point occupied by the selected one of the one or more known real world objects (107) to the one or more other extended reality devices (241).
41. The apparatus of claim 40, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: publishing (515, 617) the common point occupied by the selected one of the one or more known real world objects (107) to one or more additional extended reality devices that do not include the one or more other extended reality devices (241).
42. The apparatus of any one of claims 26 through 41, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: initiating (507, 607) performance of steps a) through e) in response to detection of a user command, wherein the user command is one or more of: a voice command;a button or switch activation; a reference hand gesture; a reference sound; a reference movement of one or more eyelids; and a steady user gaze in a reference direction for at least a reference period of time.
43. The apparatus of any one of claims 26 through 42, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: receiving a label; and associating the label with the user-specified location (225).
44. The apparatus of claim 43, wherein receiving the label comprises sensing one or more spoken words and using the one or more spoken words as the label.
45. The apparatus of any one of claims 26 through 44, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: repeating steps a) through e) for each of a plurality of locations in the extended reality environment (105, 205); and using the plurality of locations in the extended reality environment (105, 205) to define corners of a projection area in the extended reality environment (105, 205).
46. The apparatus of any one of claims 26 through 45, wherein the apparatus is configured to cause the extended reality device (103, 203, 401) to perform: accessing a 3D occupancy grid map of the extended reality environment (105, 205); and using the user-specified location in the extended reality environment (105, 205) to adjust a location in the 3D occupancy grid map to prevent the location from being inside an occupied voxel.
47. The apparatus of any one of claims 26 through 46, wherein detecting (505, 703) the user gaze direction (231) toward the viewing area (221) comprises: monitoring user gaze direction (231) over time, and detecting when the monitored user gaze has been stable for at least a reference amount of time.
48. The apparatus of any one of claims 26 through 47, wherein the shared set of one or more positions of respective known real world objects (107) in the extended reality environment (105, 205) is a shared map of the one or more positions of the respective known real world objects (107) in the extended reality environment (105, 205).
Citation Information
Patent Citations
Systems and methods for virtual whiteboards
US20220255974A1
EP2021008358W