Systems and methods for multi-user virtual and augmented reality - Patents.com

By determining anchor points aligned with the physical environment and using these points to position virtual content, the method addresses the challenge of accurately placing virtual objects in MR systems, preventing offset and drift and enhancing immersive interactions.

JP7679395B2Active Publication Date: 2025-05-19MAGIC LEAP INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022554528
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-13
Filing Date
2021-03-13
Publication Date
2025-05-19
Estimated Expiration
2041-03-13

AI Technical Summary

Technical Problem

Current optical systems in Mixed Reality (MR) struggle to accurately place virtual objects relative to the physical environment, especially when the objects are far from the user, leading to offset and drift issues.

Method used

The method involves determining anchor points aligned with the physical environment and using these anchor points to position virtual content, allowing for accurate placement of virtual objects even when they are far from the user. This is achieved by communicating between display screens worn by users and a processing unit that determines common anchor points for multiple users, enabling synchronized virtual content experiences.

Benefits of technology

This approach effectively prevents offset and drift of virtual objects, allowing for accurate and immersive interactions in MR environments, even when users are far apart, enhancing gaming and collaborative applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679395000001
    Figure 0007679395000001
  • Figure 0007679395000002
    Figure 0007679395000002
  • Figure 0007679395000003
    Figure 0007679395000003
Patent Text Reader

Abstract

1. An apparatus for providing virtual content within an environment in which a first and a second user may interact with each other, the apparatus comprising: a communication interface configured to communicate with a first display screen worn by a first user and / or a second display screen worn by a second user; and a processing unit configured to: acquire a first position of the first user; determine a first set of anchor points based on the first position of the first user; acquire a second position of a second user; determine a second set of anchor points based on the second position of the second user; determine one or more common anchor points within both the first set and the second set; and provide the virtual content for experience by the first user and / or the second user based on at least one of the one or more common anchor points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computing, learning network configurations, and connected mobile computing systems, methods, and configurations, and more particularly to mobile computing systems, methods, and configurations characterized by at least one wearable component that can be utilized for virtual and / or augmented reality operations.

Background Art

[0002] Modern computing and display technologies have facilitated the development of so-called "Mixed Reality" (MR) systems for "Virtual Reality" (VR) or "Augmented Reality" (AR) experiences, where digitally reproduced images or portions thereof are presented to a user in a manner that appears or is perceived to be real. VR scenarios typically involve the presentation of digital or virtual image information without transparency to other real-world visual inputs. AR scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user (i.e., transparency to real-world visual inputs). Thus, AR scenarios involve the presentation of digital or virtual image information with transparency to real-world visual inputs.

[0003] MR systems can generate and display color data, which increases the realism of MR scenarios. Many of these MR systems display color data by sequentially projecting sub-images within different (e.g., primary) colors or "fields" (e.g., red, green, and blue) corresponding to a color image at a high speed and in succession. Projecting the color sub-images at a sufficiently high rate (e.g., 60 Hz, 120 Hz, etc.) can result in a smooth color MR scenario for the user's memory.

[0004] Various optical systems generate images, including color images, at various depths for displaying MR (VR and AR) scenarios.

[0005] The MR system may employ at least a wearable display device (e.g., a head-mounted display, a helmet-mounted display, or smart glasses) that is loosely coupled to the user's head and thus moves as the user's head moves. If the user's head movement is detected by the display device, the displayed data may be updated (e.g., "warped") to account for changes in the head pose (i.e., the orientation and / or location of the user's head).

[0006] As an example, if a user wearing a head-mounted display device views a virtual representation of a virtual object on the display and walks around the perimeter of the area where the virtual object appears, the virtual object can be rendered per viewpoint and give the user the perception of walking around the perimeter of an object that occupies real space. If the head-mounted display device is used to present multiple virtual objects, the measurement of the head pose can be used to render the scene to match the user's dynamically changing head pose and provide an increased sense of immersion.

[0007] A head-mounted display device that enables AR provides simultaneous viewing of both real and virtual objects. By using an "optical see-through" display, the user can see through a transparent (e.g., translucent or fully transparent) element within the display system and directly view the light from real objects in the environment. The transparent element is often referred to as a "combiner" that superimposes light from the display across the user's view of the real world, and the light from the display projects an image of virtual content across the see-through view of real objects in the environment. A camera may be mounted on the head-mounted display device to capture an image or video of the scene being viewed by the user.

[0008] Current optical systems, such as those in MR systems, optically render virtual content. The content is "virtual" in that it does not correspond to an actual physical object located at an individual position in space. Instead, the virtual content only exists within the user's brain (e.g., the visual cortex) of a head-mounted display device when stimulated by a light beam directed at the user's eye.

[0009] In some cases, a head-mounted image display device may be able to display virtual objects with respect to the real environment and / or enable a user to place and / or manipulate virtual objects with respect to the real environment. In such cases, the image display device may be configured to locate the user with respect to the real environment such that the virtual objects can be correctly displaced with respect to the real environment.

[0010] An eye display for mixed reality or augmented reality desirably has a lightweight, low cost, small form factor, has a wide virtual image field of view, and is as transparent as possible. In addition, it is desirable to have a configuration that presents virtual image information within a plurality of focal planes (e.g., two or more) so as not to exceed an acceptable tolerance regarding vergence-accommodation mismatch and be practical for various use cases.

[0011] Also, it would be desirable to have a new technique for providing a virtual object to the user's view such that the virtual object can be accurately placed with respect to the physical environment when viewed by the user. In some cases, when the virtual object is virtually placed with respect to a physical environment that is located far from the user, the virtual object may be offset or "drift" away from its intended location. This can occur because the local coordinate frame for the user is correctly aligned with features within the physical environment but cannot be accurately aligned with other features within the physical environment that are further away from the user. Summary of the Invention Means for Solving the Problems

[0012] Methods and apparatus for providing virtual content, such as virtual objects, for display on one or more screens of one or more image display devices (worn by one or more users) are described herein. In some embodiments, the virtual content may be displayed so as to appear to be within the physical environment when viewed by the user through the screen. The virtual content may be provided based on one or more anchor points that are aligned with the physical environment. In some embodiments, the virtual content may be provided as a movable object, and the position of the movable object may be based on one or more anchor points that are proximate to an action of the movable object. This enables the object to be virtually and accurately placed relative to the user, even when the object is far from the user (when viewed by the user through the screen worn by the user). In a gaming application, such a feature may enable multiple users to interact with the same object, even when the users are relatively far apart. For example, in a gaming application, the virtual object may be virtually passed back and forth between users. The placement (positioning) of the virtual object based on the anchor point proximity described herein prevents problems of offset and drift, and thus enables the virtual object to be accurately positioned.

[0013] In an environment in which a first user and a second user can interact with each other, an apparatus for providing virtual content includes a communication interface configured to communicate with a first display screen worn by the first user and / or a second display screen worn by the second user, and a processing unit configured to obtain a first position of the first user, determine a first set of one or more anchor points based on the first position of the first user, obtain a second position of the second user, determine a second set of one or more anchor points based on the second position of the second user, determine one or more common anchor points that are within both the first set and the second set, and provide virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points.

[0014] Optionally, one or more common anchor points comprise a plurality of common anchor points, and the processing unit is configured to select a subset of the common anchor points from the plurality of common anchor points.

[0015] Optionally, the processing unit is configured to select a subset of the common anchor points and reduce a positioning error of the first user and the second user relative to each other.

[0016] Optionally, one or more common anchor points comprise a single common anchor point.

[0017] Optionally, the processing unit is configured to position and / or orient the virtual content based on at least one of the one or more common anchor points.

[0018] Optionally, each of the one or more anchor points within the first set is a point within a Persistent Coordinate Frame (PCF).

[0019] Optionally, the processing unit is configured to provide the virtual content as a movable virtual object within the first display screen and / or the second display screen for display.

[0020] Optionally, the processing unit is configured to provide the virtual object within the first display screen for display such that the virtual object appears to be moving within the space between the first user and the second user.

[0021] Optionally, one or more common anchor points include a first common anchor point and a second common anchor point, and the processing unit is configured to provide the movable virtual object within the first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, wherein the first object position of the movable virtual object is based on the first common anchor point and the second object position of the movable virtual object is based on the second common anchor point.

[0022] Optionally, the processing unit is configured to select a first common anchor point for placing the virtual object at the first object position based on the location where the action of the virtual object occurs.

[0023] Optionally, one or more common anchor points include a single common anchor point, and the processing unit is configured to provide the movable virtual object within the first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, wherein the first object position of the movable virtual object is based on the single common anchor point and the second object position of the movable virtual object is based on the single common anchor point.

[0024] Optionally, one or more common anchor points are provided, and the processing unit is configured to select one of the common anchor points for placing virtual content within the first display screen.

[0025] Optionally, the processing unit is configured to select one of the common anchor points for placing virtual content by selecting one of the common anchor points that is closest to the action of the virtual content or within a distance threshold from the action of the virtual content.

[0026] Optionally, the position and / or movement of the virtual content is controllable by the first handheld device of the first user.

[0027] Optionally, the position and / or movement of the virtual content is also controllable by the second handheld device of the second user.

[0028] Optionally, the processing unit is configured to localize the first user and the second user with respect to the same mapping information based on one or more common anchor points.

[0029] Optionally, the processing unit is configured to display the virtual content on the first display screen such that the virtual content will appear in a spatial relationship to physical objects within the surrounding environment of the first user.

[0030] Optionally, the processing unit is configured to obtain one or more sensor inputs, and based on the one or more sensor inputs, the processing unit is configured to assist the first user in performing an objective involving the virtual content.

[0031] Optionally, the one or more sensor inputs indicate the gaze direction, upper limb kinematics, body position, body orientation of the first user, or any combination of the foregoing.

[0032] Optionally, the processing unit is configured to assist the first user in accomplishing the goal by applying one or more limits regarding the position and / or angular velocity of the system component.

[0033] Optionally, the processing unit is configured to assist the first user in accomplishing the goal by gradually reducing the distance between the virtual content and another element.

[0034] Optionally, the processing unit comprises a first processing part that communicates with the first display screen and a second processing part that communicates with the second display screen.

[0035] A method implemented by an apparatus configured to provide virtual content in an environment in which a first user wearing the first display screen and a second user wearing the second display screen can interact with each other includes obtaining a first position of the first user, determining a first set of one or more anchor points based on the first position of the first user, obtaining a second position of the second user, determining a second set of one or more anchor points based on the second position of the second user, determining one or more common anchor points that are within both the first set and the second set, and providing virtual content for the experience of the first user and / or the second user based on at least one of the one or more common anchor points.

[0036] Optionally, the one or more common anchor points comprise a plurality of common anchor points, and the method further includes selecting a subset of the common anchor points from the plurality of common anchor points.

[0037] Optionally, a subset of the common anchor points is selected to reduce the relative positioning error between the first user and the second user.

[0038] Optionally, one or more common anchor points comprise a single common anchor point.

[0039] Optionally, the method further includes determining a position and / or orientation of virtual content based on at least one of one or more common anchor points.

[0040] Optionally, each of one or more anchor points within the first set is a point within a persistent coordinate frame (PCF).

[0041] Optionally, the virtual content is provided as a movable virtual object within the first display screen and / or the second display screen for display.

[0042] Optionally, the virtual object is provided within the first display screen for display such that the virtual object appears to be moving within the space between the first user and the second user.

[0043] Optionally, one or more common anchor points comprise a first common anchor point and a second common anchor point, and the movable virtual object is provided within the first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, the first object position of the movable virtual object is based on the first common anchor point, and the second object position of the movable virtual object is based on the second common anchor point.

[0044] Optionally, the method further includes selecting a first common anchor point for placing the virtual object at a first object position based on where the action of the virtual object occurs.

[0045] Optionally, one or more common anchor points comprise a single common anchor point, and the movable virtual object is provided within a first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, wherein the first object position of the movable virtual object is based on the single common anchor point and the second object position of the movable virtual object is based on the single common anchor point.

[0046] Optionally, one or more common anchor points comprise a plurality of common anchor points, and the method further includes selecting one of the common anchor points for placing the virtual content within the first display screen.

[0047] Optionally, the act of selecting includes selecting one of the common anchor points that is closest to the action of the virtual content or within a distance threshold from the action of the virtual content.

[0048] Optionally, the position and / or movement of the virtual content is controllable by a first handheld device of a first user.

[0049] Optionally, the position and / or movement of the virtual content is also controllable by a second handheld device of a second user.

[0050] Optionally, the method further includes localizing the first user and the second user relative to the same mapping information based on one or more common anchor points.

[0051] Optionally, the method further includes displaying virtual content on a first display screen such that the virtual content appears in a spatial relationship with a physical object within the surrounding environment of the first user.

[0052] Optionally, the method further includes obtaining one or more sensor inputs and assisting the first user in performing an objective with the virtual content based on the one or more sensor inputs.

[0053] Optionally, the one or more sensor inputs indicate the gaze direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing of the first user.

[0054] Optionally, the act of assisting the first user in performing an objective includes applying one or more limits regarding the position and / or angular velocity of a system component.

[0055] Optionally, the act of assisting the first user in performing an objective includes gradually reducing the distance between the virtual content and another element.

[0056] Optionally, the apparatus includes a first processing portion that communicates with the first display screen and a second processing portion that communicates with the second display screen.

[0057] A processor-readable non-transitory medium stores a set of instructions, the execution of which by a processing unit causes a method to be performed, the processing unit being part of an apparatus configured to provide virtual content in an environment in which a first user and a second user can interact with each other, the method comprising obtaining a first position of the first user, determining a first set of one or more anchor points based on the first position of the first user, obtaining a second position of the second user, determining a second set of one or more anchor points based on the second position of the second user, determining one or more common anchor points that are within both the first set and the second set, and providing virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points.

[0058] Additional and other objects, features, and advantages of the present disclosure are set forth in the detailed description, the drawings, and the claims. This specification also provides, for example, the following items. (Item 1) An apparatus for providing virtual content within an environment in which a first user and a second user can interact with each other, the apparatus comprising: A communication interface configured to communicate with a first display screen worn by the first user and / or a second display screen worn by the second user; A processing unit, the processing unit being configured to: Obtain a first position of the first user; Determine a first set of one or more anchor points based on the first position of the first user; Obtain a second position of the second user; Determine a second set of one or more anchor points based on the second position of the second user; Determine one or more common anchor points that are within both the first set and the second set; Provide the virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points And a processing unit configured to perform the above. An apparatus comprising the above. (Item 2) The apparatus according to item 1, wherein the one or more common anchor points comprise a plurality of common anchor points, and the processing unit is configured to select a subset of the common anchor points from the plurality of common anchor points. (Item 3) The apparatus according to item 2, wherein the processing unit is configured to select a subset of the common anchor points and reduce a positioning error of the first user and the second user relative to each other. (Item 4) The apparatus according to item 1, wherein the one or more common anchor points comprise a single common anchor point. (Item 5) The apparatus according to item 1, wherein the processing unit is configured to position and / or orient the virtual content based on at least one of the one or more common anchor points. (Item 6) The apparatus according to item 1, wherein each of the one or more anchor points within the first set is a point within a Persistent Coordinate Frame (PCF). (Item 7) The apparatus according to item 1, wherein the processing unit is configured to provide the virtual content as a movable virtual object within the first display screen and / or the second display screen for display. (Item 8) The apparatus according to item 7, wherein the processing unit is configured to provide the virtual object within the first display screen for display such that the virtual object appears to be moving within the space between the first user and the second user. (Item 9) The one or more common anchor points include a first common anchor point and a second common anchor point. The processing unit is configured to provide the movable virtual object within the first display screen for display such that the movable virtual object has a first object position with respect to the first display screen and a second object position with respect to the first display screen. The first object position of the movable virtual object is based on the first common anchor point. The second object position of the movable virtual object is based on the second common anchor point. The apparatus according to item 7. (Item 10) The apparatus according to item 9, wherein the processing unit is configured to select the first common anchor point for installing the virtual object at the first object position based on the location where the action of the virtual object occurs. (Item 11) The one or more common anchor points include a single common anchor point. The processing unit is configured to provide the movable virtual object within the first display screen for display such that the movable virtual object has a first object position with respect to the first display screen and a second object position with respect to the first display screen. The first object position of the movable virtual object is based on the single common anchor point. The second object position of the movable virtual object is based on the single common anchor point. The apparatus according to item 7. (Item 12) The one or more common anchor points comprise a plurality of common anchor points, and the processing unit is configured to select one of the common anchor points for placing the virtual content within the first display screen, the apparatus according to item 1. (Item 13) The processing unit is configured to select one of the common anchor points for placing the virtual content by selecting one of the common anchor points that is closest to the action of the virtual content or within a distance threshold from the action of the virtual content, the apparatus according to item 12. (Item 14) The position and / or movement of the virtual content is controllable by the first handheld device of the first user, the apparatus according to item 1. (Item 15) The position and / or movement of the virtual content is also controllable by the second handheld device of the second user, the apparatus according to item 14. (Item 16) The processing unit is configured to localize the first user and the second user with respect to the same mapping information based on the one or more common anchor points, the apparatus according to item 1. (Item 17) The processing unit is configured to cause the virtual content to be displayed on the first display screen such that the virtual content will appear in a spatial relationship with respect to physical objects within the surrounding environment of the first user, the apparatus according to item 1. (Item 18) The processing unit is configured to obtain one or more sensor inputs, The processing unit is configured to assist the first user in performing an objective involving the virtual content based on the one or more sensor inputs, the apparatus according to item 1. (Item 19) The one or more sensor inputs indicate the gaze direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing of the first user, the apparatus according to item 18. (Item 20) The processing unit is configured to assist the first user in performing the objective by applying one or more limits regarding the position and / or angular velocity of system components, the apparatus according to item 18. (Item 21) The apparatus according to item 18, wherein the processing unit is configured to assist the first user in accomplishing the objective by gradually reducing the distance between the virtual content and another element. (Item 22) The apparatus according to item 1, wherein the processing unit includes a first processing part that communicates with the first display screen and a second processing part that communicates with the second display screen. (Item 23) A method implemented by an apparatus configured to provide virtual content in an environment, wherein in the environment, a first user wearing a first display screen and a second user wearing a second display screen can interact with each other, and the method includes: obtaining a first position of the first user; determining a first set of one or more anchor points based on the first position of the first user; obtaining a second position of the second user; determining a second set of one or more anchor points based on the second position of the second user; determining one or more common anchor points that are in both the first set and the second set; providing the virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points The method includes. (Item 24) The method according to item 23, wherein the one or more common anchor points include a plurality of common anchor points, and the method further includes selecting a subset of the common anchor points from the plurality of common anchor points. (Item 25) The method according to item 24, wherein the subset of the common anchor points is selected to reduce the relative positioning error between the first user and the second user. (Item 26) The method according to item 23, wherein the one or more common anchor points include a single common anchor point. (Item 27) The method according to item 23, further including determining a position and / or orientation of the virtual content based on at least one of the one or more common anchor points. (Item 28) The method according to item 23, wherein each of the one or more anchor points in the first set is a point in a Persistent Coordinate Frame (PCF). (Item 29) The method according to item 23, wherein the virtual content is provided as a movable virtual object within the first display screen and / or the second display screen for display. (Item 30) The method according to item 29, wherein the virtual object is provided for display within the first display screen such that the virtual object appears to be moving within the space between the first user and the second user. (Item 31) The one or more common anchor points comprise a first common anchor point and a second common anchor point, The movable virtual object is provided for display within the first display screen such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, The first object position of the movable virtual object is based on the first common anchor point, The second object position of the movable virtual object is based on the second common anchor point, The method according to item 29. (Item 32) The method according to item 31, further comprising selecting the first common anchor point for placing the virtual object at the first object position based on where the action of the virtual object occurs. (Item 33) The one or more common anchor points comprise a single common anchor point, The movable virtual object is provided for display within the first display screen such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen, The first object position of the movable virtual object is based on the single common anchor point, The second object position of the movable virtual object is based on the single common anchor point, The method according to item 29. (Item 34) The one or more common anchor points comprise a plurality of common anchor points, and the method further includes selecting one of the common anchor points for placing the virtual content within the first display screen, the method according to item 23. (Item 35) The act of selecting includes selecting one of the common anchor points that is closest to the action of the virtual content or within a distance threshold from the action of the virtual content, the method according to item 34. (Item 36) The position and / or movement of the virtual content is controllable by the first handheld device of the first user, the method according to item 23. (Item 37) The position and / or movement of the virtual content is also controllable by the second handheld device of the second user, the method according to item 36. (Item 38) The method according to item 23 further includes localizing the first user and the second user with respect to the same mapping information based on the one or more common anchor points. (Item 39) The method according to item 23 further includes displaying the virtual content by the first display screen such that the virtual content appears in a spatial relationship with a physical object in the surrounding environment of the first user. (Item 40) acquiring one or more sensor inputs; assisting the first user in performing an objective involving the virtual content based on the one or more sensor inputs The method according to item 23 further includes. (Item 41) The one or more sensor inputs indicate the line-of-sight direction of the first user, upper limb kinematics, body position, body orientation, or any combination of the foregoing, the method according to item 40. (Item 42) The act of assisting the first user in performing the objective includes applying one or more limits regarding the position and / or angular velocity of a system component, the method according to item 40. (Item 43) The act of assisting the first user in performing the objective includes gradually reducing the distance between the virtual content and another element, the method according to item 40. (Item 44) The apparatus according to item 23, comprising a first processing part that communicates with the first display screen and a second processing part that communicates with the second display screen. (Item 45) A processor-readable non-transitory medium storing a set of instructions, the execution of the set of instructions by a processing unit causes the method to be implemented, the processing unit is part of an apparatus configured to provide virtual content in an environment, in which a first user and a second user can interact with each other, the method comprising: obtaining a first position of the first user; determining a first set of one or more anchor points based on the first position of the first user; obtaining a second position of the second user; determining a second set of one or more anchor points based on the second position of the second user; determining one or more common anchor points that are in both the first set and the second set; providing the virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points A processor-readable non-transitory medium comprising the above.

Brief Description of the Drawings

[0059] The drawings illustrate the design and utility of various embodiments of the present disclosure. Note that the figures are not drawn to exact scale, and elements of similar structure or function are represented by like reference numerals throughout the figures. To gain a deeper understanding of the foregoing and other advantages and objects of various embodiments of the present disclosure, a more detailed description of the present disclosure briefly described above will be given by referring to the specific embodiments illustrated in the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the present disclosure and are not to be considered as limiting of its scope, and that the present disclosure will be described and described with additional specificity and detail through the use of the accompanying drawings.

[0060]

Figure 1A

[0061]

Figure 1B

[0062]

Figure 2

[0063]

Figure 3

[0064]

Figure 4

[0065]

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 5F

Figure 5G

Figure 5H

Figure 5I

Figure 5J

Figure 5K

Figure 5L

[0066]

Figure 6

[0067]

Figure 7A

Figure 7B

Figure 7C

Figure 7D

[0068]

Figure 8

[0069]

Figure 9

[0070]

Figure 10

[0071]

Figure 11

[0072] Detailed Description Various embodiments of the present disclosure are directed to methods, apparatuses, and articles of manufacture for providing input for a head-mounted video image device. Other objects, features, and advantages of the present disclosure are set forth in the detailed description, the drawings, and the claims.

[0073] Various embodiments will be described hereinafter with reference to the figures. Note that the figures are not drawn to exact scale, and elements of similar structure or function are represented by like reference numerals throughout the figures. Also note that the figures are intended only to facilitate the description of the embodiments. They are not intended as an exhaustive description of the invention or as a limitation on the scope of the invention. Additionally, the illustrated embodiments need not have all the aspects or advantages shown. Aspects or advantages described in conjunction with a particular embodiment are not necessarily limited to that embodiment and may be practiced in any other embodiment, whether or not such aspects or advantages are so illustrated or explicitly described.

[0074] The following description relates to exemplary VR, AR, and / or MR systems by which the embodiments described herein can be practiced. However, the embodiments are also suitable for use in other types of display systems (including other types of VR, AR, and / or MR systems), and thus it should be understood that the embodiments are not limited only to the exemplary examples disclosed herein.

[0075] Referring to FIG. 1A, an augmented reality system 1 is illustrated, featuring a head-mounted visual component (image display device) 2, a handheld controller component 4, and an interconnected auxiliary computing or controller component 6 that can be configured to be worn on a user as a belt pack or equivalent. Each of these components may be operably coupled to each other and to other connected resources 8 such as cloud computing or cloud storage resources via wired or wireless communication configurations defined by IEEE802.11, Bluetooth® (RTM), and other connectivity standards and configurations (10, 12, 14, 16, 17, 18). Using various embodiments such as two depicted optical elements 20, the user can see the surrounding world along with visual components that can be produced by system components associated for an augmented reality experience. As shown in FIG. 1A, such a system 1 may also include various sensors configured to provide information about the user's surrounding environment, including but not limited to various camera type sensors (monochrome, color / RGB, and / or thermal imaging components, etc.) (22, 24, 26), a depth camera sensor 28, and / or a sound sensor such as a microphone 30. There is a need for compact and continuously connected wearable computing systems and assemblies such as those described herein that can be utilized to provide the user with a rich perception of the augmented reality experience.

[0076] System 1 also includes an apparatus 7 for providing input for the image display device device 2. Apparatus 7 will be described in further detail below. The image display device 2 can be any of a VR device, an AR device, an MR device, or other types of display devices. As shown in the figure, the image display device 2 includes a frame structure worn by an end user, a display subsystem carried by the frame structure so that the display subsystem is positioned in front of the end user's eyes, and a speaker carried by the frame structure so that the speaker is positioned adjacent to the end user's external auditory canal (optionally, another speaker (not shown) is positioned adjacent to the other external auditory canal of the end user to provide stereo / adjustable sound control). The display subsystem is designed to present a light pattern to the end user's eyes that can be comfortably perceived as an extension to the physical reality with a high level of image quality and three-dimensional perception, and can also present two-dimensional content. The display subsystem presents a sequence of frames at a high frequency to provide the perception of a single coherent scene.

[0077] In the illustrated embodiment, the display subsystem employs an "optical see-through" display through which the user can directly view light from real objects through a transparent (or translucent) element. The transparent element is often referred to as a "combiner" and superimposes light from the display across the user's view of the real world. To achieve this purpose, the display subsystem comprises a partially transparent display or a fully transparent display. The display is positioned within the end user's field of view between the end user's eyes and the surrounding environment such that direct light from the surrounding environment is transmitted through the display to the end user's eyes.

[0078] In the illustrated embodiment, the image projection assembly provides light to a partially transparent display, whereby it is combined with direct light from the ambient environment and transmitted from the display to the user's eyes. The projection subsystem may be a fiber optic scanning-based projection device, and the display may be a waveguide-based display into which scanned light from the projection subsystem is input to produce an image at a single optical viewing distance closer than, for example, infinity (e.g., arm's length), an image at a plurality of discrete optical viewing distances or focal planes, and / or an image layer stack at a plurality of viewing distances or focal planes representing a three-dimensional 3D object. These layers in the light field may be stacked together sufficiently close so as to continuously appear in the human peripheral vision system (i.e., one layer is within the cone of confusion of an adjacent layer). Additionally or alternatively, the picture elements may be blended across two or more layers, and even when these layers are stacked more sparsely (i.e., one layer is outside the cone of confusion of an adjacent layer), the perceived continuity of the transition between the layers in the light field may be increased. The display subsystem may be for monocular or binocular use.

[0079] The image display device 2 may also include one or more sensors mounted on the frame structure to detect the position and movement of the end user's head and / or the end user's eye position and interpupillary distance. Such sensors may include image capture devices (such as cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes, or any combination of the foregoing. Many of these sensors operate based on the assumption that the frame to which they are attached is substantially fixed to the user's head, eyes, and ears, in that order.

[0080] The image display device 2 may also include a user orientation detection module. The user orientation module may detect the instantaneous position of the end user's head (e.g., via sensors coupled to the frame) and predict the position of the end user's head based on the position data received from the sensors. Detecting the instantaneous position of the end user's head facilitates the determination of the specific actual object that the end user is looking at, thereby providing an indication of the specific virtual object to be generated in relation to that actual object, and further providing an indication of the position at which the virtual object is to be displayed. The user orientation module may also track the end user's eyes based on the tracking data received from the sensors.

[0081] The image display device 2 may also include a control subsystem that may take any of a variety of forms. The control subsystem includes several controllers, such as one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers such as application specific integrated circuits (ASICs), programmable gate arrays (PGAs), such as field PGAs (FPGAs), and / or programmable logic controllers (PLUs).

[0082] The control subsystem of the image display device 2 may include a central processing unit (CPU), a graphics processing unit (GPU), one or more frame buffers, and a three-dimensional database for storing three-dimensional scene data. The CPU may control the overall operation, while the GPU may render frames from the three-dimensional data stored in the three-dimensional database (i.e., convert the three-dimensional scene into a two-dimensional image) and store these frames in the frame buffer. One or more additional integrated circuits may control the reading of frames into and out of the frame buffer and the operation of the image projection assembly of the display subsystem.

[0083] Device 7 represents various processing components for System 1. In the figure, Device 7 is illustrated as part of the image display device 2. In other embodiments, Device 7 may be implemented within the handheld controller component 4 and / or the controller component 6. In further embodiments, the various processing components of Device 7 may be implemented within a distributed subsystem. For example, the processing components of Device 7 may be located within two or more of the image display device 2, the handheld controller component 4, the controller component 6, or another device (communicating with the image display device 2, the handheld controller component 4, and / or the controller component 6).

[0084] The couplings 10, 12, 14, 16, 17, 18 between the various components described above may include one or more wired interfaces or ports for providing wire or optical communication, or one or more wireless interfaces or ports via RF, microwave, IR, etc. for providing wireless communication. In some implementations, all communication may be wired, while in other implementations, all communication may be wireless. Thus, the specific choice of wired or wireless communication should not be considered limiting.

[0085] Some image display systems (e.g., VR systems, AR systems, MR systems, etc.) use multiple volume phase holograms, surface relief holograms, or light guiding optical elements that incorporate depth plane information for generating images that appear to originate from individual depth planes. In other words, a diffraction pattern or diffraction optical element ("DOE") is incorporated within or imprinted / embossed onto a light guiding optical element ("LOE", e.g., a planar waveguide) such that collimated light (a light beam with a substantially planar wavefront) intersects the diffraction pattern at multiple locations as it is substantially totally internally reflected along the LOE and exits towards the user's eye. The DOE is configured such that the light exiting the LOE through it is converged so as to appear to originate from a particular depth plane. The collimated light may be generated using an optical condenser lens ("condenser").

[0086] For example, a first LOE may be configured to deliver collimated light that appears to originate from an optically infinite depth plane (0 diopters) to the eye. Another LOE may be configured to deliver collimated light that appears to originate from a distance of 2 meters (1 / 2 diopter). Yet another LOE may be configured to deliver collimated light that appears to originate from a distance of 1 meter (1 diopter). It should be understood that by using a stacked LOE assembly, multiple depth planes can be created, and each LOE is configured to display an image that appears to originate from a particular depth plane. It should be understood that the stack may include any number of LOEs. However, at least N stacked LOEs are required to generate N depth planes. Further, N, 2N, or 3N stacked LOEs may be used to generate an RGB color image on N depth planes.

[0087] To present 3D virtual content to a user, an image display system 1 (e.g., a VR system, an AR system, an MR system, etc.) projects an image of the virtual content into the user's eye such that they appear to originate from various depth planes in the Z direction (i.e., orthogonally away from the user's eye). In other words, the virtual content can vary not only in the X and Y directions (i.e., the 2D plane orthogonal to the user's central line of sight), but also appear to vary in the Z direction such that the user can perceive an object as being very close, or at infinite distance, or at any distance in between. In other embodiments, the user can perceive multiple objects simultaneously at different depth planes. For example, to the user, a virtual dragon can appear to come from infinity and run towards the user. Alternatively, to the user, a virtual bird at a distance of 3 meters from the user and a virtual coffee cup at arm's length (about 1 meter) from the user can be seen simultaneously.

[0088] A multi-plane focus system creates a perception of variable depth by projecting an image onto some or all of a plurality of depth planes located at individual fixed distances in the Z direction from the user's eye. Referring now to FIG. 1B, it should be understood that the multi-plane focus system can display a frame on a fixed depth plane 150 (e.g., the six depth planes 150 shown in FIG. 1B). The MR system can include any number of depth planes 150, but one exemplary multi-plane focus system has six fixed depth planes 150 in the Z direction. When generating virtual content on one or more of the six depth planes 150, a 3D perception is created such that the user perceives one or more virtual objects at variable distances from the user's eye. Assuming that the human eye is more sensitive to objects that are closer than objects that appear to be farther away, more depth planes 150 are generated closer to the eye, as shown in FIG. 1B. In some embodiments, the depth planes 150 may be equidistantly spaced apart from each other.

[0089] The depth plane position 150 may be measured in diopters, which is a unit of refractive power equal to the reciprocal of the focal length measured in meters. For example, in some embodiments, depth plane 1 may be 1 / 3 diopter apart, depth plane 2 may be 0.3 diopter apart, depth plane 3 may be 0.2 diopter apart, depth plane 4 may be 0.15 diopter apart, depth plane 5 may be 0.1 diopter apart, and depth plane 6 may represent infinity (i.e., 0 diopter apart). It should be understood that other embodiments may generate depth plane 150 at other distances / diopters. Thus, when generating virtual content at the strategically placed depth plane 150, the user is able to perceive virtual objects in three dimensions. For example, the user may perceive a first virtual object as being nearby when it is displayed within depth plane 1 while another virtual object appears at infinity in depth plane 6. Alternatively, the virtual object may be displayed such that it first appears in depth plane 6, then in depth plane 5, and so on until the virtual object appears very close to the user. It should be understood that the above examples are significantly simplified for illustrative purposes. In another embodiment, all six depth planes may be concentrated on a particular focal length away from the user. For example, if the virtual content to be displayed is a coffee cup that is 0.5 meters away from the user, all six depth planes may be generated at various cross-sections of the coffee cup, giving the user a high granularity 3D view of the coffee cup.

[0090] In some embodiments, the image display system 1 (e.g., VR system, AR system, MR system, etc.) may function as a multi-plane focus system. In other words, all six LOEs may be illuminated simultaneously such that images that appear to originate from six fixed depth planes are generated rapidly and continuously, and the light source rapidly transmits image information to LOE1, then LOE2, then LOE3, etc. For example, a portion of a desired image, including an empty image at optical infinity, may be input at time 1, and an LOE (e.g., depth plane 6 from FIG. 1B) that retains the collimation of light may be utilized. Then, an image of a closer tree branch may be input at time 2, and an LOE (e.g., depth plane 5 from FIG. 1B) configured to create an image that appears to originate from a depth plane 10 meters away may be utilized. Then, an image of a pen may be input at time 3, and an LOE configured to create an image that appears to originate from a depth plane 1 meter away may be utilized. This type of paradigm can be repeated in a high-speed time-sequential (e.g., 360 Hz) manner such that the user's eyes and brain (e.g., visual cortex) perceive the input as being all parts of the same image.

[0091] The image display system 1 may project an image that appears to originate from various locations along the Z-axis (i.e., depth plane) to generate an image for a 3D experience / scenario (i.e., by diverging or converging a light beam). As used herein, a light beam includes a directed projection of light energy (including visible and invisible light energy) emitted from a light source, among other things. Generating an image that appears to originate from various depth planes corresponds to the convergence / divergence movement and accommodation of the user's eyes for that image, minimizing or eliminating convergence / divergence movement-accommodation conflict.

[0092] In some cases, an environmental location map is obtained to locate the user of the head-mounted image display device relative to the user's environment. In some embodiments, the location map may be stored in a non-transitory medium that is part of system 1. In other embodiments, the location map may be received wirelessly from a database. After the location map is obtained, a real-time input image from the camera system of the image display device is then matched against the location map to locate the user. For example, the corner features of the input image may be detected from the input image and matched against the corner features of the location map. In some embodiments, the image may first need to undergo corner detection to obtain an initial set of detected corners in order to obtain features from the image for use in location using a set of corners. The initial set of detected corners may then be further processed, for example, through non-maximum suppression, spatial binning, etc. to obtain a final set of detected corners for location purposes. In some cases, filtering may be performed to identify a subset of the detected corners within the initial set and obtain a final set of corners.

[0093] Also, in some embodiments, the environmental location map may be created by the user pointing the image display device 2 in different directions (e.g., by turning their head while wearing the image display device 2). As the image display device 2 is pointed at different spaces within the environment, sensors on the image display device 2 sense the characteristics of the environment, which can then be used by system 1 to create the location map. In one implementation, the sensors may include one or more cameras and / or one or more depth sensors. The camera provides a camera image, which is processed by device 7 to identify different objects within the environment. Additionally, or alternatively, the depth sensor provides depth information, which is processed by the device to determine different surfaces of the objects within the environment.

[0094] In various embodiments, a user may be wearing an augmented reality system such as that depicted in FIG. 1A, which may also be referred to as a “spatial computing” system in relation to the interaction with the three-dimensional world around the user when such a system is operated. Such a system may comprise, for example, a head-wearable display component 2 and may be characterized by environmental sensing capabilities such as various types of cameras, which may be configured to map the environment around the user or create a “mesh” of such an environment with various points representing the geometric shapes of various objects within the environment around the user, such as walls, floors, chairs, and the like. The spatial computing system may be configured to map or mesh the environment around the user and to launch or operate software such as that available from Magic Leap, Inc. (Planation, Florida), which may be configured to utilize the map or mesh of the room to assist the user in placing, manipulating, visualizing, creating, and modifying various objects and elements within the three-dimensional space around the user. Referring back to FIG. 1A, the system may be operably coupled to additional resources such as other computing systems by a cloud or other connectivity configuration. One of the challenges in spatial computing relates to the utilization of data captured by various operably coupled sensors (elements 22, 24, 26, 28, etc. of the system of FIG. 1A) in making various decisions useful and / or important to the user, such as in computer vision and / or object recognition challenges, which may be related to the three-dimensional world around the user.

[0095] For example, referring to FIG. 2, a typical spatial computing scenario is illustrated that utilizes a system (which may also be referred to as "ML1", representing the Magic Leap One (RTM) system available from Magic Leap, Inc. (Plantation, Florida)) such as that illustrated in FIG. 1A. A first user (who may be referred to as "User 1") boots up the ML1 system, mounts the head-mounted component 2 on their head, and ML1 scans the local environment around the user's head 1 using sensors that make up the head-mounted component 2 and performs a simultaneous localization and mapping ("SLAM") activity to create a local map or mesh (which may be referred to as "Local Map 1" in this scenario) of the environment around User 1's head, and User 1 may be "localized" within this Local Map 1 by the SLAM activity such that their real or near-real-time position and orientation are determined with respect to the local environment (40). Referring again to FIG. 2, User 1 may navigate around the environment, visually recognize and interact with real and virtual objects, and continue the SLAM activity to continue mapping / meshing the vicinity of the environment and generally enjoy the benefits of their spatial computing (42). Referring to FIG. 3, additional steps and configurations may be added such that User 1 may encounter one of a number of predefined anchor points or points within what may be known as a "persistent coordinate frame" or "PCF", and these anchor points and / or PCFs may be known to User 1's local ML1 system by virtue of previous installation and / or may be known via cloud connectivity (i.e., by element 8, which may be connected resources such as those illustrated in FIG. 1A, i.e., edge computing, cloud computing, and other connected resources) in a state where User 1 is localized within a cloud-based map (which may be larger and / or more refined than Local Map 1) (44).Referring again to FIG. 3, the anchor point and / or PCF may be utilized to assist the spatial computing task of user 1, such as by displaying various virtual objects or assets (e.g., virtual signs, etc. indicating that there is a sinkhole within a hiking course at a given fixed location near the user for a user during hiking) intentionally placed by others for user 1 (46).

[0096] Referring to FIG. 4, a multi - user (or "multi - player" in a game scenario) configuration similar to that described above with reference to FIG. 3 is illustrated. User 1 boots up ML1 and mounts it in a head - mounted configuration. ML1 scans the environment around the user's head 1, performs SLAM activity, creates a local map or mesh ( "Local Map 1") of the environment around the user 1's head, and User 1 is "localized" within Local Map 1 by the SLAM activity such that their real or near - real - time position and orientation are determined with respect to the local environment (40). User 1 may encounter one of a number of predetermined anchor points or points within what may be known as "persistent coordinate frames" or "PCFs". These anchor points and / or PCFs may be known to User 1's local ML1 system by previous installation and / or may be known via cloud connectivity in a state where User 1 is localized within a cloud - based map (larger than and / or more refined than Local Map 1) (48). A separate user, "User 2", may boot up a separate ML1 system and mount it in a head - mounted configuration. This second ML1 system scans the environment around User 2's head, performs SLAM activity, and may create a local map or mesh ( "Local Map 2") of the environment around User 2's head. User 2 is "localized" within Local Map 2 by the SLAM activity such that their real or near - real - time position and orientation are determined with respect to the local environment (50). Similar to User 1, User 2 may encounter one of a number of predetermined anchor points or points within what may be known as "persistent coordinate frames" or "PCFs". These anchor points and / or PCFs may be known to User 2's local ML1 system by previous installation and / or may be known via cloud connectivity in a state where User 2 is localized within a cloud - based map (larger than Local Map 1 or Local Map 2 and / or more refined) (52).Referring again to FIG. 4, user 1 and user 2 may be physically close enough that their ML1 systems begin to encounter a common anchor point and / or PCF. From the overlapping set between the two users, the system, which uses resources such as cloud computing connection resources 8, may be configured to select a subset of the anchor point and / or PCF, which minimizes the positioning error of the users relative to each other, and this subset of the anchor point and / or PCF may be utilized to position and orient virtual content for the users within the shared experience. Certain content and / or virtual assets may be experienced by both users from their respective unique lines of sight, along with the positioning and orientation positioning with respect to the handheld 4 and other components configured to form part of such a shared experience.

[0097] Referring to FIGS. 5A - 5L, an exemplary common, collaborative, or multi - user experience is shown in the form of a pancake - flipping game that may be available from Magic Leap, Inc. (Plantation, Florida) under the trademark name “Pancake Pals” (registered trademark). Referring to FIG. 5A, an element 60, which is a first user (also referred to as “User 1”), is shown holding a handheld component 4 that, at one end of an office environment 66, wears a head - mounted component 2 and interconnected auxiliary computing or controller components 6, forming an ML1 system similar to that illustrated in FIG. 1A. The system of User 1 (60) is configured to display, in virtual form, a virtual frying - pan element 62 that extends from that handheld component 4 as if the handheld component 4 were the handle of a frying - pan element 62 for him. The system of User 1 (60) is configured to re - position and re - orient the virtual frying - pan 62 as User 1 re - orients and re - positions that handheld component. Both the head - mounted 2 and handheld 4 components may be tracked from the perspective of their position and orientation relative to each other in real or near - real time, for example, using the tracking features of the system. The system is configured to enable interaction with the virtual element of the pancake 64 using simulated physics, for example, using the software physics capabilities of an environment such as Unity(RTM), such that other elements controlled by the user, such as the virtual frying - pan 62, or other virtual elements or other actual elements (such as one of the user's two hands, or perhaps an actual frying - pan or paddle, etc., configured to be trackable by the system), allow User 1 to flip the pancake 64, land it in its frying - pan 62, and / or throw or fling the virtual pancake 64 away from User 1 along a certain trajectory. The software physics capabilities may be configured to make the pancake bounce around the perimeter of any edge of the virtual frying - pan 62 and fly in orbits and patterns that an actual pancake could take.Referring to FIG. 5B, the virtual pancake 64 may have animated characteristics and be configured to provide the user with a perception of scoring, making sounds, music, and heart or rainbow images, and being pleased with the flipping and / or landing of the pancake 64 by equivalents, etc. FIG. 5C illustrates a user 1 (60) who has successfully landed the flipped virtual pancake 64 into its virtual frying pan 62. FIG. 5D illustrates the user 1 (60) preparing to throw the virtual pancake 64 in front of the user 1 (60). FIG. 5E illustrates the thrown virtual pancake 64 flying away from the user 1 (60).

[0098] Referring to FIG. 5F, in a multi-user experience such as that described above with reference to FIG. 4, the system is configured to locate two players within the same environment 66 relative to each other. Here, User 1 (60) and User 2 (61), who occupy two different ends of the same hallway in an office, are depicted (referring first to FIGS. 5K and 5L, a view showing both users can be seen), and as the virtual pancake 64 is thrown by User 1 (60) in FIG. 5E, the same virtual pancake 64, along with a simulated physics trajectory, is also thrown towards User 2 (61), who has its own virtual frying pan element 63 and is able to interact with the virtual pancake 64 using the ML1 system with the head-mounted 2, hand-held 4, and compute pack 6 components, using simulated physics. Referring to FIGS. 5G and 5H, User 2 (61) aligns its virtual frying pan 63 and successfully catches the virtual pancake 64, and referring to FIGS. 5I and 5J, User 2 (61) can throw the virtual pancake 64 back towards User 1 (60), and as shown in FIG. 5K, User 1 attempts to align its virtual frying pan 62 for another successful catch of the virtual pancake 64, or alternatively, in FIG. 5L, User 1 does not have enough time to position and orient its virtual frying pan 62 and the virtual pancake 64 appears to be headed towards the floor. Thus, a multi-user or multi-player configuration is presented in which two users can collaborate or interact with various elements such as virtual dynamic elements.

[0099] Regarding the PCF and the anchor point, as described above, a local map, such as one created by a local user, may contain a persistent anchor point or coordinate frame, which may correspond to the location and / or orientation of various elements. Maps and other external resources 8 that can be elevated to cloud-based computing resources, stored, or created may be merged with maps generated by other users. In fact, a given user may be located within the cloud or a portion thereof, which may be larger or more refined than what was generated in-situ by the user, as described above. Further, similar to the local map, the cloud map may be configured to contain a persistent anchor point or PCF, which may correspond to real-world locations and / or orientations, and which may be agreed upon by various devices within the same area or portion of the map or environment. When a user is located (e.g., first, after booting up, starting a scan using the ML1 system, or after loss of cooperation with the local map, or after walking a distance within the environment such that the SLAM activity aids in locating the user), the user can be located based on neighboring map features that correspond to observable features within the real world. The persistent anchor point and the PCF may correspond to real-world locations, but they may also be fixed relative to each other until the map itself is updated. For example, if PCF-A and PCF-B are 5 meters apart, they can be configured to remain 5 meters apart even if the user is re-located (i.e., the system can be configured such that the individual PCFs do not move, i.e., only the user's estimated map alignment and the user's location within it move). Referring previously to FIGS. 6 and 7A-7C, the further the user is from a high-confidence PCF (i.e., near the user's location point), the greater the error will be in terms of location and orientation.For example, the second-degree error in PCF alignment on any axis can be associated with an offset of 35 centimeters at 10 meters (tan(2 degrees)×10), i.e., as discussed previously with reference to FIGS. 6 and 7A-7C, it is preferable to utilize the persistent anchor points and PCF close to the user or object being tracked.

[0100] As described with reference to FIG. 4, neighboring users may be configured to receive the same neighboring PCF and persistent anchor point information. From the overlapping set of PCFs and persistent anchor points between two players, the system may be configured to select one or more PCFs and / or persistent anchor points, which minimizes the average error, and these shared anchors / PCFs may be used to set the position / orientation of the shared virtual content. In one embodiment, a host system such as a cloud computing resource may be configured to provide and transmit mesh / mapping information to all users when starting a new collaborative session or game. The mesh may be positioned by a local user within a suitable installation location using its common anchor with the host, and the system may be configured to specifically "cut" or "crop" off a vertically tall mesh portion so that the player may have a perception of a very large "headroom" or ceiling height (useful for scenarios such as flipping a pancake between two users when a maximum "communication time" is desired). Various aspects of the mapping or mesh-based limits such as the ceiling height, etc., may optionally be avoided or ignored (e.g., in one embodiment, only the floor mesh / map may be utilized to confirm that a particular virtual element such as a flying pancake has impacted the ground surface).

[0101] As described above, soft physics simulations may be utilized. In various embodiments, such collisional elements are configured to have extensions that grow on opposite sides of the pancake so that when particles land between the virtual pans 62, 63 or other collisional elements, they are prevented from becoming solid so that the collisions can be properly resolved.

[0102] The user's system and associated connected resources 8 may be configured to enable the user to start a game through associated social network resources (such as a predetermined group of friends, etc. that also have the ML1 system) and by such user's geographical location. For example, when a particular user attempts to play a game as illustrated in FIGS. 5A-5L, the associated system may be configured to automatically discourage the selection of play partners from among the people within the particular user's social network in the same building or the same room.

[0103] Referring back to FIGS. 5A-5L, the system may be configured to deal with only one virtual pancake 64 at a time so that computational efficiency can be obtained (i.e., the entire game state may be packed within a single packet. In various embodiments, the player / user closest to a given virtual pancake 64 at any given time is given authority over such virtual pancake 64, which may mean that they control where it is for other users). In various embodiments, the networked state updates the physics and renders everything at an update frequency within the range of 60 Hz each, which is selected to maintain visual presentation consistently even when crossing authority boundaries (i.e., the authority of a pancake from one user to another). In various embodiments, a Tnet configuration (packet switching two-point system area network) may be utilized along with some custom packet configurations and other more standard packet configurations such as those described in known RFC publications.

[0104] Referring to FIG. 6, two users (User 1 (60), User 2 (61)) are positioned very close to each other within the same local environment 66 with one PCF 68 located immediately beside them. With such a configuration and using only one reliable PCF 68, games such as those illustrated in FIGS. 5A - 5L can be played with both users being located relative to the same mapping information by the same PCF. However, referring to FIGS. 7A - 7C, the users (User 1 (60), User 2 (61)) are positioned relatively far apart, but fortunately, a plurality of PCFs (69, 70, 71, 72) are positioned between the two users. As described above, in order to prevent offset and / or drift, it is desirable to use a PCF that is as close to the action as possible. Thus, in FIG. 7A, when User 2 (61) is on the side of flipping / throwing and generally has proximity and its control with the virtual pancake 64, neighboring PCFs (e.g., 70, 69, or both 69 and 70, etc.) may be utilized. Referring to FIG. 7B, as the virtual pancake 64 is thrown between the two users (61, 60), the PCF (e.g., 70, 71, or both 70 and 71) closest to the flying virtual pancake may be utilized. As shown in FIG. 7C, when the virtual pancake is released at a lower position, other PCFs (e.g., 69, 70, or both 69 and 70) closest to the flying virtual pancake may be utilized. Referring to FIG. 7D, as the virtual pancake (64) continues its trajectory towards User 1 (60), the PCF (e.g., 69, 70, or both 69 and 70) closest to the flying virtual pancake may be utilized. Thus, a dynamic PCF selection configuration can assist in minimizing drift and maximizing accuracy.

[0105] Referring to FIG. 8, users 1 and 2 may be located within the same map (80) such that certain content and / or virtual assets can be experienced by both users from their respective lines of sight so that they can engage in a common space computing experience. A connected system such as local computing capabilities or certain cloud computing resources 8 that can be interconnected and resident within the ML1 system of the user may be configured to attempt to infer the intended aspects from the user's activities such as line of sight, upper limb kinematics, and body position and orientation with respect to the local environment (82). For example, in one embodiment, the system may be configured to utilize the captured line of sight information of the first user to infer a destination for releasing virtual elements such as a pancake 64, such as an approximate location of the second user that the first user may attempt to target. Similarly, the system may be configured to infer targeting or other variables from upper limb position, orientation, velocity, and / or angular velocity. The information utilized to assist the user may be from real or near real-time sampling or may be based on sampling from a larger time domain such as the play or participation history of a particular user (e.g., a convolutional neural network (“CNN”) configuration may be utilized to understand that a particular user always glances in a particular way or always moves their arm in a particular way when attempting to hit a particular target or in a particular manner, and such a configuration may be utilized to assist the user). At least in part, based on one or more interferences regarding the user's intent, the system may be configured to assist the user in accomplishing the intended purpose (84).For example, in one embodiment, the system provides functional limits on the position or angular velocity of a given system component relative to the local environment when tremors are detected by the user's hand or arm (i.e., to reduce effects that deviate from the norm of the tremors, thereby smoothing the user's commands to the system), or when it is determined that the collision is desired, causes one or more elements to be slightly drawn towards each other over time (i.e., in an example of a young child with relatively undeveloped gross motor skills, the system may be configured to assist the child in aiming or moving in a game or use case so that the child is more successful. For example, with respect to a child who continuously fails to catch the virtual pancake 64, by placing the virtual frying pan 62 extremely to the right, the system may be configured to guide the pancake towards the frying pan, or the frying pan towards the pancake, or both).

[0106] In various other embodiments, the system may be configured to have other interactions with the local physical world. For example, in one embodiment, a user playing a game as illustrated in FIGS. 5A - 5L can intentionally slap a virtual pancake 64 towards a wall, and if the trajectory is straight enough and at sufficient speed, the virtual pancake may remain stuck to the wall for the remainder of the game, or even permanently, as part of the local mapping information, or associated with a neighboring PCF.

[0107] Processing unit

[0108] FIG. 9 illustrates a processing unit 1002 according to some embodiments. The processing unit 1002 may be, in some embodiments, an example of the apparatus 7 described herein. In other embodiments, the processing unit 1002 or any part of the processing unit 1002 may be implemented using separate devices that communicate with each other. As illustrated, the processing unit 1002 includes a communication interface 1010, a positioner 1020, a graphic generator 1030, a non-transitory medium 1040, a controller input 1050, and a task assistant 1060. In some embodiments, the communication interface 1010, the positioner 1020, the graphic generator 1030, the non-transitory medium 1040, the controller input 1050, the task assistant 1060, or any combination of the foregoing may be implemented using hardware. By way of non-limiting example, the hardware may include one or more FPGA processors, one or more ASIC processors, one or more signal processors, one or more math processors, one or more integrated circuits, or any combination of the foregoing. In some embodiments, any component of the processing unit 1102 may be implemented using software.

[0109] In some embodiments, processing unit 1002 may be implemented as separate components that are communicatively coupled together. For example, processing unit 1002 may have a first substrate that carries communication interface 1010, positioner 1020, graphic generator 1030, controller input 1050, task assistant 1060, and another substrate that carries non-transitory medium 1040. As another example, all components of processing unit 1002 may be carried by the same substrate. In some embodiments, any, some, or all of the components of processing unit 1002 may be implemented in image display device 2. In some embodiments, any, some, or all of the components of processing unit 1002 may be implemented in a device remote from image display device 2, such as handheld control component 4, control component 6, mobile phone, server, etc. In further embodiments, processing unit 1002 or any of the components of processing unit 1002 (such as positioner 1020, etc.) may be implemented in different display devices worn by different individual users, or may be implemented in different devices associated with (e.g., in proximity to) different individual users.

[0110] The processing unit 1002 is configured to receive position information (e.g., from a sensor in the image display device 2 or from an external device) and / or control information from the controller component 4, and based on the position information and / or control information, provide virtual content within the screen of the image display device 2 for display. For example, as illustrated with reference to FIGS. 5A - 5L, the position information may indicate the position of the user 60, and the control information from the controller 4 may indicate the position of the controller 4 and / or an action being performed by the user 60 via the controller 4. In such a case, the processing unit 1002 generates an image of a virtual object (e.g., the pancake 64 in the above embodiment) based on the position of the user 60 and the control information from the controller 4. In one embodiment, the control information indicates the position of the controller 4. In such a case, the processing unit 1002 generates an image of the pancake 64 such that the position of the pancake 64 is related to the position of the controller 4 (as shown in FIG. 5D). For example, the movement of the pancake 64 will follow the movement of the controller 4. In another embodiment, when the user 60 uses the controller 4 to perform maneuvers such as throwing the virtual pancake 64 using the controller 4, the control information will include information regarding the direction of movement of the controller 4 and the speed and / or acceleration associated with the movement. In such a case, the processing unit 1002 then generates a graphic indicating the movement of the virtual pancake 64 (as shown in FIGS. 5E - 5G). The movement may be along a movement trajectory calculated by the processing unit 1002 based on the position where the pancake 64 exits the virtual frying pan 63 and based on a movement model (which receives as input the direction of movement of the controller 4 and the speed and / or acceleration of the controller 4). The movement model will be described in more detail herein.

[0111] Returning to FIG. 9, the communication interface 1010 is configured to receive location information. As used herein, the term "location information" refers to any information that represents the location of an entity, or any information that can be used to derive the location of an entity. In some embodiments, the communication interface 1010 is communicatively coupled to the camera and / or depth sensor of the image display device 2. In such embodiments, the communication interface 1010 directly receives images from the camera and / or depth signals from the depth sensor. In some embodiments, the communication interface 1010 may be coupled to another device, such as another processing unit, that processes the images from the camera and / or the depth signals from the depth sensor before the location information passes through the communication interface 1010. In other embodiments, the communication interface 1010 may be configured to receive GPS information, or any information that can be used to derive a location. Also, in some embodiments, the communication interface 1010 may be configured to obtain location information that is output wirelessly or via a physical conductive transmission line.

[0112] In some embodiments, where different sensors are present in the image display device 2 to provide different types of sensor outputs, the communication interface 1010 of the processing unit 1002 may have different individual sub - communication interfaces to receive the different individual sensor outputs. In some embodiments, the sensor output may include an image captured by a camera in the image display device 2. Alternatively, or in addition, the sensor output may include distance data captured by a depth sensor in the image display device 2. The distance data may be data generated based on time - of - flight techniques. In such a case, a signal generator in the image display device 2 transmits a signal, and the signal reflects from objects in the environment around the user. The reflected signal is received by a receiver in the image display device 2. Based on the time it takes for the signal to reach the object and reflect back to the receiver, the sensor or the processing unit 1002 may then determine the distance between the object and the receiver. In other embodiments, the sensor output may include any other data that can be processed to determine the location of entities (users, objects, etc.) in the environment.

[0113] The positioner 1020 of the processing unit 1002 is configured to determine the position of a user of the image display device and / or to determine the position of a virtual object to be displayed within the image display device. In some embodiments, the position information received by the communication interface 1010 may be a sensor signal, and the positioner 1020 is configured to process the sensor signal to determine the position of the user of the image display device. For example, the sensor signal may be a camera image captured by one or more cameras of the image display device. In such a case, the positioner 1020 of the processing unit 1002 is configured to determine a positioning map based on the camera image and / or to match features in the camera image with features in the created positioning map for user positioning. In one implementation, the positioner 1020 is configured to perform the actions described with reference to FIGS. 2 and / or 3 for user positioning. In other embodiments, the position information received by the communication interface 1010 may already indicate the position of the user. In such a case, the positioner 1020 then uses the position information as the position of the user.

[0114] As shown in FIG. 9, the positioner 1020 includes an anchor point module 1022 and an anchor point selector 1024. The anchor point module 1022 is configured to determine one or more anchor points, which may be utilized by the processing unit 1002 to locate the user and / or place virtual objects relative to the environment surrounding the user. In some embodiments, the anchor points may be points within a positioning map, and each point within the positioning map may be a feature (e.g., a corner, an edge, an object, etc.) identified within the physical environment. Also, in some embodiments, each anchor point may be a persistent coordinate frame (PCF) determined previously or in the current session. In some embodiments, the communication interface 1010 may receive previously determined anchor points from another device. In such a case, the anchor point module 1022 may obtain the anchor points by receiving them from the communication interface 1010. In other embodiments, the anchor points may be stored within the non-transitory medium 1040. In such a case, the anchor point module 1022 may obtain the anchor points by reading them from the non-transitory medium 1040. In further embodiments, the anchor point module 1022 may be configured to determine anchor points within a mapping session. In the mapping session, a user wearing an image display device walks around within the environment and / or orientates the image display device at different viewing angles such that the camera of the image display device captures images of different features within the environment. The processing unit 1002 may then perform feature identification and identify one or more features within the environment for use as anchor points. In some embodiments, the anchor points for a certain physical environment have already been determined within a previous session. In such a case, when the user enters the same physical environment, the camera in the image display device worn by the user will capture an image of the physical environment.The processing unit 1002 may identify features in the physical environment and check whether one or more of the features match previously determined anchor points. If so, the matched anchor points will be made available by the anchor point module 1022 for the processing unit 1002 to use them for user location identification and / or virtual content placement.

[0115] Also, in some embodiments, as the user moves around in the physical environment, the anchor point module 1022 of the processing unit 1002 will identify additional anchor points. For example, when the user is at a first position in the environment, the anchor point module 1022 of the processing unit 1002 may identify anchor points AP1, AP2, AP3 that are in close proximity to the user's first position in the environment. When the user moves from the first position to a second position in the physical environment, the anchor point module 1022 of the processing unit 1002 may identify anchor points AP3, AP4, AP5 that are in close proximity to the user's second position in the environment.

[0116] In addition, in some embodiments, the anchor point module 1022 is configured to obtain anchor points associated with multiple users. For example, two users in the same physical environment may be standing apart from each other. The first user may be at a first location with a first set of anchor points associated with it. Similarly, the second user may be at a second location with a second set of anchor points associated with it. Since the two users are standing apart from each other, initially, the first set of anchor points and the second set of anchor points may not have any overlap. However, when one or both of the users move towards each other, the composition of the anchor points within the individual first and second sets will change. When they get close enough, the first and second sets of anchor points will begin to have an overlap.

[0117] The anchor point selector 1024 is configured by the processing unit 1002 to select a subset of anchor points (provided by the anchor point module 1022) for use in locating a user and / or placing virtual objects in an environment surrounding the user. In some embodiments, the anchor point module 1022 provides a plurality of anchor points associated with a single user, and when no other users are involved, the anchor point selector 1024 may select one or more of the anchor points for user location and / or placement of virtual content in the physical environment. In other embodiments, the anchor point module 1022 may provide a set of a plurality of anchor points associated with different individual users (e.g., users wearing individual image display devices) who desire to interact virtually with each other within the same physical environment. In such a case, the anchor point selector 1024 is configured to select one or more common anchor points that are common among different sets of anchor points. For example, as shown in FIG. 6, one common anchor point 68 may be selected to enable users 60, 61 to interact with the same virtual content (pancake 64). FIG. 7A shows another example in which four common anchor points 69, 70, 71, 72 are selected to enable users 60, 61 to interact with virtual content (pancake 64). The processing unit 1002 may then utilize the selected common anchor points for placement of virtual content so that users can interact with the virtual content within the same physical environment.

[0118] In some embodiments, the anchor point selector 1024 may be configured to perform the actions described with reference to FIG. 4.

[0119] Returning to FIG. 9, the controller input 1050 of the processing unit 1002 is configured to receive an input from the controller component 4. The input from the controller component 4 may be position information regarding the position and / or orientation of the controller component 4, and / or control information based on a user action performed via the controller component 4. As a non-limiting example, the control information from the controller component 4 may be generated based on the user translating the controller component 4, rotating the controller component 4, pressing one or more buttons on the controller component 4, actuating a knob, trackball, or joystick on the controller component 4, or any combination of the foregoing. In some embodiments, the user input is utilized by the processing unit 1002 to insert and / or move a virtual object presented on the screen of the image display device 2. For example, if the virtual object is the virtual pancake 64 described with reference to FIGS. 5A-5L, the handheld controller component 4 may be operated by the user to catch the virtual pancake 64, move the virtual pancake 64 using the frying pan 62, and / or throw the virtual pancake 64 away from the frying pan 62 such that the virtual pancake 64 appears to be moving within the physical environment as it is visually perceived by the user through the screen of the image display device 2. In some embodiments, the handheld controller component 4 may be configured to move the virtual object within the two-dimensional display screen such that the virtual object appears to be moving within a virtual three-dimensional space. For example, in addition to moving the virtual object up, down, left, and right, the handheld controller component 4 may also move the virtual object in and out of the user's visual depth.

[0120] The graphic generator 1030 is configured to generate graphics on the screen of the image display device 2, at least in part, based on the output from the positioner 1020 and / or the output from the controller input 1050 for display. For example, the graphic generator 1030 may control the screen of the image display device 2 and display virtual objects so that they appear to be within the environment when viewed by the user through the screen. As non-limiting examples, the virtual objects may be virtual movable objects (e.g., balls, shuttles, bullets, missiles, fires, heat waves, energy waves), weapons (e.g., swords, axes, hammers, knives, bullets, etc.), any objects that can be found inside a room (e.g., pencils, rolled-up paper balls, cups, chairs, etc.), any objects that can be found outside a building (e.g., rocks, tree branches, etc.), vehicles (e.g., cars, airplanes, space shuttles, rockets, submarines, helicopters, motorcycles, bicycles, tractors, all-terrain vehicles, snowmobiles, etc.), and the like. Also, in some embodiments, the graphic generator 1030 may generate an image of a virtual object on the screen for display such that the virtual object will appear to interact with actual physical objects within the environment. For example, the graphic generator 1030 may display an image of a virtual object on the screen in a moving configuration such that it appears to be moving through the space within the environment when viewed by the user through the screen of the image display device 2. Also, in some embodiments, the graphic generator 1030 may display an image of a virtual object on the screen such that it appears to deform or damage an actual physical object within the environment or another virtual object when viewed by the user through the screen of the image display device 2. In some cases, this may be accomplished by the graphic generator 1030 generating interaction images such as images of deformation marks (e.g., dents, creases, etc.), burn marks, images indicating thermal changes, images of fires, explosion images, debris images, etc. for display on the screen of the image display device 2.

[0121] As described above, in some embodiments, the graphic generator 1030 may be configured to provide virtual content as a movable virtual object such that the virtual object appears to move within the three-dimensional space of the physical environment surrounding the user. For example, the movable virtual object may be the flying pancake 64 described with reference to FIGS. 5A-5L. In some embodiments, the graphic generator 1030 may be configured to generate the graphics of the flying pancake 64 based on an orbit model and based on one or more anchor points provided by the anchor point module 1022 or the anchor point selector 1024. For example, the processing unit 1002 may determine an initial orbit for the flying pancake 64 based on the orbit model, and the initial orbit indicates the location where the action of the flying pancake 64 is desired to be performed. The graphic generator 1030 then generates a sequence of images of the pancake 64 and forms a video of the flying pancake 64 corresponding to (e.g., following as closely as possible) the initial orbit. The position of each image of the pancake 64 to be presented in the video may be determined by the processing unit 1002 based on the proximity of the action of the pancake 64 to one or more neighboring anchor points. For example, as discussed with reference to FIG. 7A, when the action of the pancake 64 approaches close to the anchor points 69, 70, one or both of the anchor points 69, 70 may be used by the graphic generator 1030 to place the pancake 64 at a desired position relative to the display screen such that when the users 60, 61 view the pancake 64 in relation to the physical environment, the pancake 64 will be in the correct position relative to the physical environment. As shown in FIG. 7B, as the pancake 64 further progresses along its orbit, the action of the pancake 64 approaches close to the anchor points 70, 71.In such a case, one or both of the anchor points 70, 71 may be utilized by the graphic generator 1030 to place the pancake 64 at a desired position relative to the display screen such that when the user 60, 61 views the pancake 64 in relation to the physical environment, the pancake 64 will be in the correct position relative to the physical environment. As shown in FIG. 7D, as the pancake 64 further progresses along its trajectory, the action of the pancake 64 approaches the anchor points 69, 72. In such a case, one or both of the anchor points 69, 72 may be utilized by the graphic generator 1030 to place the pancake 64 at a desired position relative to the display screen such that when the user 60, 61 views the pancake 64 in relation to the physical environment, the pancake 64 will be in the correct position relative to the physical environment.

[0122] As illustrated in the above embodiments, the actual positioning of the pancake 64 at various locations along its trajectory is based on different anchor points (e.g., features identified within a physical environment, PCF, etc.). Therefore, as the pancake 64 moves across space, the pancake 64 is accurately positioned relative to the anchor point that the moving pancake 64 (where the action of the pancake 64 takes place) approaches. This feature is advantageous as it prevents the pancake 64 from being inaccurately positioned relative to the environment, which could otherwise occur if the pancake 64 were positioned relative to only one anchor point close to the user. For example, if the positioning of the pancake 64 is based only on the anchor point 70, as the pancake 64 moves further away from the user 61, the distance between the pancake 64 and the anchor point 70 increases. If there are some errors in the anchor point 70, such as incorrect positioning and / or orientation of the PCF, this results in the pancake 64 being offset or drifting away from its intended position, and the magnitude of the offset or drift becomes larger as the pancake 64 moves further away from the anchor point 70. The above technique of selecting different anchor points that approach the pancake 64 to install the pancake 64 addresses the offset and drifting problems. The above feature is also advantageous as it enables multiple users who are separated (e.g., more than 5 feet, more than 10 feet, more than 15 feet, more than 20 feet, etc.) to accurately interact with each other and / or interact with the same virtual content. In a gaming application, the above technique may enable multiple users to accurately interact with the same object even when the users are relatively far apart. For example, in a gaming application, a virtual object may be virtually passed back and forth between remote users.As used herein, the term "proximity" refers to the distance between two items that meets a criterion such as distance that is less than a pre-defined value (e.g., less than 15 feet, 12 feet, 10 feet, 8 feet, 6 feet, 4 feet, 2 feet, 1 foot, etc.).

[0123] Note that the above-described technique of placing virtual content based on an anchor point in proximity to an action of the virtual content is not limited to games involving two users. In other embodiments, the above-described technique of placing virtual content may be applied to any application, with only a single user or more than two users (which may or may not be a game application). For example, in other embodiments, the above-described technique of placing virtual content may be utilized in an application that enables a user to place virtual content that is remote from the user within a physical environment. The above-described technique of placing virtual content is advantageous because it enables the virtual content to be placed virtually accurately with respect to the user even when the virtual content is remote from the user (e.g., more than 5 feet, more than 10 feet, more than 15 feet, more than 20 feet, etc.) (when viewed by the user through a screen worn by the user).

[0124] As discussed with reference to FIG. 9, the processing unit 1002 includes a non-transitory medium 1040 configured to store anchor point information. As a non-limiting example, the non-transitory medium 1040 may store the positions of anchor points, different sets of anchor points associated with different users, a set of common anchor points, common anchor points selected for user localization and / or virtual content placement, etc. In other embodiments, the non-transitory medium 1040 may store other information. In some embodiments, the non-transitory medium 1040 may store different virtual content, which may be read by the graphic generator 1030 for presentation to the user. In some cases, a certain virtual content may be associated with a game application. In such a case, when the game application is activated, the processing unit 1002 may then access the non-transitory medium 1040 and obtain the corresponding virtual content for the game application. In some embodiments, the non-transitory medium may also store the game application and / or parameters associated with the game application.

[0125] In addition, as disclosed herein, in some embodiments, the virtual content may be a movable object that moves within the screen based on an orbit model. The orbit model may be stored within the non-transitory medium 1040 in some embodiments. In some embodiments, the orbit model may be a straight line. In such a case, when the orbit model is applied to the movement of the virtual object, the virtual object will move within the straight line path defined by the straight line of the orbit model. As another example, the orbit model may be a parabola equation that defines a path based on the initial velocity Vo and the initial movement direction of the virtual object, and also based on the weight of the virtual object. Thus, different virtual objects with different individually assigned weights will move along different parabolic paths.

[0126] The non-transitory medium 1040 may include a plurality of memory units that are not limited to a single memory unit and may be integrated or separated, but are communicatively connected (e.g., wirelessly or by a wire).

[0127] In some embodiments, as the virtual object moves virtually through the physical environment, the processing unit 1002 tracks the position of the virtual object relative to one or more objects identified within the physical environment. In some cases, when the virtual object contacts or approaches a physical object, the graphics generator 1030 may generate graphics indicating an interaction between the virtual object and the physical object within the environment. For example, the graphics may indicate that the virtual object is deflected from a physical object (e.g., a wall) or another virtual object by changing the virtual object's path of travel. As another example, when the virtual object contacts or approaches a physical object (e.g., a wall) or another virtual object, the graphics generator 1030 may place an interaction image spatially associated with the location where the virtual object contacts the physical object or another virtual object. The interaction image may indicate that the wall has cracks, dents, scratches, stains, etc.

[0128] In some embodiments, different interaction images may be stored in the non-transitory medium 1040 and / or may be stored in a server that communicates with the processing unit 1002. The interaction images may be stored in association with one or more attributes related to the interaction of two objects. For example, an image of wrinkles may be stored in association with the attribute "blanket". In such a case, when a virtual object is displayed as being supported on a physical object that can be identified as a "blanket", when viewed through the screen of the image display device 2, the graphic generator 1030 may display an image of wrinkles between the virtual object and the physical object such that the virtual object appears to cause the wrinkles on the blanket by sitting on top of the blanket.

[0129] Note that the virtual content that can be virtually displayed with respect to the physical environment based on one or more anchor points is not limited to the described examples, and the virtual content may be other items. Also, as used herein, the term "virtual content" is not limited to virtualized physical items, and may refer to the virtualization of any item such as virtualized energy (e.g., laser beam, sound wave, energy wave, heat, etc.). The term "virtual content" may also refer to any content such as text, symbols, comics, animations, etc.

[0130] Task Assistant

[0131] As shown in FIG. 9, the processing unit 1002 also includes a task assistant 1060. The task assistant 1060 of the processing unit 1002 is configured to receive one or more sensor information based on one or more sensor inputs and assist a user of the image display device in performing an objective with virtual content. For example, in some embodiments, the one or more sensor inputs may indicate the user's line of sight direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing. In some embodiments, the processing unit 1002 is configured to assist the user in performing the objective by applying one or more limits regarding the position and / or angular velocity of system components. Alternatively, or in addition, the processing unit 1002 may be configured to assist the user in performing the objective by gradually reducing the distance between the virtual content and another element. Using the above pancake tossing game, described with reference to FIGS. 5A-5L as an example, the processing unit 1002 may detect that the user is attempting to catch the pancake 64 based on the movement of the just-occurred controller 4, based on the current direction and speed of the controller 4, and / or based on the trajectory of the pancake 64. In such a case, the task assistant 1060 may gradually reduce the distance between the pancake 64 and the frying pan 62 by moving the pancake 64 away from the determined trajectory so that the pancake 64 approaches the frying pan 62, by deviating from the determined trajectory with respect to the pancake 64, and the like. Alternatively, the task assistant 1060 may discretely increase the size of the pancake 64 (i.e., computationally, not graphically), and / or increase the size of the frying pan 62 (i.e., computationally, not graphically), thereby enabling the user to more easily catch the pancake 64 with the frying pan 62.

[0132] In some embodiments, assisting the user in performing a task involving virtual content may be performed in response to meeting a criterion. For example, using the pancake catching game described herein, in some embodiments, the processing unit 1002 may be configured to determine (e.g., predict) whether the user will approach a state of catching the pancake 64 based on the trajectory of the moving pancake 64 and the movement trajectory of the controller 4 (e.g., within a distance threshold such as 5 inches, 3 inches, 1 inch, etc.). If so, the task assistant 1060 will control the graphic generator 1030 to output a graphic indicating that the pancake 64 is caught by the frying pan 62. On the other hand, if the processing unit 1002 determines (e.g., predicts) that the user will not approach a state of catching the pancake 64 with the frying pan 62, the task assistant 1060 will not perform any action to assist the user in performing the task.

[0133] Note that the task that the task assistant 1060 can assist the user in performing is not limited to the example of catching a flying virtual object. In other embodiments, the task assistant 1060 may assist the user in performing other tasks when the processing unit 1002 determines (e.g., predicts) that the task will be very close to being completed (e.g., exceeding 80%, 85%, 90%, 95%, etc.). For example, in other embodiments, the task may involve the user throwing or shooting a virtual object at a destination, such as an object (e.g., a target at a shooting range) through an opening (e.g., a basketball hoop) at another user.

[0134] In other embodiments, the task assistant 1060 is optional and the processing unit does not include the task assistant 1060.

[0135] Methods implemented by the processing unit and / or an application within the processing unit

[0136] Figure 10 illustrates method 1100 according to some embodiments. Method 1100 may be implemented by an apparatus configured to provide virtual content within a virtual or augmented environment. Also, in some embodiments, method 1100 may be implemented by an apparatus configured to provide virtual content within a virtual or augmented reality environment in which a first user wearing a first display screen and a second user wearing a second display screen can interact with each other. In some embodiments, each image display device may be the image display device 2. In some embodiments, method 1100 may be implemented by any of the image display devices described herein or by a plurality of image display devices. Also, in some embodiments, at least a portion of method 1100 may be implemented by processing unit 1002 or by a plurality of processing units (e.g., processing units within individual image display devices). Further, in some embodiments, method 1100 may be implemented by an apparatus separate from a server or an image display device worn by an individual user.

[0137] As shown in FIG. 10, method 1100 includes obtaining a first position of a first user (item 1102), determining a first set of one or more anchor points based on the first position of the first user (item 1104), obtaining a second position of a second user (item 1106), determining a second set of one or more anchor points based on the second position of the second user (item 1108), determining one or more common anchor points that are within both the first set and the second set (item 1110), and providing virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points (item 1112).

[0138] Optionally, in method 1100, one or more common anchor points comprise a plurality of common anchor points, and the method further includes selecting a subset of the common anchor points from the plurality of common anchor points.

[0139] Optionally, in method 1100, the subset of common anchor points is selected to reduce the relative positioning error between the first user and the second user.

[0140] Optionally, in method 1100, one or more common anchor points comprise a single common anchor point.

[0141] Optionally, method 1100 further includes determining a position and / or orientation of virtual content based on at least one of one or more common anchor points.

[0142] Optionally, in method 1100, each of one or more anchor points within the first set is a point within a persistent coordinate frame (PCF).

[0143] Optionally, in method 1100, the virtual content is provided as a movable virtual object within the first display screen and / or the second display screen for display.

[0144] Optionally, in method 1100, the virtual object is provided within the first display screen for display such that the virtual object appears to move within the space between the first user and the second user.

[0145] Optionally, in method 1100, one or more common anchor points include a first common anchor point and a second common anchor point, and the movable virtual object is provided within the first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen. The first object position of the movable virtual object is based on the first common anchor point, and the second object position of the movable virtual object is based on the second common anchor point.

[0146] Optionally, method 1100 further includes selecting a first common anchor point for placing the virtual object at the first object position based on where the action of the virtual object occurs.

[0147] Optionally, in method 1100, one or more common anchor points include a single common anchor point, and the movable virtual object is provided within the first display screen for display such that the movable virtual object has a first object position relative to the first display screen and a second object position relative to the first display screen. The first object position of the movable virtual object is based on the single common anchor point, and the second object position of the movable virtual object is based on the single common anchor point.

[0148] Optionally, in method 1100, one or more common anchor points include a plurality of common anchor points, and the method further includes selecting one of the common anchor points for placing the virtual content within the first display screen.

[0149] Optionally, in method 1100, the act of selecting includes selecting one of the common anchor points that is closest to the action of the virtual content or within a distance threshold from the action of the virtual content.

[0150] Optionally, in method 1100, the position and / or movement of virtual content can be controlled by a first handheld device of a first user.

[0151] Optionally, in method 1100, the position and / or movement of virtual content can also be controlled by a second handheld device of a second user.

[0152] Optionally, method 1100 further includes localizing a first user and a second user with respect to the same mapping information based on one or more common anchor points.

[0153] Optionally, method 1100 further includes displaying virtual content on a first display screen such that the virtual content appears in a spatial relationship to physical objects within the surrounding environment of the first user.

[0154] Optionally, method 1100 further includes obtaining one or more sensor inputs and assisting the first user in performing an objective with the virtual content based on the one or more sensor inputs.

[0155] Optionally, in method 1100, the one or more sensor inputs indicate the line-of-sight direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing of the first user.

[0156] Optionally, in method 1100, the act of assisting the first user in performing an objective includes applying one or more limits regarding the position and / or angular velocity of system components.

[0157] Optionally, in method 1100, the act of assisting the first user in performing an objective includes gradually reducing the distance between the virtual content and another element.

[0158] Optionally, in method 1100, the apparatus comprises a first processing part that communicates with a first display screen and a second processing part that communicates with a second display screen.

[0159] In some embodiments, method 1100 may be implemented in response to a processing unit executing instructions stored in a non-transitory medium. Thus, in some embodiments, the non-transitory medium contains the stored instructions, and their execution by the processing unit causes a method to be implemented. The processing unit may be part of an apparatus configured to provide virtual content within a virtual or augmented reality environment in which a first user and a second user can interact with each other. The method (caused to be implemented by the processing unit executing instructions) includes obtaining a first position of the first user, determining a first set of one or more anchor points based on the first position of the first user, obtaining a second position of the second user, determining a second set of one or more anchor points based on the second position of the second user, determining one or more common anchor points that are in both the first set and the second set, and providing virtual content for an experience by the first user and / or the second user based on at least one of the one or more common anchor points.

[0160] Special processing system

[0161] In some embodiments, the method 1100 described herein may be implemented by a system 1 (e.g., a processing unit 1002) that executes an application, or by the application. The application may contain a set of instructions. In one implementation, a special processing system may be provided that has a non-transitory medium storing the set of instructions for the application. Execution of the instructions by the processing unit 1102 of the system 1 will cause the processing unit 1102 and / or the image display device 2 to implement the features described herein. For example, in some embodiments, execution of the instructions by the processing unit 1102 will cause the method 1100 to be implemented.

[0162] In some embodiments, the system 1, the image display device 2, or the apparatus 7 may also be regarded as a special processing system. In particular, the system 1, the image display device 2, or the apparatus 7 is a special processing system in that it contains instructions stored in its non-transitory medium for execution by the processing unit 1102 to provide a unique tangible effect in the real world. The features provided by the image display device 2 (as a result of the processing unit 1102 executing instructions) provide improvements within the technical fields of augmented reality and virtual reality.

[0163] FIG. 11 is a block diagram illustrating an embodiment of a special processing system 1600 that may be used to implement the various features described herein. For example, in some embodiments, the processing system 1600 may be used to implement at least a portion of the system 1, such as the image display device 2, the processing unit 1002, etc. Also, in some embodiments, the processing system 1600 may be used to implement the processing unit 1102 or one or more components therein (e.g., the positioner 1020, the graphic generator 1030, etc.).

[0164] The processing system 1600 includes a bus 1602 or other communication mechanism for communicating information, and a processor 1604 coupled to the bus 1602 for processing information. The processor system 1600 also includes a main memory 1606, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 1602 for storing information and instructions to be executed by the processor 1604. The main memory 1606 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 1604. The processor system 1600 further includes a read-only memory (ROM) 1608 or other static storage device coupled to the bus 1602 for storing static information and instructions for the processor 1604. A data storage device 1610, such as a magnetic disk, solid state disk, or optical disk, is provided and coupled to the bus 1602 for storing information and instructions.

[0165] The processor system 1600 may be coupled via the bus 1602 to a display 1612, such as a screen, for displaying information to a user. In some cases, where the processing system 1600 is part of a device that includes a touch screen, the display 1612 may be a touch screen. An input device 1614, including alphanumeric and other keys, is coupled to the bus 1602 for communicating information and command selections to the processor 1604. Another type of user input device is a cursor control 1616, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 1604 and for controlling cursor movement on the display 1612. This input device typically has two degrees of freedom in two axes, i.e., a first axis (e.g., x) and a second axis (e.g., y), which allows the device to define a position within a plane. In some cases, where the processing system 1600 is part of a device that includes a touch screen, the input device 1614 and the cursor control may be the touch screen.

[0166] In some embodiments, the processor system 1600 can be used to implement the various functions described herein. According to some embodiments, such use is provided by the processor system 1600 in response to the processor 1604 executing one or more sequences of one or more instructions contained within the main memory 1606. Those skilled in the art will appreciate how to prepare such instructions based on the functions and methods described herein. Such instructions may be read into the main memory 1606 from another processor-readable medium such as the storage device 1610. Execution of the sequence of instructions contained within the main memory 1606 causes the processor 1604 to perform the process steps described herein. One or more processors in a multi-processing arrangement may also be employed to execute the sequence of instructions contained within the main memory 1606. In alternative embodiments, a wired circuitry configuration may be used, instead of or in combination with software instructions, to implement the various embodiments described herein. Accordingly, embodiments are not limited to any specific combination of hardware circuitry and software.

[0167] The term "processor-readable medium" as used herein refers to any medium involved in providing instructions to the processor 1604 for execution. Such a medium may take many forms including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical, solid-state, or magnetic disks such as the storage device 1610. Non-volatile media may be regarded as an example of non-transitory media. Volatile media includes dynamic memory such as the main memory 1606. Volatile media may be regarded as an example of non-transitory media. Transmission media includes coaxial cables, copper wire, and fiber optics including the wires that make up the bus 1602. Transmission media can also take the form of acoustic or light waves such as those generated during radio-wave and infrared data communications.

[0168] A processor-readable medium in general form includes, for example, a flexible disk, a hard disk, a magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, a solid state disk, any other memory chip or cartridge, a carrier wave as described hereinafter, or any other medium from which a processor can read.

[0169] Various forms of processor-readable media may be involved in carrying one or more sequences of one or more instructions to processor 1604 for execution. For example, the instructions may first be carried on a magnetic disk or solid state disk of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions via a network such as the Internet. Processing system 1600 can receive data regarding the network line. Bus 1602 carries the data to main memory 1606, from which processor 1604 reads and executes the instructions. The instructions received by main memory 1606 may optionally be stored on storage device 1610 either before or after execution by processor 1604.

[0170] Processing system 1600 also includes a communication interface 1618 coupled to bus 1602. Communication interface 1618 provides a two-way data communication coupling to network link 1620 connected to local network 1622. For example, communication interface 1618 may be a local area network (LAN) card for providing a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 1618 transmits and receives electrical, electromagnetic, or optical signals that carry a data stream representing various types of information.

[0171] Network link 1620 typically provides data communication to other devices through one or more networks. For example, network link 1620 may provide a connection to host computer 1624 or device 1626 through local network 1622. The data stream transported via network link 1620 can include electrical, electromagnetic, or optical signals. Signals through various networks, and signals on network link 1620 and through communication interface 1618 that carry data to and from processing system 1600 are exemplary forms of carrier waves that transport information. Processing system 1600 can send messages and receive data, including program code, through the network, network link 1620, and communication interface 1618.

[0172] Note that the term "image" as used herein can refer to both displayed images and / or images that are not in a display form (e.g., images or image data stored in a medium or being processed).

[0173] Also, as used herein, the term "action" of virtual content is not limited to moving virtual content, but can also refer to stationary virtual content that can be moved (e.g., virtual content that can be "dragged" or is being "dragged" by a user using a pointer), or any virtual content on or beside which an action can be performed.

[0174] Various exemplary embodiments are described herein. These examples are referenced in a non-limiting sense. They are provided to illustrate more broadly applicable aspects of the invention. Various changes may be made to the described embodiments, and equivalents may be substituted without departing from the true spirit and scope of the invention. Additionally, many modifications may be made to adapt a particular situation, material, composition, process, process act, or step to the purposes, spirit, or scope of the invention. Further, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein can be readily separated from or combined with features of any of the other several embodiments without departing from the scope or spirit of the invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.

[0175] This embodiment includes a method that can be implemented using the device of the subject matter. The method may comprise the act of providing such a suitable device. Such providing may be performed by an end user. In other words, the act of "providing" simply requires that the end user act to obtain, access, approach, locate, configure, activate, power on, or otherwise provide the device required in the method of the subject matter. The methods recited herein may be performed in any order of the recited logical events and in the recited order of events.

[0176] Unless such exclusive terms are used, the term "comprising" in the claims associated with the present invention shall be taken to allow the inclusion of any additional elements, whether or not a given number of elements are recited in the claims, or the addition of features may be regarded as transforming the nature of the elements recited in such claims. Unless specifically defined herein, all technical and scientific terms used herein shall be given the broadest generally understood meaning possible while maintaining the validity of the claims.

[0177] Exemplary aspects of the present disclosure have been described above, along with details regarding material selection and manufacturing. Regarding other details of the present disclosure, these are understood in relation to the patents and publications referenced above and are generally known or understandable by those skilled in the art. The same can apply to the method-based aspects of the present disclosure from the perspective of additional operations that are generally or logically employed.

[0178] In addition, the present disclosure has been described with reference to several embodiments that optionally incorporate various features, but the present disclosure is not limited to what is described or illustrated as being considered for each variation of the disclosure. Various modifications may be made to the present disclosure described, and equivalents (whether or not recited herein or included for some brevity purposes) may be substituted without departing from the true spirit and scope of the present disclosure. Additionally, when a range of values is provided, it is to be understood that all intervening values between the upper and lower limits of that range and any other recited values or intervening values within the recited range are included within the present disclosure.

[0179] Also, any optional feature of the described variations of the present invention may be considered to be described and claimed independently or in combination with any one or more of the features described herein. References to singular items include the possibility that multiple identical items exist. More specifically, as used in this specification and the claims associated herewith, the singular forms "a", "an", "said", and "the" include plural references unless specifically stated otherwise. Further, note that any claim may be drafted to exclude any optional element. Thus, the text is intended to serve as a preamble for the use of exclusive terms such as "merely", "only", and equivalents in connection with the recitation of claim elements, or for the use of "negative" limitations.

[0180] In addition, as used herein, a phrase referring to a list of items "at least one of" refers to one item or any combination of items, including a single element. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Connective phrases such as "at least one of X, Y, and Z" are generally understood in a context such that, unless specifically described otherwise, they are used to convey that an item, term, etc. can be at least one of X, Y, or Z. Thus, such connective phrases are generally not intended to suggest that an embodiment requires that at least one of each of X, at least one of Y, and at least one of Z be present.

[0181] The scope of the present disclosure should not be limited to the provided examples and / or this specification, but rather should be limited only by the scope of the terms of the claims associated with the present disclosure.

[0182] In the foregoing specification, the present disclosure has been described with reference to its specific embodiments. However, it will be apparent that various modifications and changes may be made therein without departing from the broader spirit and scope of the present disclosure. For example, the foregoing process flow is described with reference to a particular order of process actions. However, many of the orders of the process actions described may be changed without affecting the scope or operation of the present disclosure. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a limiting sense.

Claims

1. 1. An apparatus for providing virtual content in an environment in which a first user and a second user may interact with each other, the apparatus comprising: a communications interface configured to communicate with a first display screen worn by the first user and / or a second display screen worn by the second user; and A processing unit, the processing unit comprising: obtaining a first location of the first user; determining a first set of anchor points based on a first location of the first user; obtaining a second location of the second user; determining a second set of anchor points based on a second location of the second user; identifying anchor points in both the first set and the second set as common anchor points, the common anchor points comprising a first common anchor point and a second common anchor point; providing the virtual content for experience by the first user and / or the second user based on at least one of the common anchor points; [0023] The processing unit is configured to provide the virtual content for display within the first display screen and / or the second display screen as a movable virtual object along a trajectory such that the virtual content moves from a first location closer to the first common anchor point than the second common anchor point to a second location closer to the second common anchor point than the first common anchor point, the processing unit determining the first location of the virtual content based on the first common anchor point for display of the virtual content at a first time point; and determining a second location of the virtual content based on the second common anchor point for display of the virtual content at a second time point, the first location of the virtual content corresponding to a first portion of the trajectory and the second location of the virtual content corresponding to a second portion of the trajectory, the first portion of the trajectory comprising the first location associated with the first user and based on the first common anchor point away from the trajectory, and the second portion of the trajectory comprising the second location associated with the second user and based on the second common anchor point away from the trajectory. a processing unit configured to use the second common anchor point to present the moveable virtual object to the first user when the moveable virtual object is on the second portion of the trajectory to prevent inaccurate placement of the moveable virtual object relative to the environment due to a local coordinate frame associated with the first user not being correctly aligned with an environmental feature closer to the second user than the first user; and An apparatus comprising:

2. The apparatus of claim 1, wherein the processing unit is configured to select a subset of the common anchor points.

3. The apparatus of claim 2 , wherein the processing unit is configured to select a subset of the common anchor points to reduce a location error of the first user and the second user relative to one another.

4. The apparatus of claim 1 , wherein the processing unit is configured to position and / or orient the virtual content based on at least one of the common anchor points.

5. The apparatus of claim 1 , wherein each of the anchor points in the first set is a point in a persistent coordinate frame (PCF).

6. 2. The device of claim 1, wherein the processing unit is configured to provide the moveable virtual object for display within the first display screen such that the moveable virtual object appears to be moving within a space between the first user and the second user.

7. The apparatus of claim 1 , wherein the position and / or movement of the virtual content is controllable by a first handheld device of the first user.

8. The apparatus of claim 7 , wherein the position and / or movement of the virtual content is also controllable by a second handheld device of the second user.

9. The apparatus of claim 1 , wherein the processing unit is configured to locate the first user and the second user relative to the same mapping information based on the common anchor point.

10. 2. The device of claim 1, wherein the processing unit is configured to cause the first display screen to display the virtual content such that the virtual content would appear in a certain spatial relationship to physical objects in the first user's surrounding environment.

11. the processing unit is configured to acquire one or more sensor inputs; the processing unit is configured to assist the first user in accomplishing an objective involving the virtual content based on the one or more sensor inputs.

2. The apparatus of claim 1.

12. 12. The apparatus of claim 11, wherein the one or more sensor inputs are indicative of the first user's eye gaze direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing.

13. 12. The apparatus of claim 11, wherein the processing unit is configured to assist the first user in achieving the objective by applying one or more limits on positions and / or angular velocities of system components.

14. 12. The apparatus of claim 11, wherein the processing unit is configured to assist the first user in achieving the goal by gradually reducing a distance between the virtual content and another virtual content, the goal comprising catching the virtual content by the first user.

15. 10. The apparatus of claim 1, wherein the processing unit comprises a first processing portion in communication with the first display screen and a second processing portion in communication with the second display screen.

16. 1. A method implemented by an apparatus configured to provide virtual content in an environment in which a first user wearing a first display screen and a second user wearing a second display screen may interact with each other, the apparatus comprising a processing unit, the method comprising: obtaining a first location of the first user using the processing unit; determining, using the processing unit, a first set of anchor points based on a first position of the first user; obtaining a second location of the second user using the processing unit; and determining, using the processing unit, a second set of anchor points based on a second location of the second user; using the processing unit to identify anchor points in both the first set and the second set as common anchor points, the common anchor points comprising a first common anchor point and a second common anchor point; providing, using the processing unit, the virtual content for experience by the first user and / or the second user based on at least one of the common anchor points; Including, The processing unit is configured to provide the virtual content for display within the first display screen and / or the second display screen as a movable virtual object along a trajectory such that the virtual content moves from a first location closer to the first common anchor point than the second common anchor point to a second location closer to the second common anchor point than the first common anchor point, the processing unit determining the first location of the virtual content based on the first common anchor point for display of the virtual content at a first time point; and determining a second location of the virtual content based on the second common anchor point for display of the virtual content at a second time point, the first location of the virtual content corresponding to a first portion of the trajectory and the second location of the virtual content corresponding to a second portion of the trajectory, the first portion of the trajectory comprising the first location associated with the first user and based on the first common anchor point away from the trajectory, and the second portion of the trajectory comprising the second location associated with the second user and based on the second common anchor point away from the trajectory. the processing unit is configured to use the second common anchor point to present the moveable virtual object to the first user when the moveable virtual object is on the second portion of the trajectory to prevent inaccurate placement of the moveable virtual object relative to the environment due to a local coordinate frame associated with the first user not being correctly aligned to environmental features closer to the second user than the first user.

17. The method of claim 16, further comprising using the processing unit to select a subset of the common anchor points.

18. The method of claim 17 , wherein the subset of common anchor points is selected to reduce a location error of the first user and the second user relative to one another.

19. The method of claim 16 , further comprising determining a position and / or an orientation for the virtual content based on at least one of the common anchor points.

20. The method of claim 16 , wherein each of the anchor points in the first set is a point in a persistent coordinate frame (PCF).

21. 17. The method of claim 16, wherein the moveable virtual object is provided for display within the first display screen such that the moveable virtual object appears to be moving within a space between the first user and the second user.

22. The method of claim 16 , wherein the position and / or movement of the virtual content is controllable by a first handheld device of the first user.

23. The method of claim 22 , wherein the position and / or movement of the virtual content is also controllable by a second handheld device of the second user.

24. The method of claim 16 , further comprising locating the first user and the second user to the same mapping information based on the common anchor point.

25. 17. The method of claim 16, further comprising displaying, by the first display screen, the virtual content such that the virtual content appears in a spatial relationship to physical objects in the first user's surrounding environment.

26. Obtaining one or more sensor inputs; assisting the first user in accomplishing an objective involving the virtual content based on the one or more sensor inputs.

20. The method of claim 16, further comprising:

27. 27. The method of claim 26, wherein the one or more sensor inputs are indicative of the first user's eye gaze direction, upper limb kinematics, body position, body orientation, or any combination of the foregoing.

28. 27. The method of claim 26, wherein the action of assisting the first user in achieving the goal comprises applying one or more limits on positions and / or angular velocities of system components.

29. 27. The method of claim 26, wherein the act of assisting the first user in accomplishing the goal includes gradually reducing a distance between the virtual content and another virtual content, and the goal includes catching the virtual content by the first user.

30. 17. The method of claim 16, wherein the device comprises a first processing portion in communication with the first display screen and a second processing portion in communication with the second display screen.

31. 1. A processor-readable non-transitory medium storing a set of instructions, execution of the set of instructions by a processing unit causing a method to be performed, the processing unit being part of an apparatus configured to provide virtual content within an environment in which a first user and a second user may interact with each other, the method comprising: obtaining a first location of the first user; determining a first set of anchor points based on a first location of the first user; obtaining a second location of the second user; determining a second set of anchor points based on a second location of the second user; identifying anchor points in both the first set and the second set as common anchor points, the common anchor points comprising a first common anchor point and a second common anchor point; providing the virtual content for experience by the first user and / or the second user based on at least one of the common anchor points; Including, The processing unit is configured to provide the virtual content for display within the first display screen and / or the second display screen as a movable virtual object along a trajectory such that the virtual content moves from a first location closer to the first common anchor point than the second common anchor point to a second location closer to the second common anchor point than the first common anchor point, the processing unit determining the first location of the virtual content based on the first common anchor point for display of the virtual content at a first time point; and determining a second location of the virtual content based on the second common anchor point for display of the virtual content at a time point, the first location of the virtual content corresponding to a first portion of the trajectory, the second location of the virtual content corresponding to a second portion of the trajectory, the first portion of the trajectory comprising the first location associated with the first user and based on the first common anchor point away from the trajectory, the second portion of the trajectory comprising the second location associated with the second user and based on the second common anchor point away from the trajectory. the processing unit is configured to use the second common anchor point to present the moveable virtual object to the first user when the moveable virtual object is in the second portion of the trajectory to prevent inaccurate placement of the moveable virtual object relative to the environment due to a local coordinate frame associated with the first user not being correctly aligned to an environmental feature closer to the second user than the first user.

Citation Information

Patent Citations

  • Composite reality simulation device and composite reality simulation program

    JP2019008473A

  • Remote communication system and the like

    JP2020035392A

  • A cross reality system

    WO2020036898A1