Hand lock rendering of virtual objects in artificial reality
Patent Information
- Application Number
- CN202180094974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2021-12-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-12-30
Smart Images

Figure CN116897326B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to generating graphics for artificial reality environments. Background Technology
[0002] Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. This artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in single-channel or multi-channel (e.g., stereoscopic video providing a three-dimensional effect to the viewer). Artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or for use in artificial reality (e.g., performing activities in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers. Summary of the Invention
[0003] In one aspect of the invention, a method is provided, comprising: using one or more computing devices to: determine a user's first viewpoint and a first hand pose of the user's hand based on first tracking data associated with a first time; generate a virtual object in a virtual environment based on the first hand pose and a predetermined spatial relationship between the virtual object and the user's hand; render a first image of the virtual object as seen from the first viewpoint; determine a user's second viewpoint and a second hand pose of the hand based on second tracking data associated with a second time; adjust the first image of the virtual object based on the change from the first hand pose to the second hand pose; render a second image based on the adjusted first image as seen from the second viewpoint; and display the second image.
[0004] In some embodiments, the predetermined spatial relationship between the virtual object and the user's hand can be based on one or more anchor points relative to the user's hand.
[0005] The method may further include: determining a first head pose of the user's head based on first tracking data associated with a first time; and wherein virtual objects generated in the virtual environment may also be based on the first head pose.
[0006] The method may further include: determining a second head pose of the user's head based on second tracking data associated with a second time; and wherein adjusting the first image of the virtual object may also be based on the change from the first head pose to the second head pose.
[0007] Adjusting the first image of the virtual image may include: calculating a three-dimensional transformation of the first image of the virtual object in the virtual environment based on the change from a first hand pose to a second hand pose; and applying the three-dimensional transformation to the first image of the virtual object in the virtual environment; and wherein rendering the second image based on the adjusted first image may include: generating a projection of the adjusted first image of the virtual object based on a second viewpoint.
[0008] Adjusting a first image of a virtual object may include: calculating a visual deformation of the first image of the virtual object based on a change from a first hand pose to a second hand pose; and applying the visual deformation to the first image of the virtual object; and wherein rendering a second image based on the adjusted first image may include: rendering the deformed first image of the virtual object based on a second viewpoint.
[0009] Visual distortion can include: scaling the first image of the virtual object; tilting the first image of the virtual object; rotating the first image of the virtual object; or shearing the first image of the virtual object.
[0010] Rendering a second image based on an adjusted first image seen from a second viewpoint may include updating the visual appearance of the first image of the virtual object based on changes in hand pose from a first hand pose to a second hand pose.
[0011] The method may further include: determining a portion of the virtual object that is obscured by a hand from a second viewpoint; and wherein rendering a second image based on an adjusted first image viewed from the second viewpoint includes: generating instructions to not display the portion of the second image corresponding to the portion of the virtual object obscured by the hand.
[0012] Determining the portion of a virtual object obscured by a hand from a second viewpoint may include: generating a virtual object representation of the hand in a virtual environment, the position and orientation of which are based on second tracking data and a second viewpoint; projecting a ray having an origin and trajectory based on the second viewpoint into the virtual environment; and determining the intersection point of the ray with the virtual object representation of the hand, wherein the ray intersects the virtual object representation of the hand before intersecting with another virtual object.
[0013] A first image of a virtual object viewed from a first viewpoint may be rendered by a first computing device among one or more computing devices; and a second image based on the adjusted first image viewed from a second viewpoint may be rendered by a second computing device among one or more computing devices.
[0014] One of the computing devices may include a head-mounted display.
[0015] The method may also include allocating multiple steps of the method between a computing device including a head-mounted display and another computing device among one or more computing devices, based on one or more metrics of available computing resources.
[0016] The method may further include: after rendering a second image based on an adjusted first image viewed from a second viewpoint; determining a third viewpoint and a third hand pose of the user based on third tracking data associated with a third time; adjusting the first image of the virtual object based on the change from the first hand pose to the third hand pose; rendering a third image based on the adjusted first image viewed from a third viewpoint; and displaying the third image.
[0017] In one aspect of the invention, a system is provided, comprising: one or more processors; and one or more computer-readable non-transitory storage media communicating with the one or more processors and including instructions configured, when executed by the one or more processors, to cause the system to perform a plurality of operations: the plurality of operations including: determining a first viewpoint of a user and a first hand pose of a user's hand based on first tracking data associated with a first time; generating a virtual object in a virtual environment based on the first hand pose and a predetermined spatial relationship between a virtual object and the user's hand; rendering a first image of the virtual object as seen from the first viewpoint; determining a second viewpoint of a user and a second hand pose of a user's hand based on second tracking data associated with a second time; adjusting the first image of the virtual object based on a change from the first hand pose to the second hand pose; rendering a second image based on the adjusted first image as seen from the second viewpoint; and displaying the second image.
[0018] The predetermined spatial relationship between a virtual object and the user's hand can be based on one or more anchor points relative to the user's hand.
[0019] The instruction can also be configured to cause the system to perform a plurality of operations, including: determining a first head pose of the user’s head based on first tracking data associated with a first time; and wherein generating virtual objects in the virtual environment can also be based on the first head pose.
[0020] In one aspect of the invention, one or more computer-readable non-transitory storage media are provided, the one or more computer-readable non-transitory storage media comprising instructions configured, when executed by one or more processors of a computing system, to cause the one or more processors to perform a plurality of operations: the plurality of operations including: determining a user's first viewpoint and a user's first hand pose based on first tracking data associated with a first time; generating a virtual object in a virtual environment based on the first hand pose and a predetermined spatial relationship between a virtual object and the user's hand; rendering a first image of the virtual object as seen from the first viewpoint; determining a user's second viewpoint and a second hand pose based on second tracking data associated with a second time; adjusting the first image of the virtual object based on a change from the first hand pose to the second hand pose; rendering a second image based on the adjusted first image as seen from the second viewpoint; and displaying the second image.
[0021] The predetermined spatial relationship between a virtual object and the user's hand can be based on one or more anchor points relative to the user's hand.
[0022] The instructions can also be configured to cause one or more processors to perform a plurality of operations, the plurality of operations further including: determining a first head pose of the user’s head based on first tracking data associated with a first time; and wherein generating virtual objects in the virtual environment can also be based on the first head pose.
[0023] Various embodiments of the present invention may include an artificial reality system or a combination thereof. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. The artificial reality may, for example, include virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in single-channel or multi-channel (e.g., stereoscopic video providing a three-dimensional effect to the viewer). Furthermore, in certain embodiments, the artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in the artificial reality and / or for use in the artificial reality (e.g., performing activities in the artificial reality). Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including head-mounted displays (or any other hardware platform) connected to a host computer system.
[0024] The various embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited to these embodiments. Specific embodiments may include, or may not include, all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the invention are specifically disclosed in the appended claims for methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen merely for formal reasons. However, protection may also be claimed for any subject matter arising from the intentional reference (in particular multiple dependencies) to any plurality of prior claims, thereby disclosing any combination of multiple claims and their features, and protection may be claimed for any combination of multiple claims and their features regardless of the dependencies chosen in the appended claims. The subject matter for which protection can be claimed includes not only various combinations of the features set forth in the appended claims, but also any other combination of the features in the appended claims, wherein each feature mentioned in the appended claims may be combined with any other feature in the appended claims, or with a combination of multiple other features. Furthermore, protection may be claimed in a single claim for any embodiment and feature of the multiple embodiments and features described or depicted herein, and / or protection may be claimed for any combination of any embodiment and feature of the multiple embodiments and features described or depicted herein with any embodiment or feature described or depicted herein, or protection may be claimed for any combination of any embodiment and feature of the multiple embodiments and features described or depicted herein with any feature of the appended claims. Attached Figure Description
[0025] Figure 1A An example artificial reality system is shown.
[0026] Figure 1B An example eye display system of a headset system is shown.
[0027] Figure 2 A system diagram of the display engine is shown.
[0028] Figure 3 Example images viewed through an artificial reality system are shown.
[0029] Figure 4 Example images viewed through an artificial reality system are shown.
[0030] Figures 5A to 5BThis demonstrates an example method for providing hand-locked rendering of virtual objects using subframe rendering in artificial reality.
[0031] Figure 6 An example is shown where the first image is adjusted to be used to render the second image.
[0032] Figure 7 An example is shown where the first image is adjusted to be used to render the second image.
[0033] Figure 8 A visual representation of the image used to generate the virtual environment is shown.
[0034] Figure 9 An example computer system is shown. Detailed Implementation
[0035] In a particular embodiment, a method is performed by one or more computing systems of an artificial reality system. These computing systems may be included in a head-mounted display or a less portable computing system. The method includes: determining a user's first viewpoint and a first hand pose of the user's hand based on first tracking data associated with a first time. The method may include: generating a virtual object in a virtual environment based on the first hand pose and a predetermined spatial relationship between the virtual object and the user's hand. The method may include: rendering a first image of the virtual object as seen from the first viewpoint. The method may include: determining a user's second viewpoint and a second hand pose of the hand based on second tracking data associated with a second time. The method may include: adjusting the first image of the virtual object based on changes from the first hand pose to the second hand pose. The method may include: rendering a second image based on the adjusted first image as seen from the second viewpoint. The method may include: displaying the second image. In a particular embodiment, the predetermined spatial relationship between the virtual object and the user's hand may be based on one or more anchor points relative to the user's hand.
[0036] In a particular embodiment, the method may further include: determining a first head pose of the user's head based on first tracking data associated with a first time. Generating virtual objects in the virtual environment may also be based on this first head pose. In a particular embodiment, the method may further include: determining a second head pose of the user's head based on second tracking data associated with a second time. Adjusting a first image of the virtual object may also be based on a change from the first head pose to the second head pose. In a particular embodiment, adjusting the first image of the virtual object may include: calculating a three-dimensional transformation of the first image of the virtual object in the virtual environment based on a change from a first hand pose to a second hand pose; and applying the three-dimensional transformation to the first image of the virtual object in the virtual environment. Rendering a second image based on the adjusted first image may include: generating a projection of the adjusted first image of the virtual object based on a second viewpoint.
[0037] In a particular embodiment, adjusting the first image of the virtual object may include: calculating a visual deformation of the first image of the virtual object based on a change from a first hand pose to a second hand pose; and applying the visual deformation to the first image of the virtual object. Rendering a second image based on the adjusted first image may include: rendering the deformed first image of the virtual object based on a second viewpoint. In a particular embodiment, the visual deformation may include: scaling the first image of the virtual object; tilting the first image of the virtual object; rotating the first image of the virtual object; or shearing the first image of the virtual object.
[0038] In a particular embodiment, rendering a second image based on an adjusted first image viewed from a second viewpoint may include updating the visual appearance of the first image of the virtual object based on a change from a first hand pose to a second hand pose. In a particular embodiment, the method further includes determining a portion of the virtual object occluded by the hand from the second viewpoint, wherein rendering the second image based on the adjusted first image viewed from the second viewpoint may include generating instructions to not display the portion of the second image corresponding to the occluded portion of the virtual object. In a particular embodiment, determining the occluded portion of the virtual object from the second viewpoint may include generating a virtual object representation of the hand in a virtual environment, the position and orientation of which are based on second tracking data and the second viewpoint. The method may include projecting a ray having an origin and trajectory based on the second viewpoint into the virtual environment. The method may include determining the intersection point of the ray with the virtual object representation of the hand, wherein the ray intersects the virtual object representation of the hand before intersecting with another virtual object.
[0039] In a particular embodiment, a first image of a virtual object viewed from a first viewpoint may be rendered by a first computing device among one or more computing devices, and a second image based on the adjusted first image viewed from a second viewpoint may be rendered by a second computing device among the one or more computing devices. In a particular embodiment, one of the one or more computing devices may include a head-mounted display. In a particular embodiment, the method may include allocating multiple steps of the method between a computing device including a head-mounted display and another computing device among the one or more computing devices based on one or more metrics of available computing resources.
[0040] In a particular embodiment, the method may further include: after rendering a second image based on an adjusted first image viewed from a second viewpoint, determining a user's third viewpoint and a third hand pose based on third tracking data associated with a third time. The method may include: adjusting a first image of a virtual object based on changes from a first hand pose to a third hand pose. The method may include: rendering a third image based on an adjusted first image viewed from a third viewpoint. The method may include: displaying the third image.
[0041] In certain embodiments, this disclosure relates to the problem of associating virtual objects in a virtual environment with physical objects for presentation to a user. In augmented reality (AR), a virtual environment can be displayed to a user as an augmentation layer on top of a real environment. This can be achieved by creating a correspondence or mapping between viewpoints in the real environment and viewpoints in the virtual avatar. Embodiments of this disclosure relate to the task of efficiently rendering virtual objects associated with physical objects, such as a user's hand. Specifically, embodiments of this disclosure relate to a technique commonly referred to as "hand-locked rendering," by which the pose of a virtual object presented to a user is determined based on a predetermined relationship with the pose of the user's hand or the pose of parts of the user's hand. This technique can also be applied to other physical objects that can be detected by a tracking system used for virtual reality or augmented reality. Although the examples herein may be described as relating to a user's hand, those skilled in the art will recognize that the technique can be applied, by way of example and not limitation, to: other body parts of a user operating an artificial reality system; bodies of other users detected by the artificial reality system; physical objects manipulated by a user detected by the artificial reality system; or other physical objects. Therefore, in the context of the use of the phrase "hand-locked rendering" in this disclosure, the technique should be understood to include techniques for properly rendering virtual objects anchored to any of the aforementioned object categories.
[0042] In a particular embodiment, this disclosure provides a technique for a system capable of generating a hand-locked virtual object at a subframe display rate based on hand and head pose information (which can be updated at or faster than the frame generation rate). Using the rendering techniques described herein, the artificial reality system can update an image of the virtual object as information about the user's hand and head pose is updated, the image of which can be rendered at a first frame rate. Updating the image of the virtual object can be used to enable position locking or pose locking using anchor points or other predetermined relationships with the user's hand. Using the subframe generation and rendering architecture described herein, updating the image of the virtual object can occur at a rate faster than the original frame rate at which the image was generated. This faster subframe rate can also be used to display the update, thereby improving the user experience of the artificial reality system by increasing the tracking fidelity of the virtual object with respect to physical objects manipulated by the user, including the user's hand.
[0043] The combined effect of the technologies described in this paper is to allow application developers using artificial reality systems (especially those with more limited access to the programming interfaces provided by artificial reality systems than first-party developers) to develop artificial reality experiences featuring accurate and intuitive hand-locked virtual objects anchored to specific locations on the user's hand. For example, virtual objects such as rings can be anchored to the user's fingers. As another example, animation effects can be anchored to the user's fingertips. Yet another example, lighting effects can be anchored to specific locations on objects held by the user. Currently, schemes for rendering virtual objects may require all virtual objects to be oriented relative to world coordinates or a fixed origin or anchor point in the world. This makes the procedures used by virtual experience developers to attempt to create dynamic effects associated with physical objects with high rates of motion (e.g., the user's hand) extremely complex.
[0044] Furthermore, due to input and display latency, developers of artificial reality experiences often struggle to support smooth animation of virtual objects that interact with the user's hands. The complexity of virtual objects limits the rendering rate of those objects using typical methods. For example, to save power and ensure a consistent rendering rate, the rendering rate of virtual objects may be limited to an acceptable, but far from optimal, standard rate (e.g., 30 frames per second (fps)). Since it typically takes time for an artificial reality system to render a frame, by the time the frame is ready to be displayed, the pose of the physical objects (e.g., the user's hands and head) that must be considered when displaying that frame has likely changed. This means that the rendered frame may be outdated when displayed. This latency between input, rendering, and display can be particularly noticeable in artificial reality applications where the user's physical hands are visible (whether in passthrough or animated form). The system and technology described in this article provide a comprehensive solution that can convincingly display objects associated with a user's hand (such as a pen or cup held in the user's hand) or facilitate advanced interaction with the user's hand (such as a virtual orange (putty) that morphs based on the tracked hand position).
[0045] In certain embodiments, this disclosure relates to a system that provides a dedicated application programming interface (API) to applications running on a framework provided by an artificial reality system. For example... Figure 1A As shown, the artificial reality system includes a desktop computing system and a mobile computing system, such as a head-mounted device system, which also provides the user with the display of virtual objects. Through an API, the application can specify that the virtual object will be anchored to a specific point on the user's hand (e.g., the first joint of the user's left ring finger), or the application can specify other predetermined relationships with a specified physical object. Through the API, the application can also interact with callback functions to further customize the interaction between the virtual object and the user's hand. When rendering frames on the desktop computing system, the artificial reality system receives hand tracking information and head tracking information, associates the virtual object with the specified anchor point based on the hand tracking information, and renders an image of the virtual object as seen from the viewpoint defined by the head tracking information. In a particular embodiment, the image of the virtual object may be referred to as a surface object primitive or surface of the virtual environment.
[0046] The rendered image can be sent to a head-mounted device capable of displaying subframes at a higher rate than the rate at which the virtual object's image is rendered (e.g., 100 fps compared to 30 fps). At each subframe, the image of the hand-locked virtual object can be warped based on updates to the user's viewpoint (e.g., based on head rotation and head translation) and updates from the hand-tracking subsystem (e.g., hand and finger rotation and hand and finger translation) until the next frame, which includes a new image of the virtual object, is rendered.
[0047] The deformation and translation can be calculated in various ways, for example, based on available power and hardware resources, the complexity of the virtual objects and the virtual environment, etc. As an example, in a particular embodiment, for each subframe, the head-mounted device can first calculate a 3D transformation of the image of the virtual object in the virtual environment based on the difference between the current hand pose and the hand pose used to render the image. Then, the head-mounted device can transform the image of the virtual object in the virtual environment based on this 3D transformation. Subsequently, the head-mounted device can (e.g., based on head pose tracking information) project the transformed surface to the user's current viewpoint. This embodiment can specifically consider other virtual objects, including other surfaces that may occlude the image of the hand-locked virtual object, or other surfaces occluded by the image of the hand-locked virtual object. Specifically, the third step can allow direct calculation of the relative visibility of the surface and the virtual object.
[0048] As another example, in a particular embodiment, a two-dimensional transformation (e.g., relative to the user's viewpoint) for the image of the virtual object can be calculated to simplify the process described above. For each subframe, the head-mounted device can first calculate the two-dimensional transformation for the image of the virtual object based on the difference between the current head pose and the head pose used to render the image of the virtual object, and the difference between the current hand pose and the hand pose used to render the image of the virtual object. The head-mounted device can then transform the image of the virtual object using a two-dimensional transformation function. In both example scenarios, the appearance (e.g., color, text, etc.) of the hand-locked virtual object shown in the image of the hand-locked virtual object may not be updated until another frame is rendered. Instead, the image of the virtual object is manipulated as a virtual object in the virtual environment. At an appropriate time when another frame is rendered, the pose and appearance of the virtual object can be updated based on the newly reported hand pose and the newly reported head pose.
[0049] In one embodiment, the hand-locked virtual object is first rendered by a desktop computing system, and the deformation operation is performed by a head-mounted device. In other embodiments, the head-mounted device is capable of rendering the hand-locked virtual object and deforming it. In this embodiment, the desktop computing system may continue to render a world-locked object at a first frame rate (e.g., 30 fps), and the head-mounted device may deform the world-locked object. The hand-locked virtual object may be rendered directly at the display rate (e.g., 100 fps), and the head-mounted device may deform it. In this case, the image of the hand-locked virtual object is rendered at the display rate, so the hand-locked virtual object can be transformed to adapt to changes in the user's hand posture. However, the hand-locked virtual object can still be provided to the deformation engine of the head-mounted device so that the image of the hand-locked virtual object can be corrected for, for example, lens distortion, display inhomogeneity, chromatic aberration, etc.
[0050] In certain embodiments, artificial reality systems supporting hand-locked virtual object rendering can also facilitate the occlusion of virtual objects by the user's hand. For example, this can be used to show only portions of a hand-locked virtual object that are visible if the hand-locked virtual object were actually held by the user (e.g., showing only the portion of a hand tool not obscured by the user's fingers, or only showing the front of a ring). The hand occlusion subsystem can receive hand tracking data from the hand tracking subsystem. The hand occlusion subsystem can update a 3D hand model representing the user's current hand pose in the virtual environment at a display rate (e.g., 100 fps). The hand occlusion subsystem can use the hand model and the user's current viewpoint to generate a mask for the occluded portions of the user's hand with visual spatial depth. During each subframe prepared for display, the display system can check against the hand mask and compare the depth of surfaces and virtual objects near the user's hand with the depth associated with the hand mask. If the hand mask is closer to the user's viewpoint, the portion of the hand-locked virtual object obscured by the hand mask is not displayed during the sub-virtual object frame. This could involve providing instructions to the display components to turn off the pixels corresponding to the hand mask, thereby preventing the virtual object from being displayed and making the physical world (including the user's hand) visible. Allowing the user to see or interact with their own hand in front of a virtual object can help the user feel comfortable in the augmented reality environment. For example, embodiments of this disclosure can help reduce motion sickness or simulation-related illness in users.
[0051] Existing AR technologies fail to effectively address these issues. In a common approach to presenting AR experiences, the user views the environment through a standard screen (e.g., a smartphone screen). The virtual environment is overlaid on an image of the environment captured by a camera. This requires significant computational resources, as the image of the environment must be captured and processed rapidly, quickly draining the mobile device's battery. Furthermore, this type of experience is not particularly immersive for the user, as they can only view the environment through a small screen. In related approaches, many current systems struggle to accurately detect the user's hands using existing camera technology, preventing the user's hands from manipulating virtual objects within the virtual environment. Advanced technologies (such as those disclosed herein) are also lacking to accurately simulate the user's hands in the virtual environment and the effects of those hands within the virtual environment, and to render the virtual environment based on these effects. As an additional example, current AR system rendering schemes cannot render most virtual environments at sufficiently high frame rates and quality levels, making it difficult for users to comfortably experience them for extended periods. As discussed herein, high frame rates can be particularly advantageous for mixed or augmented reality experiences because the juxtaposition of virtual objects with the user's real environment allows the user to quickly identify any technical flaws in the rendering. The solution described in this article solves all and more of these technical problems.
[0052] In a particular embodiment, to support occlusion of virtual objects by a user's hand, the computing device can determine the distance from the user's hand to the user's viewpoint in the real environment. The computing device can correlate this distance with the distance from a virtual object representation of the hand to the user's viewpoint in the virtual environment. The computing device can also create a grid for storing height information of different regions of the hand (e.g., corresponding to the height of the region above or below a reference plane), or a grid for storing depth information of different regions of the hand (e.g., corresponding to the distance between that region of the hand and the user's viewpoint). Based on one or more images and a 3D grid, the computing device determines the height of various points on the hand (e.g., the height of a specific point on the hand relative to an average or median height, or relative to a reference point on the hand). The determined height indicates the position of each point on the hand relative to the rest of the hand. Combined with the determined distance, this height can be used to determine the precise position of various parts of the hand relative to the user. Alternatively, as described above, the user's hand can be stored as a 3D geometric model, and relevant distances can be determined based on this 3D representation.
[0053] When rendering and presenting a virtual environment to a user, the portion of the hand closer to the user than any virtual object should be visible, while the portion of the hand behind at least one virtual object should be occluded. The user's actual physical hand can be made visible to the user via the HMD of the AR system. In certain embodiments, light-emitting components (e.g., light-emitting diodes, LEDs) in the HMD that display virtual objects in the virtual environment can be selectively disabled to allow light from the real environment to pass through the HMD and reach the user's eyes. That is, the computing device creates a truncated region in the rendered image where the hand will appear. Thus, for example, by instructing light emitters not to illuminate the position on the display corresponding to the thumb, the user's actual physical thumb can be made visible via the HMD. Since these light emitters are turned off, any portion of the virtual object behind the thumb will not be displayed.
[0054] To correctly render real-world object occlusion, the light emitter corresponding to a portion of an object (e.g., a portion of a user's hand) should be turned off or instructed not to emit light when that portion of the hand is closer to the user in the real environment than any virtual object in the virtual environment, and should be turned on when a virtual object exists between the user and that portion of the hand. The light emitter displaying the virtual object will show portions of the object (fingers on the hand) behind the virtual object: these portions are farther from the user in the real environment than the virtual object is from the user in the virtual environment. Comparing the distances to real-world objects and to virtual objects is possible because, for example, these distances are determined based on hand tracking information. The distance to the virtual object is known by the application or scene running on the AR system, and this distance is available to both the AR system and the HMD.
[0055] Considering the virtual object representation of the user's hand and the known height of various positions of that hand (or, alternatively, the depth of a portion of the hand from the user's viewpoint), the HMD can make the portions of the hand visible to the user as follows: A frame depicting the virtual environment can be rendered by the main rendering unit (such as a desktop computing device) based on the user's current pose (e.g., position and orientation) at a first frame rate (e.g., 30fps). As part of rendering this frame, two items are generated: (1) a two-dimensional opaque texture for the hand based on a 3D mesh, and (2) a height map (or depth map, if applicable) of the hand. This two-dimensional opaque texture is stored as the texture of the image of the virtual object. These operations can be performed by the HMD or by a separate computing device communicating with the HMD hardware. By using a specially designated color for this two-dimensional opaque texture, a light emitter can be instructed not to emit light. In AR, it may be desirable to have a default background that allows light from the real environment to penetrate. In this way, the immersive experience of virtual objects presented in the real environment can be greatly enhanced. Therefore, the background can be associated with a color that is converted to indicate that the light emitter is not emitting light.
[0056] The HMD can then render multiple subframes at a second rate (e.g., 100fps) based on previously generated frames (e.g., frames generated at 30fps based on images including virtual objects). For each subframe, the AR system can perform a preliminary visibility test for each individual pixel or block of pixels entering the virtual environment by projecting one or more rays from the user's current viewpoint into the virtual environment based on the user's pose, which may differ from the pose used to generate the main 30fps frame. In a particular embodiment, the virtual environment may have a limited number of virtual objects. For example, the AR system may limit the number of discrete objects in the virtual environment (including objects created to represent the user's hand). Virtual objects can be represented by surfaces with corresponding heightmap information (or depthmap information if depth is used). If the projected ray intersects the surface corresponding to the user's hand, the ray samples the associated texture to render the subframe. Since the texture is opaque and black (indicating that the light emitter is not emitting light), the actual physical hand can be seen through the transparent, unlit areas of the HMD. More information regarding hand occlusion achieved by modeling the user's hand and performing depth testing on the model and other virtual objects can be found in U.S. Patent Application No. 16 / 805,484, filed February 28, 2020, which has been incorporated herein by reference.
[0057] In certain embodiments, the primary rendering device (e.g., a computing device other than the HMD) may be equipped with greater processing power and power supply than the HMD. Therefore, the primary rendering device can perform certain parts of the techniques described herein. For example, the primary rendering device can perform hand tracking calculations, generating a 3D mesh, a 2D opaque texture, and a height map (or depth map, if applicable) for the hand. For example, the primary rendering device can receive images created by the HMD's camera and can use dedicated computing hardware designed to be more efficient or powerful for a particular task to perform the necessary processing. However, if the HMD has sufficient processing power and power, it can (e.g., using its own onboard computing components) perform one or more of these steps to reduce latency. In certain embodiments, all of the steps described herein are performed on the HMD.
[0058] Figure 1A An example artificial reality system 100 is illustrated. In a particular embodiment, the artificial reality system 100 may include a head-mounted device system 110 (which may be included in an HMD), a desktop computing system 120 (which may include a body-worn computing system, a laptop computer system, a desktop computer system, etc.), a cloud computing system 132 in a cloud computing environment 130, etc. In a particular embodiment, the head-mounted device system 110 may include various subsystems that perform the functions described herein. The various subsystems may include dedicated hardware and integrated circuits to facilitate the functionality of these subsystems. In a particular embodiment, these subsystems may be functional subsystems running on one or more processors or integrated circuits of the head-mounted device system 110. Therefore, given the various systems of the head-mounted device system 110 described in this application, it should be understood that these systems are not limited to a single hardware component of the head-mounted device system 110.
[0059] In a particular embodiment, the head-mounted device system 110 may include a tracking system 140 that tracks information about the wearer's physical environment. For example, the tracking system 140 may include a hand tracking subsystem that tracks the wearer's hands. The hand tracking subsystem may employ optical sensors, light projection sensors, or other types of sensors to perform hand tracking. The hand tracking subsystem may also utilize hand tracking algorithms and models to generate hand tracking information. The hand tracking information may include the hand's pose (e.g., position and orientation) and various anchor points or portions of the hand. In a particular embodiment, the tracking system 140 may include a head tracking subsystem that tracks the pose (e.g., position and orientation) of the wearer's head. The head tracking subsystem may determine the pose using sensors (such as optical sensors) mounted on or included in the head-mounted device system 110, or it may calculate the pose using external sensors throughout the physical environment of the head-mounted device system 110. In a particular embodiment, the head-mounted device system 110 may further include an occlusion system 145, which determines, based on information provided by the tracking system 140, whether one or more virtual objects in the virtual environment of the artificial reality experience presented by the head-mounted device system 110 should be occluded by the user or by physical objects manipulated by the user.
[0060] In a particular embodiment, the head-mounted device system 110 may further include a display engine 112 connected via a data bus 114 to two eye display systems 116A and 116B (also referred to as the first eye display system 116A and the second eye display system 116B). The head-mounted device system 110 may be a system including a head-mounted display (HMD) that can be fixed to a user's head to provide the user with artificial or augmented reality. The head-mounted device system 110 may be designed to be lightweight and highly portable. Therefore, the available power in the head-mounted device system's power supply (e.g., a battery) may be limited. The display engine 112 can provide display data to the eye display systems 116A and 116B via the data bus 114 at a relatively high data rate (e.g., adapted to support refresh rates of 200 Hz or higher). The display engine 112 may include one or more controller blocks, texel memories, transform blocks, pixel blocks, etc. The texels stored in the texel memory can be accessed by the pixel blocks and can be provided to the eye display systems 116A and 116B for display. Further information regarding the described display engine 112 can be found in U.S. Patent Application No. 16 / 657,820, filed October 1, 2019; U.S. Patent Application No. 16 / 586,590, filed September 27, 2019; and U.S. Patent Application No. 16 / 586,598, filed September 27, 2019, all of which are incorporated herein by reference.
[0061] In certain embodiments, the stage computing system 120 may be worn on a user's body. In certain embodiments, the stage computing system 120 may be a computing system not worn on a user's body (e.g., a laptop computing system, a desktop computing system, a mobile computing system). The stage computing system 120 may include one or more GPUs, one or more intelligent video decoders, memory, a processor, and other modules. The stage computing system 120 may have more computing resources than the display engine 112, but in some embodiments, the power supply (e.g., a battery) of the stage computing system may still be limited. Although not shown, in certain embodiments, the stage computing system may have subsystems for supporting tracking of the wearer of the head-mounted device system 110 and physical objects in the environment of the head-mounted device system 110, or for supporting occlusion computing. The stage computing system 120 may be coupled to the head-mounted device system 110 via a wireless connection 144. The cloud computing system 132 may include a high-performance computer (e.g., a server) and may communicate with the stage computing system 120 via a wired or wireless connection 142. In some embodiments, the cloud computing system 132 can also communicate wirelessly with the head-mounted device system 110 (not shown). The desktop computing system 120 can generate data for rendering at a standard data rate (e.g., adapted to support refresh rates of 30 Hz or higher). The display engine 112 can upsample the data received from the desktop computing system 120 to generate frames that will be displayed by the eye-mounted display systems 116A and 116B at a higher frame rate (e.g., 200 Hz or higher).
[0062] Figure 1B An example eye-mounted display system (e.g., 116A or 116B) is shown for a head-mounted device system 110. In a particular embodiment, the eye-mounted display system 116A may include a driver 154, a pupil display 156, etc. The display engine 112 may provide display data to the pupil display 156, the data bus 114, and the driver 154 at a high data rate (e.g., adapted to support a refresh rate of 200 Hz or higher).
[0063] Figure 2 A system diagram of display engine 112 is shown. In a particular embodiment, display engine 112 may include control block 210, transformation blocks 220A and 220B, pixel blocks 230A and 230B, display blocks 240A and 240B, etc. One or more components of display engine 112 may be configured to communicate via a high-speed bus, shared memory, or any other suitable method. Figure 2As shown, the control block (also called the controller block) 210 of the display engine 112 can be configured to communicate with the transformation blocks 220A and 220B, the pixel blocks 230A and 230B, and the display blocks 240A and 240B. As explained in more detail herein, this communication may include data as well as control signals, interrupts, and other instructions.
[0064] In a particular embodiment, control block 210 can be controlled from a desktop computing system (e.g., Figure 1A The control block 210 receives input and initializes the pipeline in the display engine 112 to ultimately complete rendering for display. In a particular embodiment, the control block 210 may receive data and control packets from the desktop computing system at a first data rate or frame rate. The data and control packets may include information such as one or more data structures, including texture data and position data, and additional rendering instructions. In a particular embodiment, the data structure may include two-dimensional rendering information. The data structure may be referred to herein as a “surface” and may represent an image of a virtual object to be rendered in the virtual environment. The control block 210 may distribute data to one or more other blocks of the display engine 112 as needed. The control block 210 may initiate pipeline processing for one or more frames to be displayed. In a particular embodiment, each of the eye display systems 116A and 116B may include its own control block 210. In a particular embodiment, one or both of the eye display systems 116A and 116B may share the control block 210.
[0065] In a particular embodiment, transform blocks 220A and 220B can determine initial visibility information of the surfaces to be displayed in an artificial reality scene. Typically, transform blocks 220A and 220B can project light rays with an origin based on the pixel positions in the image to be displayed and generate filtering commands (e.g., filtering based on bilinear or other types of interpolation techniques) to be sent to pixel blocks 230A and 230B. Transform blocks 220A and 220B can project light rays into the user's real or virtual environment based on the user's current viewpoint. Head-mounted device sensors (such as one or more cameras (e.g., monochrome, full-color, depth sensing), inertial measurement units, eye trackers) and / or any suitable tracking / localization algorithm (such as simultaneous localization and mapping (SLAM) embedded in the environment and / or virtual scene in which the surfaces are located) can be used to determine the user's viewpoint, and the results can be sent to pixel blocks 230A and 230B.
[0066] Typically, according to a specific embodiment, transform blocks 220A and 220B may each include a four-stage pipeline. The stages of transform blocks 220A or 220B can be performed as follows: A ray projector can emit a ray beam corresponding to an array of one or more aligned pixels, referred to as a tile (e.g., each tile may include 16×16 aligned pixels). The ray beam can be deformed according to one or more distortion meshes before entering the artificial reality scene. The distortion meshes can be configured to correct for geometric distortion effects from at least the eye display systems 116A and 116B of the head-mounted device system 110. In a specific embodiment, transform blocks 220A and 220B can determine whether each ray beam intersects a surface in the scene by comparing the bounding box of each tile with the bounding box of a surface. If a ray beam does not intersect an object, it can be discarded. Tile-surface intersections are detected, and the corresponding tile-surface pairs are transmitted to pixel blocks 230A and 230B.
[0067] Typically, according to a particular embodiment, pixel blocks 230A and 230B can determine color values based on tile-surface pairs to generate pixel color values. The color value of each pixel can be sampled from the texture data of the surface received and stored by control block 210. Pixel blocks 230A and 230B can receive tile-surface pairs from transform blocks 220A and 220B and can schedule bilinear filtering. For each tile-surface pair, pixel blocks 230A and 230B can sample the color information of the pixel corresponding to that tile using the color value corresponding to the location where the projected tile intersects with the surface. In a particular embodiment, pixel blocks 230A and 230B can process the red, green, and blue components of each pixel separately. In a particular embodiment, as described herein, the pixel blocks can employ one or more processing shortcuts based on colors associated with the surface (e.g., color and opacity). In a particular embodiment, the pixel block 230A of the display engine 112 of the first eye display system 116A can travel independently and can travel in parallel with the pixel block 230B of the display engine 112 of the second eye display system 116B. The pixel block can then output its color measurement to the display block.
[0068] Typically, display blocks 240A and 240B can receive pixel color values from pixel blocks 230A and 230B, convert the data format to be more suitable for the display (e.g., if the display requires a specific data format such as that in a scanline display), apply one or more brightness corrections to the pixel color values, and prepare the pixel color values for output to the display. Display blocks 240A and 240B can convert the tile-sequence pixel color values generated by pixel blocks 230A and 230B into scanline data or line-sequence data that the physical display may require. Brightness correction may include any necessary brightness correction, gamma mapping, and dithering. Display blocks 240A and 240B can output the corrected pixel color values directly to the physical display (e.g., directly via driver 154). Figure 1B The display system 116A and 116B or the head-mounted device system 110 may include additional hardware or software to further customize back-end color processing, support wider display interfaces, or optimize display speed or fidelity.
[0069] In a particular embodiment, controller block 210 may include microcontroller 212, texel memory 214, memory controller 216, data bus 217 for I / O communication, data bus 218 for input stream data (also referred to as input stream) 205, etc. Memory controller 216 and microcontroller 212 can be coupled via data bus 217 for I / O communication with other modules of the system. Microcontroller 212 can receive control packets such as position data and surface information via data bus 217. Input stream data 205 can be input from a desktop computing system to controller block 210 after being configured by microcontroller 212. Input stream data 205 can be converted to the required texel format by memory controller 216 and stored in texel memory 214 by memory controller 216. In a particular embodiment, texel memory 214 may be static random-access memory (SRAM).
[0070] In a particular embodiment, the desktop computing system 120 and other subsystems of the head-mounted device system 110 can send input stream data 205 to a memory controller 216, which can convert the input stream data into texels with a desired format and store the texels in a texel memory 214 using a swizzle pattern. The texel memory organized in these swizzle patterns allows pixel blocks 230A and 230B to retrieve texels (e.g., in 4x4 texel blocks) using a single read operation. These texels need to be used to determine at least one color component (e.g., red, green, and / or blue) for each pixel in all pixels associated with a tile (e.g., a "tile" refers to an aligned pixel block, such as a 16x16 pixel block). As a result, the display engine 112 can avoid the excessive complexity typically required to read and assemble the texel array when it is not stored in the appropriate pattern, and thus reduce the computational resource requirements and power consumption of the display engine 112 and the head-mounted device system as a whole.
[0071] In a particular embodiment, pixel blocks 230A and 230B can generate pixel data for display based on texels retrieved from texel memory 212. Memory controller 216 can be coupled to pixel blocks 230A and 230B via two 256-bit data buses (also referred to as 256-bit buses) 204A and 204B, respectively. Pixel blocks 230A and 230B can receive tile / surface pairs 202A and 202B from their respective transform blocks 220A and 220B, and can identify texels that need to be used to determine at least one color component of all pixels associated with the tile. Pixel blocks 230A and 230B can retrieve the identified texels (e.g., a 4x4 texel array) in parallel from texel memory 214 via memory controller 216 and 256-bit data buses 204A and 204B (also referred to as 256-bit buses 204A and 204B). For example, a 4x4 texel array required to determine at least one color component of all pixels associated with a tile can be stored in a memory block and retrieved using a single memory read operation. Pixel blocks 230A and 230B can use multiple sampling filter blocks (e.g., one sampling filter block per color component) to perform interpolation on different texel groups in parallel to determine the corresponding color component of the corresponding pixel. Pixel values 203A and 203B for each eye can be sent to display blocks 240A and 240B for further processing, respectively, before the eye display systems 116A and 116B display pixel values 203A and 203B for each eye.
[0072] In certain embodiments, the artificial reality system 100, and particularly the head-mounted device system 110, can be used to render an augmented reality environment to a user. The augmented reality environment can include elements of a virtual reality environment (e.g., virtual reality objects) rendered for the user, such that virtual elements appear on or within the user's real environment. For example, the user can wear an HMD (e.g., head-mounted device system 110) incorporating features of the various techniques disclosed herein. The HMD can include a display that allows light from the user's environment to continue reaching the user's eyes normally. However, when the light-emitting components of the display (e.g., LEDs, organic light-emitting diodes (OLEDs), micro-LEDs, etc.) emit light at specific locations within the display, the color of the LEDs can be superimposed onto the user's environment. Thus, when the light emitter of the display emits light, virtual objects can appear in front of the user's real environment, and when the light-emitting components at a specific location are not emitting light, the user's environment can be made visible at that location via the display. This can be achieved without re-rendering the environment using a camera. This process can improve the user's immersion and comfort when using the artificial reality system 100, while reducing the computational power and battery consumption required to render the virtual environment.
[0073] In certain embodiments, selectively emitting displays can be used to facilitate user interaction with virtual objects in a virtual environment. For example, the position of a user's hand can be tracked. In previous systems, even hand-tracking systems, the only way to represent a user's hand in a virtual environment was to generate and render some kind of virtual representation of the hand. This can be unsuitable in many use cases. For example, in a workplace environment, it may be desirable for the user's hand itself to be visible to the user when interacting with an object. For example, when an artist holds a virtual paintbrush or other tool, they may want to see their own hand. The techniques disclosed herein for achieving such advanced displays are...
[0074] The principles behind the technology described in this article will now be explained. Figure 3 An example of a user's internal view of a head-mounted device is shown, illustrating an example of a user viewing their own hand through an artificial reality display system. Figure 3The diagram also illustrates how the difference between the pose of the virtual object displayed to the user and the user's current hand pose can degrade the user's experience interacting with the artificial reality system 100. The head-mounted device's internal view 300 shows a composite (e.g., mixed) reality display including a view of the virtual object 304 and a view of the user's physical hand 302. The virtual object 304 is anchored to a predetermined position (also referred to as an anchor position, anchor point, or point) 306 on the user's physical hand 302. In this example, the virtual object 304 is a loop configured to be anchored to the first joint of the user's hand. The anchor position 306 can be configured to be adjustable, for example, by the user or by the developers of the artificial reality experience, allowing the user to move the anchor point 306, or relationship, to the user's hand 302. Note that because the virtual object 304 is designed to appear wrapped around the user's hand 302 when viewed by the user, a portion of the virtual object 304 is obscured by the user's hand 302 and should not be displayed.
[0075] To determine the occluded portions of virtual object 304, in certain embodiments, the artificial reality system 100 can access or generate a virtual representation of a user's hand. View 310 illustrates a virtual representation 312 of the user's hand based on view 300. The artificial reality system 100 can use the virtual representation 312 of the hand to determine which portions of the virtual object 304 should be occluded. For example, in some embodiments, the artificial reality system 100 can perform ray casting or ray tracing operations from the user's viewpoint. The artificial reality system 100 can insert the virtual object 304 and the virtual representation 312 of the user's hand into the virtual environment based on the positions and specified relationships (e.g., relationships with rings, appropriate fingers, and appropriate points) specified by anchor position 306. The artificial reality system 100 can then project light from the user's viewpoint into the virtual environment and determine whether the first visible virtual object is virtual object 304 or the virtual representation 312 of the user's hand. Based on this result, the artificial reality system 100 can prepare a rendered image of the virtual object 304, selectively omitting the portions occluded by the user's hand. In some embodiments, ray casting or ray tracing operations can be enhanced by using a depth map (or height map) associated with the virtual object 304 and a virtual representation 312 of the user's hand.
[0076] As will be described in this article, there may be situations where changes in a user's hand posture may outpace the speed at which traditional techniques update and present an image of the user's hand to the user. This is considered in the context of... Figure 4 Examples of potential differences are shown. Figure 4 It shows including from Figure 3The view 300 shows the user's hand 302 and the ring virtual object 304. After the user's hand 302 has moved for a period of time, views 410 and 420 show the user different presentations. In this example, the artificial reality system 100 is showing a mixed reality view, meaning the user can perceive their own hand 302. Any difference between the position where the virtual object is shown and the position where the virtual object is believed to belong can have a particularly large impact on hand-locked rendering, because how a hand-locked virtual object should be presented can be immediately intuitive for the user. In view 410, the user's hand 302 has moved to a slightly rotated position. However, the pose and appearance of the virtual object 304 have not changed. Therefore, the user will immediately notice that the virtual object 304 is not rendered correctly, thus destroying the immersive nature of the hand-locked rendering of the virtual object 304. The situation in view 410 could occur if the virtual object 304 is not re-rendered at a sufficiently high rate.
[0077] As a solution to this problem, and as disclosed herein, the artificial reality system 100 can provide subframe updates of the rendered frames, rather than simply attempting to increase the rate at which an image of a virtual object is generated from a virtual model and presented to the user. Subframe rendering may involve adjusting or otherwise modifying a rendered image of the virtual object (e.g., a “first image”) before rendering (e.g., placing the first image in a virtual environment or directly modifying the first image) and providing the adjusted image (e.g., a “second image”) to the user. View 420 shows a view of the user’s hand 302 and the virtual object 304 after the user’s hand pose has changed and the virtual object 304 has been updated to reflect the detected change. When viewed through the artificial reality system 100, the virtual object 304 has been rotated or otherwise moved to maintain its anchor point 306 on the user’s hand 302.
[0078] Figures 5A to 5BAn example method 500 for hand-locked rendering of virtual objects using subframe rendering in artificial reality is illustrated. In a particular embodiment, steps 510 through 540 may involve generating a first image to be displayed to the wearer of head-mounted device system 110 based on the user's hand and the position of the virtual object in the virtual environment. In a particular embodiment, one or more of steps 510 through 540 may be performed by desktop computing system 120, by another suitable computing system having more computing resources than head-mounted device system 110 and communicatively coupled to head-mounted device system 110, or by head-mounted device system 110 itself. In a particular embodiment, the workload performed in each step may be distributed among suitable computing systems by the workload controller of artificial reality system 100, for example, based on one or more metrics of available computing resources, such as battery power, processor (e.g., graphics processing unit (GPU), central processing unit (CPU), other application-specific integrated circuit (ASIC)) utilization, memory availability, hot water level, etc.
[0079] The method may begin at step 510, where at least one camera of the head-mounted device system 110 (e.g., a head-mounted display) captures an image of the environment of the head-mounted device system 110. In a particular embodiment, the image may be a standard panchromatic image, a monochrome image, a depth-sensing image, or any other suitable type of image. The image may include objects that the artificial reality system 100 determines should be used to occlude virtual objects in a virtual or mixed reality view. Furthermore, the image may include the user's hand, but may also include various suitable objects as described above.
[0080] The method can continue at step 515, where the computing system can determine the user's viewpoint in the environment of the head-mounted device system 110. In a particular embodiment, the viewpoint can be determined entirely based on captured environmental images using appropriate positioning and / or mapping techniques. In a particular embodiment, the viewpoint can be determined using data retrieved from other sensors of the head-mounted device system 110, such as head pose tracking sensors integrated into the head-mounted device system 110 that provide head tracking information (e.g., from tracking system 140). Based on this head tracking information, the artificial reality system 100 can determine the user's viewpoint and transpose the user's viewpoint to a virtual environment containing virtual objects. For example, the artificial reality system 100 can treat the user's viewpoint as a camera viewing the virtual environment.
[0081] In step 520, the artificial reality system 100 may determine one or more hand poses (e.g., position and orientation) of the user's hand based on hand tracking information. The artificial reality system 100 may communicate with, or may include, a hand tracking component that detects the user's hand pose (e.g., from tracking system 140). In a particular embodiment, the hand tracking component may include a sensor attached to an object held by the user to interact with the virtual environment. In a particular embodiment, the hand tracking component may include a hand tracking model that receives captured images (including portions of the captured images that include the user's hand) as input and uses hand tracking algorithms and machine learning models to generate hand poses.
[0082] The computing system can execute one of several algorithms to identify the presence of a user's hand in a captured image. After confirming that the hand is indeed present in the image, the computing system can perform hand tracking analysis to identify the presence and location of several discrete locations on the user's hand in the captured image. In a particular embodiment, a depth-tracking camera can be used to facilitate hand tracking. In a particular embodiment, hand tracking can be performed without standard depth tracking and model-based tracking using deep learning. A deep neural network can be trained and used to predict the location of a person's hand and landmarks (e.g., the joints and fingertips of the hand). These landmarks can be used to reconstruct high-DOF poses (e.g., 26-DOF poses) of the hand and fingers. The prediction model for the pose of the person's hand can be used to reduce latency between pose detection and rendering or display. Additionally, detected hand poses (e.g., detected using a hand tracking system or using a prediction model) can be stored for further use in hand tracking analysis. As an example, the detected hand pose can be used as a base pose and inferred from it to generate potential predicted poses in subsequent frames or at later times. Therefore, the speed of hand pose prediction and detection can be improved. This pose provides the position of the hand relative to the user's viewpoint in the environment. In some embodiments, the hand tracking component may include a variety of hand tracking systems and technologies to improve the adaptability and accuracy of the hand tracking system.
[0083] In step 525, the artificial reality system 100 may identify one or more anchor points on the user's hand after determining the viewpoint and hand pose. In a particular embodiment, a hand tracking component may provide these anchor points as part of hand tracking information. As an example, hand tracking information may specify discrete locations on the user's hand, such as fingertips, knuckles, palm, or back of the hand. In a particular embodiment, for example, the artificial reality system 100 may determine the anchor points based on the hand tracking information, for example, based on the operating mode of the artificial reality system 100 or based on the type of virtual object that the artificial reality system 100 will generate.
[0084] In step 530, the artificial reality system 100 can generate a virtual object based on hand pose, head pose, viewpoint, and a predetermined spatial relationship between the user's hand and an anchor point. To prepare the virtual object 304, the artificial reality system 100 can access a model or other representation of the virtual object 304. As part of the virtual object's specification, the model of the virtual object may include information about how the virtual object will be positioned relative to the user's hand. For example, the virtual object's designer may have specified that the virtual object will be displayed at or within a specified distance from an anchor point or other location relative to the user's hand. The specification may also include a specified orientation of the object relative to the anchor point. Figure 3 In the example shown, the model of the loop virtual object 304 can specify that the virtual object will be anchored to point 306 at the first joint of the first finger of the user's hand, and the loop virtual object 304 should be oriented such that the loop virtual object 304 surrounds the user's finger (e.g., such that the plane including the virtual object 304 intersects the anchor point 306 and is perpendicular to the direction indicated by the user's finger). The artificial reality system 100 can use viewpoint and hand tracking information to automatically transform the model to a position in the virtual avatar corresponding to the pose specified by the model. Therefore, the virtual object is automatically placed in the correct position relative to the anchor point by the artificial reality system 100, and the designer of the virtual object does not need to know the world coordinates or camera coordinates of the virtual object.
[0085] In step 535, the artificial reality system 100 (e.g., occlusion system 145) can determine whether any part of a virtual object is occluded by the user's hand. As an example, the head-mounted device system 110 can generate a virtual object representation of a hand in the virtual environment. The position and orientation of the virtual object representation of the hand can be based on updated hand tracking data and a second viewpoint. The head-mounted device system 110 can project one or more rays having an origin and trajectory based on the second viewpoint into the virtual environment. For example, the head-mounted device system 110 can simulate projected rays for each pixel of a display component of the head-mounted device system 110, or it can simulate pixel-projected rays for each of several pixel groups. If a virtual object is occluded by the user's hand when viewed from a particular pixel, the head-mounted device system 110 can determine the point of intersection of the ray with the virtual object representation of the hand before it intersects with another virtual object.
[0086] In step 540, the artificial reality system 100 can render a first image of the virtual object as seen from a first viewpoint, using the model of the virtual object at an appropriate location and orientation in the virtual environment. As described herein, the first image may include a two-dimensional representation of the virtual object as seen from the first viewpoint. The two-dimensional representation may be assigned a location in the virtual environment (e.g., based on the first viewpoint) and may optionally be assigned depth (e.g., to facilitate object occlusion, as described herein). The first image may also be referred to herein as a surface (generated by the components of the artificial reality system 100) to improve computational efficiency for updating the pose of the virtual object at subsequent higher frame rates.
[0087] The first image can be rendered at a first frame rate (e.g., 30 frames per second) by a desktop computing system of the artificial reality system 100. As discussed herein, the process of generating virtual objects and rendering the first image can be computationally expensive and relatively power-intensive. Therefore, in certain embodiments, the desktop computing system may be a more powerful computing system than the head-mounted device system 110 that actually displays the virtual environment (including virtual objects) to the user. The desktop computing system can send the first image to the head-mounted device system 110, which is configured to display a virtual reality environment to the user, via wired or wireless communication. However, with the introduction of transmitting the first image to the user, and especially if the first image is only reproduced at the first frame rate (e.g., along with the pose of the virtual objects), a significant delay may occur between the generation of the virtual objects and the display of the first image. Worse still, the artificial reality system 100 may not be able to quickly adapt to minute changes in viewpoint or hand pose that can be detected by a hand or head tracking system. The difference between the user's hand position when generating virtual objects and the user's hand position when displaying the first image can significantly reduce the quality of the user's experience using the artificial reality system 100. Besides reducing quality, it may also cause physical discomfort or lead to abnormal virtual experiences to some extent.
[0088] To correct this potential problem, the head-mounted device system 110 can be configured to update the pose of a first image before the first image is re-rendered, and to render one or more second images based on the first image (e.g., subsequent first images of virtual objects can be rendered using updated viewpoint and hand tracking information). In a particular embodiment, systems for determining the user's viewpoint and for performing head and hand tracking can be integrated into the head-mounted device system 110. Furthermore, these head and hand tracking systems can report updated hand tracking information at a rate much faster than the frame rate at which the first image is generated (e.g., faster than 30 Hz).
[0089] Returning to method 500, this method can proceed to... Figure 5BIn step 550, as shown, the head-mounted device system 110 determines a second viewpoint for the user in both the real and virtual environments. Steps 550 through 595 may generate and provide a second image of the virtual object based on adjustments made to a first image of the virtual object (these adjustments are themselves based on updated hand pose and viewpoint information). The second image may be the result of the head-mounted device system 110 using new tracking data generated after the first image was rendered, updating the data as needed, and actually determining the image to be displayed to the user by the display components (e.g., an array of light-emitting components) of the head-mounted device system 110. The second image (also called a subframe) may be prepared and displayed at a rate much higher than that of preparing the first image (e.g., frames) (e.g., 200 frames per second or higher). Due to the higher rate of preparing and displaying the second frame, it may be necessary to limit inter-system communication. Therefore, in a particular embodiment, the computations required for steps 550 to 595 can be performed by the head-mounted device system 110 itself, rather than by a desktop computing system 120 or other computing system communicatively coupled to the head-mounted device system 110. In a particular embodiment, the head-mounted device system 110 can be configured to perform multiple operations for each received first image: receiving updated tracking information (e.g., steps 550 to 555), calculating and applying various adjustments (e.g., steps 560 to 570), rendering a second image (e.g., step 580), and providing the rendered second image for display (e.g., step 585). This is because the operations of receiving updated tracking information, calculating any differences, calculating any adjustments, and applying these adjustments can all be performed quickly as discrete tasks.
[0090] In step 550, the head-mounted device system 110 can determine a second (updated) viewpoint of the user in both the real environment and the virtual environment surrounding the user. This second viewpoint can be used to account for minute movements of the user (e.g., head movements, eye movements, etc.). The second viewpoint can be determined in the same manner as described above and can be received from the tracking system 140.
[0091] In step 555, the head-mounted device system 110 can determine a second hand posture of the user's hand. The head-mounted device system 110 can receive updated tracking information and determine the second hand posture based on the updated tracking information as described above (the updated tracking information may be received from the tracking system 140).
[0092] In step 560, the head-mounted device system 110 can determine the difference between the second hand posture and the first hand posture, and optionally, determine the difference between the second viewpoint and the first viewpoint.
[0093] In step 565, the head-mounted device system 110 may determine one or more adjustments to be performed on the first image to take into account differences between hand pose and viewpoint. As further discussed herein, various techniques may be used to determine and perform these adjustments based on the fact that the first image is represented as a two-dimensional object placed in a three-dimensional virtual environment.
[0094] In a particular embodiment, adjustments for rendering the second image can be calculated by directly treating the first image as an object in a 3D virtual avatar. For each subframe to be generated, the head-mounted device system 110 can first calculate a 3D transformation of the image for the virtual object in the virtual environment based on the difference between the current hand pose and the hand pose used to render the first image. For example, the head-mounted device system 110 can identify a new location of the anchor point or a predefined spatial relationship between the user's hand and the virtual object. The head-mounted device system 110 can calculate the difference between the location used to generate the first image and the current location. This difference can include differences along the x-axis, y-axis, or z-axis. Furthermore, the head-mounted device system 110 can determine differences between the orientation of the user's hand, such as roll rotation, pitch rotation, or yaw rotation. The head-mounted device system 110 can calculate corresponding changes in the pose of the first image as 3D transformations based on these differences. For example, the head-mounted device system 110 can translate or rotate the first image in the virtual environment. The head-mounted device system 110 can then (e.g., based on head pose tracking information) project the transformed first image onto the updated user viewpoint and render a second image for display based on the adjusted first image and viewpoint.
[0095] Figure 6 An example is shown of calculating adjustments for rendering a second image by moving a first image in a 3D virtual avatar. View 600 shows a conceptual diagram of an environment including a representation of a user's hand 602 and an image of a virtual object 604 shown as being held by the user's hand 602. View 610 shows the pose of the user's hand 602 shortly after view 600 is shown to the user. Next, view 610 shows the conceptual environment after the head-mounted device has determined updated hand pose information. View 620 shows the translation and rotation of the image of the virtual object 604 from position 606 shown in view 610 to position 608 based on the updated hand pose information. Finally, view 630 shows a composite view of the representation of the user's hand 602 at position 608 and the image of the virtual object 604. For simplicity, view 630 is shown from the same viewpoint, although this viewpoint can change based on updated head pose information. View 630 can be used to render a second image for display based on the adjusted image of the virtual object 604.
[0096] In a particular embodiment, adjustments used to render the second image can be calculated as a two-dimensional transformation of the first image (e.g., relative to the user's viewpoint). For each subframe to be generated, the head-mounted device system 110 can use visual warping or two-dimensional transformations to simulate the effect of moving the first image in a three-dimensional virtual environment, based on the difference between the current hand pose and the hand pose used to render the first image. For example, the head-mounted device system 110 can simply scale and translate the first image (instead of moving the first image in the virtual environment) while rendering the second image in a manner that approximates the effect of moving the first image in the virtual environment. Adjustments applied to the first image before rendering the second image can include, by way of example only and not limitation, scaling the first image, translating the first image along the x-axis or y-axis of the frame, rotating the first image along either axis, or shearing the first image. The head-mounted device system 110 can render and display the second image after applying visual warping.
[0097] Figure 7 An example of adjustments that can be applied to a first image of a virtual object 704, which has been rendered based on a predetermined spatial relationship with the user's hand 702, is shown. As shown, the first image of the virtual object 704 has been rendered and is ready for display. The head-mounted device system 110 can determine and apply multiple adjustments as two-dimensional transformations of the first image of the virtual object 704 in response to updated hand pose information. Adjusted image 710 shows the first image 704 after it has been magnified so that it appears closer to the viewpoint, for example, in response to the head-mounted device system 110 determining, based on hand pose information, that the user's hand 702 has moved closer to the user. Adjusted image 712 shows the first image 704 after it has been rotated along an axis perpendicular to the viewpoint, for example, in response to the head-mounted device system 110 determining, based on hand pose information, that the user's hand 702 has been rotated in a similar manner. Adjusted images 714 and 716 show the first image 704 after being rotated along multiple axes while applying horizontal or vertical tilt, or horizontal or vertical shearing to simulate the appearance of the first image 704, for example, in response to the head-mounted device system 110 determining, based on hand posture information, that the user's hand 702 has been rotated in a similar manner. Adjusted images 718 and 720 show the first image 704 after it has been scaled in both vertical and horizontal directions so that the first image 704 appears to have been rotated along the x and y axes of the user's display, for example, in response to the head-mounted device system 110 determining, based on hand posture information, that the user's hand 702 has been rotated along an axis similar to the user's display. Although Figure 7Several adjustments that can be made to the first image 704 are shown, but this should be illustrative rather than exclusive. Furthermore, these and other adjustments can be combined into any suitable combination where appropriate.
[0098] In both example scenarios, the appearance of the virtual object shown in the first image may not be updated until another first image (e.g., another frame) is rendered. For example, the colors, text, and positions of the parts included in the first image may not be updated when preparing and rendering each second image. In contrast, the first image of the virtual object can be manipulated as a virtual object within a virtual environment. At the appropriate time for rendering the other frame, the pose and appearance of the virtual object can be updated based on the pose of the new report's hand and the pose of the new report's head.
[0099] In step 570, the head-mounted device system 110 may apply the determined adjustments to the first image. For example, if the adjustment involves a three-dimensional transformation of the first image, the head-mounted device system 110 may move the first image within a virtual environment. As another example, if the adjustment involves a two-dimensional transformation, the head-mounted device system 110 may apply the two-dimensional transformation to the first image. Applying the determined adjustments may include creating a copy of the first image to store the first image for determining and applying subsequent adjustments.
[0100] Optionally, in step 575, the head-mounted device system 110 can be configured to update the appearance of the first image based on context. For example, if such an update is detected as needed, the head-mounted device system 110 can update lighting conditions, shaders, and the brightness of the first image. As another example, if the head-mounted computing device has available computing resources and energy, it can be configured to update the appearance of virtual objects. Because the head-mounted computing device can be worn by a user, the artificial reality system 100 can be configured to distribute certain rendering tasks between, for example, a desktop computing system and the head-mounted device system 110 based on metrics used to track operational efficiency and available resources.
[0101] In step 580, the head-mounted device system 110 may render a second image based on the adjusted first image viewed from a second viewpoint. In step 585, the head-mounted device system 110 may provide the second image for display. Based on the second image, the head-mounted device system 110 may generate instructions for the light emitters of its display. These instructions may include color values, color brightness or intensity, and any variables that affect the display of the second image. These instructions may also include instructions for certain light emitters to remain off or to be turned off when needed. The intended effect is that the user can see multiple parts of the virtual environment (e.g., virtual objects in the virtual environment) blended together with multiple parts of the real environment (including the user's hands, which may be interacting with virtual objects in the virtual environment) from a viewpoint in the virtual environment.
[0102] In step 590, the head-mounted device system 110 can determine whether a new first image has been provided (e.g., whether a new frame has been provided by the desktop computing system 120). If not, the method can return to step 550 and repeat the process to render an additional second image (e.g., continue subframe rendering). If a new first image exists, the head-mounted device system 110 can load data from the new first image at step 595 before proceeding to step 550.
[0103] Where appropriate, certain embodiments may be repeated. Figures 5A to 5B One or more steps in the method. Although this disclosure will Figures 5A to 5B Specific steps in the method are described and shown as occurring in a specific order, but this disclosure contemplates... Figures 5A to 5B Any suitable steps occurring in any suitable order within the method. Furthermore, although this disclosure describes and illustrates example methods for providing hand-locked rendering of virtual objects in artificial reality using subframe rendering (including...) Figures 5A to 5B The specific steps in the method are not considered here, but this disclosure contemplates any suitable method for providing hand-locked rendering of virtual objects in artificial reality using subframe rendering, including any suitable steps, and where appropriate, the method may include... Figures 5A to 5B All steps, some steps, or may be excluded from the method Figures 5A to 5B Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figures 5A to 5B The method may refer to a specific component, device, or system in a particular step, but this disclosure is contemplated for the execution of... Figures 5A to 5B Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0104] Figure 8The process of generating images of a virtual scene for a user according to embodiments discussed herein is illustrated. First, a camera of head-mounted device system 110 captures images of the user's environment. These images include objects (such as the user's hands) that will be used to anchor the virtual object's position within the scene and may occlude the virtual object in the scene. Head-mounted device system 110 also determines a first user pose 800a. The first user pose 800a can be determined based on one or more captured images (e.g., using SLAM or another positioning technique). The first user pose 800a can be determined based on one or more onboard sensors of head-mounted device system 110 (e.g., typically an inertial measurement unit or tracking system 140) to determine head tracking information. Head-mounted device system 110 or desktop computing system 120 can also perform hand tracking and generate or receive first hand tracking information 801a (e.g., from tracking system 140). The first hand tracking information 801a can be used to determine a first hand pose of the user's hand. The first user pose 800a and the first hand pose can be passed to the frame renderer 806 component for rendering the first image (e.g., a frame) of the virtual environment 816.
[0105] Frame renderer 806 (or related components) can generate a virtual object representation 808 of the user's hand to be used in frame 802 based on the first hand pose as described above. The virtual object representation 808 can be represented by a surface virtual object and includes a two-dimensional opaque texture, a height map, and other information required to represent the user's hand in the virtual environment 816 (such as the position of the user's hand in the environment, the boundaries of surface 808, etc.). As described above, based on the determination of one or more anchor points or one or more predetermined spatial relationships with the user's hand, the pose of virtual object 814 can be anchored to the virtual object representation 808 of the user's hand in the virtual environment 816.
[0106] Frame renderer 806 can also generate a first image 818 of virtual object 814 in virtual environment 816. To generate the first image 818 of virtual object 814, frame renderer 806 can perform initial visibility determination to determine which parts of virtual object 814 are visible based on a first user pose 800a. Visibility determination may include performing ray casting into virtual environment 816 based on the first user pose 800a, wherein the origin of each of a plurality of rays (e.g., rays 820a, 820b, 820c, and 820d) is based on a position in the display of head-mounted device system 110 and the first user pose 800a. In a particular embodiment, ray casting may be similar to ray casting performed by transform blocks 220A and 220B of display engine 112, or may be performed together with ray casting performed by transform blocks 220A and 220B of display engine 112. In a particular embodiment, occlusion determination may be supported or performed by occlusion system 145.
[0107] For each ray used for visibility determination, frame renderer 806 can project the ray into virtual environment 816 and determine whether the ray intersects a surface in the virtual environment. In a particular embodiment, a depth test (e.g., determining which surface intersects first) can be performed at the level of each surface. That is, each surface can have a single height or depth value that allows frame renderer 806 (or, for example, transform blocks 220A and 220B) to quickly identify the interactive surface. For example, frame renderer 806 can project ray 820c into virtual environment 816 and determine that ray 820c intersects virtual object 814 first. Frame renderer 806 can project ray 820a into virtual environment 816 and determine that ray 820a intersects virtual object 814 at a point near surface 808 corresponding to the user's hand. Frame renderer 806 can project ray 820d into virtual environment 816 and determine that ray 820d intersects surface 808 corresponding to the user's hand first. The frame renderer 806 can project the ray 820b into the virtual environment 816 and determine that the ray 820b does not intersect with any object in the virtual environment 816.
[0108] Each ray of light projected into the virtual environment 816 may correspond to one or more pixels of an image to be displayed to the user. Color values corresponding to the ray can be assigned based on the surfaces intersecting with it. For example, color values associated with rays 820a and 820c can be assigned based on the virtual object 814 (e.g., by sampling texture values associated with the surface). Color values associated with ray 820d can be assigned based on surface 808. In certain embodiments, color values may be specified as both opaque (e.g., no blending occurs, or light will pass through) and dark or black to provide indication to the light-emitting components of the display that will ultimately display the rendered image. A default color can be assigned to the pixels associated with ray 820b. In certain embodiments, this default color may be similar to, or the same as, the value assigned to surface 808. This default color can be selected to allow the user's environment to be visible when no virtual object is to be displayed (e.g., if there is empty space).
[0109] Frame renderer 806 can use visibility determination and color values to render a first image of the virtual object 818 to be used as frame 802. The first image of the virtual object 818 can be a surface object that includes occlusion information related to the user's hand through analysis of the virtual object representation involving hand 808. Frame renderer 806 can perform calculations required to generate this information to support a first frame rate (e.g., 30fps or 60fps). Frame renderer 806 can pass all this information to subframe renderer 812 (e.g., if frame renderer 806 is included in desktop computing system 120, frame renderer 806 can pass all this information to subframe renderer 812 via a wireless connection).
[0110] Subframe renderer 812 may receive a first image of the virtual object 818 and a virtual object representation 808 of the user's hand. Subframe renderer 812 may also (e.g., from tracking system 140) receive or determine a second user pose 800b. The second user pose 800b may differ from the first user pose 800a because even though the first user pose 800a is updated with each generated frame (e.g., each first image), the user may move slightly before the frame or associated subframe is generated and displayed. When using the artificial reality system 100, disregarding the updated user pose may significantly increase user discomfort. Similarly, subframe renderer 812 may (e.g., from tracking system 140) receive or determine second hand tracking information 801b. Like the user pose, the user's hand pose may change after the first image of the virtual object 818 is generated, but before the image and associated frame (e.g., frame 802) can be generated and displayed. The second hand tracking information 801b may include newly updated hand tracking data, or it may include the difference (e.g., increment) between the first hand tracking information 801a and the hand tracking data collected since the generation frame 802.
[0111] Subframe renderer 812 (or related components) can determine and apply adjustments to a first image of virtual object 818 to generate a related second image 826 of the virtual object. The second image 826 of the virtual object may take into account differences between a first user pose 800a and a second user pose 800b, as well as differences between first user hand tracking information 801a and second user hand tracking information 801b. As shown in the second virtual environment 830, slight movements of the user's hand, based on the second hand tracking information 801b, allow for the calculation of appropriate adjustments to the first image of virtual object 818. In a particular embodiment, color values associated with the first image 818 may also be updated. These adjustments can be applied to the first image 818, and subframe renderer 812 can render the second image 826 for display using color values determined for each ray in each ray. In a particular embodiment, this may include appropriate steps performed by pixel blocks 230A and 230B and display blocks 240A and 240B of display engine 112. Subframe renderer 812 can synthesize the determined pixel color values to prepare subframe 828 for display to the user. Subframe 828 may include a second image 826 of a virtual object that appears to include a cutout for the user's hand. When displayed by the display component of the head-mounted device system 110, this cutout allows the user's hand to actually appear in the position of the virtual object 808. Therefore, the user will be able to perceive their actual hand interacting with the virtual object 814, and the virtual object will be properly anchored based on a predetermined spatial relationship. The subframe renderer can repeat the subframe renderer process several times for each resulting frame 802 and the first image 818 of the virtual object, wherein the second hand tracking information 801b and the second user pose 800b are updated for each subframe.
[0112] Figure 9 An example computer system 900 is illustrated. In a particular embodiment, one or more computer systems 900 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 900 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 900 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular embodiments include one or more portions of one or more computer systems 900. Throughout this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.
[0113] This disclosure contemplates any suitable number of computer systems 900. This disclosure contemplates computer systems 900 employing any suitable physical form. By way of example and not limitation, computer system 900 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (such as computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service machine, a mainframe, a network of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these computer systems. Where appropriate, computer system 900 may include one or more computer systems 900; computer system 900 may be single or distributed; spanning multiple locations; spanning multiple machines; spanning multiple data centers; or residing in the cloud (which may include one or more cloud components in one or more networks). Where appropriate, one or more computer systems 900 can perform one or more steps of the methods described or illustrated herein without significant space or time constraints. By way of example and not limitation, one or more computer systems 900 can perform one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 900 can perform one or more steps of the methods described or illustrated herein at different times or in different locations.
[0114] In a particular embodiment, computer system 900 includes a processor 902, memory 904, storage device 906, input / output (I / O) interface 908, communication interface 910, and bus 912. Although this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of components in any suitable arrangement.
[0115] In a particular embodiment, processor 902 includes hardware for executing a plurality of instructions, such as those that constitute a computer program. By way of example, and not limitation, to execute the plurality of instructions, processor 902 may retrieve (or read) these instructions from internal registers, internal cache, memory 904, or storage device 906; decode and execute these instructions; and then write one or more results to internal registers, internal cache, memory 904, or storage device 906. In a particular embodiment, processor 902 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 902 including any suitable number of suitable internal caches. By way of example, and not limitation, processor 902 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The plurality of instructions in the instruction cache may be copies of the plurality of instructions in memory 904 or storage device 906, and the instruction cache may accelerate the retrieval of those instructions by processor 902. The data in the data cache may be a copy of the data in memory 904 or storage device 906 for operation by instructions executed at processor 902; the result of a previous instruction executed at processor 902 for access by subsequent instructions executed at processor 902, or for writing to memory 904 or storage device 906; or the data in the data cache may be other suitable data. The data cache can accelerate read or write operations of processor 902. Multiple TLBs can accelerate virtual address translation of processor 902. In a particular embodiment, processor 902 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 902 including any suitable number of suitable internal registers. Where appropriate, processor 902 may include one or more arithmetic logic units (ALUs); processor 902 may be a multi-core processor, or may include one or more processors 902. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0116] In a particular embodiment, memory 904 includes main memory for storing instructions to be executed by processor 902 or data to be operated by processor 902. By way of example and not limitation, computer system 900 may load multiple instructions from storage device 906 or another source (e.g., another computer system 900) into memory 904. Processor 902 may then load these instructions from memory 904 into internal registers or internal cache memory. To execute these instructions, processor 902 may retrieve and decode these instructions from internal registers or internal cache memory. During or after the execution of these instructions, processor 902 may write one or more results (which may be intermediate or final results) into internal registers or internal cache memory. Processor 902 may then write one or more of those results into memory 904. In a particular embodiment, processor 902 executes only instructions in one or more internal registers or one or more internal cache memories, or in memory 904 (different from memory device 906 or other locations), and operates only on data in one or more internal registers or one or more internal cache memories, or in memory 904 (different from memory device 906 or other locations). One or more memory buses (each memory bus may include an address bus and a data bus) couple processor 902 to memory 904. As described below, bus 912 may include one or more memory buses. In a particular embodiment, one or more memory management units (MMUs) are located between processor 902 and memory 904 and facilitate access to memory 904 requested by processor 902. In a particular embodiment, memory 904 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port RAM or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 904 may include one or more memory modules. Although this disclosure describes and illustrates specific memories, it contemplates any suitable memory.
[0117] In a particular embodiment, storage device 906 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 906 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. Where appropriate, storage device 906 may include removable or non-removable (or fixed) media. Where appropriate, storage device 906 may be internal or external to computer system 900. In a particular embodiment, storage device 906 is a non-volatile solid-state memory. In a particular embodiment, storage device 906 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. This disclosure contemplates a large-capacity storage device 906 in any suitable physical form. Where appropriate, storage device 906 may include one or more storage control units that facilitate communication between processor 902 and storage device 906. Where appropriate, storage device 906 may include one or more storage devices 906. Although this disclosure describes and illustrates specific storage devices, it contemplates any suitable storage device.
[0118] In a particular embodiment, I / O interface 908 includes hardware, software, or both hardware and software that provide one or more interfaces for communication between computer system 900 and one or more I / O devices. Where appropriate, computer system 900 may include one or more of these I / O devices. These one or more I / O devices enable communication between a person and computer system 900. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, input pad, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 908 for such I / O devices. Where appropriate, I / O interface 908 may include one or more device or software drivers that enable processor 902 to drive one or more of these I / O devices. Where appropriate, I / O interface 908 may include one or more I / O interfaces 908. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.
[0119] In a particular embodiment, the communication interface 910 includes hardware, software, or both, providing one or more interfaces for communication (e.g., packet-based communication) between the computer system 900 and one or more other computer systems 900 or one or more networks. By way of example, and not limitation, the communication interface 910 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wire-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi networks. This disclosure contemplates any suitable network and any suitable communication interface 910 for that network. By way of example, and not limitation, the computer system 900 may communicate with one or more portions of an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or a combination of two or more of these networks. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 900 may communicate with a wireless PAN (WPAN) (e.g., Bluetooth WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communication (GSM) network), or other suitable wireless networks, or a combination of two or more of these networks. Where appropriate, computer system 900 may include any suitable communication interface 910 for any of these networks. Where appropriate, communication interface 910 may include one or more communication interfaces 910. Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is contemplated in this disclosure.
[0120] In a particular embodiment, bus 912 includes hardware, software, or both hardware and software that couple multiple components of computer system 900 to each other. By way of example and not limitation, bus 912 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 912 may include one or more buses 912. Although this disclosure describes and illustrates a particular bus, this disclosure considers any suitable bus or interconnect.
[0121] In this document, where appropriate, a computer-readable non-transitory storage medium may include one or more semiconductor-based integrated circuits (ICs) or other integrated circuits (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs (ODDs), magneto-optical disks (MODs), floppy disks (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards (SD cards), any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, a computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile computer-readable non-transitory storage media.
[0122] In this document, unless otherwise expressly indicated or the context otherwise indicates, “or” is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A or B” means “A, B, or both A and B”. Furthermore, unless otherwise expressly indicated or the context otherwise indicates, “and” is both common and separate. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A and B” means “A and B, commonly or separately”.
[0123] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates various embodiments herein as including specific components, elements, features, functions, operations, or steps, those skilled in the art will understand that any embodiment in these embodiments may include any combination or arrangement of any component, element, feature, function, operation, or step described or illustrated anywhere herein. Moreover, references in the appended claims to apparatus or systems, or components in apparatus or systems (that are adapted, arranged, enabled, configured, implemented, operable, or usable to perform a particular function) cover that apparatus, system, or component (whether or not the apparatus, system, component, or the particular function is activated, turned on, or unlocked), provided that the apparatus, system, or component is so adapted, arranged, enabled, configured, implemented, operable, or usable. Furthermore, although this disclosure describes or illustrates specific embodiments to provide particular advantages, specific embodiments may not provide these advantages, or may provide some or all of these advantages.
Claims
1. A method for hand-locked rendering of virtual objects in an artificial reality environment, comprising: From one or more computing devices: Based on the first tracking data associated with the first moment, the user's first viewpoint and the user's first hand posture are determined; Based on the first hand posture and the predetermined spatial relationship between the virtual object and the user's hand, the virtual object in the virtual environment is generated; Render a first image of the virtual object as seen from the first viewpoint; Based on second tracking data associated with a second time, the user's second viewpoint and the second hand posture of the hand are determined; Based on the change from the first hand posture to the second hand posture, the first image of the virtual object is adjusted; Identify the portion of the virtual object that is obscured by the hand from the second viewpoint; Rendering a second image based on an adjusted first image viewed from the second viewpoint, wherein rendering the second image based on the adjusted first image viewed from the second viewpoint includes: generating instructions to not display the portion of the second image corresponding to the portion of the virtual object obscured by the hand; and The second image is displayed.
2. The method according to claim 1, wherein, The predetermined spatial relationship between the virtual object and the user's hand is based on one or more anchor points relative to the user's hand.
3. The method according to claim 1, further comprising: Based on the first tracking data associated with the first time, a first head pose of the user's head is determined; and The generation of the virtual object in the virtual environment is also based on the first head pose.
4. The method according to claim 3, further comprising: A second head pose of the user's head is determined based on the second tracking data associated with the second time.
5. The method according to claim 4, wherein, The adjustment of the first image of the virtual object is also based on the change from the first head pose to the second head pose.
6. The method according to claim 1, wherein, Adjusting the first image of the virtual object includes: Based on the change from the first hand pose to the second hand pose, a three-dimensional transformation of the first image of the virtual object in the virtual environment is calculated; and The three-dimensional transformation is applied to the first image of the virtual object in the virtual environment; and Rendering the second image based on the adjusted first image includes generating a projection of the adjusted first image of the virtual object based on the second viewpoint.
7. The method according to claim 1, wherein, Adjusting the first image of the virtual object includes: Based on the change from the first hand pose to the second hand pose, a visual deformation of the first image for the virtual object is calculated; and The visual distortion is applied to the first image of the virtual object; and Rendering the second image based on the adjusted first image includes: rendering the deformed first image of the virtual object based on the second viewpoint.
8. The method according to claim 7, wherein, The visual distortions include: The first image of the virtual object is scaled; Tilt the first image of the virtual object; Rotate the first image of the virtual object; or The first image of the virtual object is sliced.
9. The method according to claim 1, wherein, Rendering the second image based on the adjusted first image seen from the second viewpoint includes: The visual appearance of the first image of the virtual object is updated based on the change from the first hand pose to the second hand pose.
10. The method according to claim 1, wherein, Determining the portion of the virtual object obscured by the hand from the second viewpoint includes: A virtual object representation of the hand is generated in the virtual environment, the position and orientation of the virtual object representation of the hand being based on the second tracking data and the second viewpoint; Projecting a ray with an origin and trajectory based on the second viewpoint into the virtual environment; and Determine the intersection point between the ray and the virtual object representation of the hand, wherein the ray intersects the virtual object representation of the hand before intersecting with another virtual object.
11. The method of claim 1, wherein: The first image of the virtual object as seen from the first viewpoint is rendered by the first computing device among the one or more computing devices; and The second image, based on the adjusted first image viewed from the second viewpoint, is rendered by a second computing device among the one or more computing devices.
12. The method according to claim 1, wherein, One of the one or more computing devices includes a head-mounted display.
13. The method according to claim 12, further comprising: Based on one or more metrics of available computing resources, multiple steps of the method are allocated between the computing device including the head-mounted display and another computing device among the one or more computing devices.
14. The method according to claim 1, further comprising: After rendering the second image based on the adjusted first image seen from the second viewpoint: Based on third tracking data associated with a third time, the user's third viewpoint and the third hand gesture of the hand are determined; Based on the change from the first hand posture to the third hand posture, the first image of the virtual object is adjusted; The third image is rendered based on the adjusted first image as seen from the third viewpoint; and The third image is displayed.
15. A system for hand-locked rendering of virtual objects in an artificial reality environment, comprising: One or more processors; and one or more computer-readable non-transitory storage media, the one or more computer-readable non-transitory storage media communicating with the one or more processors, and including instructions that, when executed by the one or more processors, are configured to cause the system to perform a plurality of operations, the plurality of operations including: Based on the first tracking data associated with the first moment, the user's first viewpoint and the user's first hand posture are determined; Based on the first hand posture and the predetermined spatial relationship between the virtual object and the user's hand, the virtual object in the virtual environment is generated; Render a first image of the virtual object as seen from the first viewpoint; Based on second tracking data associated with a second time, the user's second viewpoint and the second hand posture of the hand are determined; Based on the change from the first hand posture to the second hand posture, the first image of the virtual object is adjusted; Identify the portion of the virtual object that is obscured by the hand from the second viewpoint; Rendering a second image based on an adjusted first image viewed from the second viewpoint, wherein rendering the second image based on the adjusted first image viewed from the second viewpoint includes: generating instructions to not display the portion of the second image corresponding to the portion of the virtual object obscured by the hand; and The second image is displayed.
16. The system according to claim 15, wherein, The predetermined spatial relationship between the virtual object and the user's hand is based on one or more anchor points relative to the user's hand.
17. The system according to claim 15, wherein, The instructions are also configured to cause the one or more processors to perform a plurality of operations including: Based on the first tracking data associated with the first time, a first head pose of the user's head is determined; and The generation of the virtual object in the virtual environment is also based on the first head pose.
18. One or more computer-readable non-transitory storage media, the one or more computer-readable non-transitory storage media comprising instructions, the instructions being configured, when executed by one or more processors of a computing system, to cause the one or more processors to perform a plurality of operations, the plurality of operations including: Based on the first tracking data associated with the first moment, the user's first viewpoint and the user's first hand posture are determined; Based on the first hand posture and the predetermined spatial relationship between the virtual object and the user's hand, the virtual object in the virtual environment is generated; Render a first image of the virtual object as seen from the first viewpoint; Based on second tracking data associated with a second time, the user's second viewpoint and the second hand posture of the hand are determined; Based on the change from the first hand posture to the second hand posture, the first image of the virtual object is adjusted; Identify the portion of the virtual object that is obscured by the hand from the second viewpoint; Rendering a second image based on an adjusted first image viewed from the second viewpoint, wherein rendering the second image based on the adjusted first image viewed from the second viewpoint includes: generating instructions to not display the portion of the second image corresponding to the portion of the virtual object obscured by the hand; and The second image is displayed.
19. One or more computer-readable non-transitory storage media according to claim 18, wherein, The predetermined spatial relationship between the virtual object and the user's hand is based on one or more anchor points relative to the user's hand.
20. One or more computer-readable non-transitory storage media according to claim 18, wherein, The instructions are also configured to cause the one or more processors to perform a plurality of operations including: Based on the first tracking data associated with the first time, a first head pose of the user's head is determined; and The generation of the virtual object in the virtual environment is also based on the first head pose.
Citation Information
Patent Citations
Interpolation optimizations for a display engine for post-rendering processing
US11138747B1
Display engine for post-rendering processing
US11403810B2
Generating and Modifying Representations of Objects in an Augmented-Reality or Virtual-Reality Scene
US20200134923A1
Occlusion of virtual objects in augmented reality by physical objects
US20210272361A1
Selective hand occlusion over virtual projections onto physical surfaces using skeletal tracking
US20120249590A1