Occlusion of virtual objects in augmented reality by physical objects
By generating a virtual object representation of the user's hand and selectively controlling the illumination of the light emitter, the problem of hand occlusion of virtual objects in AR is solved, improving the user experience and reducing the consumption of computing resources.
Patent Information
- Application Number
- CN202180010621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2021-01-08
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2041-01-08
AI Technical Summary
Existing AR technologies struggle to efficiently address the issue of users' hands obscuring virtual objects in virtual environments, resulting in poor user experience and excessive consumption of computing resources and battery power.
By generating a virtual object representation of the user's hand, the position of the hand in the virtual environment is determined based on a 3D mesh and height map. The illumination of the light emitter is selectively controlled to ensure that the hand is visible in front of the virtual object while other parts are occluded.
It enhances the user's immersion and comfort in the AR environment, reduces computing resources and battery consumption, and enables high frame rate virtual environment rendering.
Smart Images

Figure CN115244492B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to generating graphics for artificial reality environments. Background Technology
[0002] Artificial reality is a form of reality that has been modulated in some way before being presented to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (such as stereoscopic video that produces a three-dimensional effect for the viewer). Artificial reality may be associated with, for example, applications, products, accessories, services, or some combination thereof used to create content in and / or for artificial reality (e.g., performing activities within it). Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, standalone HMDs, mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers. Summary of the Invention
[0003] The present invention relates to methods, computer-readable non-transitory storage media and systems according to the appended claims.
[0004] In a particular embodiment, a method is performed by one or more computing systems of an artificial reality system. The computing system may be embodied in a head-mounted display or a less portable computing system. The method includes accessing an image of a user's hand, including the head-mounted display. The image may also include the user's environment. The image may be generated by one or more cameras of the head-mounted display. The method may include generating at least a virtual object representation of the user's hand based on the image, the virtual object representation of the hand being defined in a virtual environment. The virtual object representation of the user's hand may be generated based on a three-dimensional mesh representing the user's hand in the virtual environment, the three-dimensional mesh having been prepared based on detected poses of the user's hand in the environment. The method may include rendering an image of the virtual environment from the user's viewpoint to the virtual environment, based on the virtual object representation of the hand and at least one other virtual object in the virtual environment. The viewpoint to the virtual environment may be determined based on a correspondence between the user's viewpoint to the real environment and the user's viewpoint to the virtual environment. The rendered image may include a set of pixels corresponding to a portion of the virtual object representation of the hand visible from the user's viewpoint. This set of pixels may be determined by determining that the virtual object representation of the user's hand is at least partially in front of other virtual objects in the virtual environment from the user's viewpoint. The method may include: providing instructions to a set of light emitters of a head-mounted display for displaying an image of a virtual environment, wherein a set of pixels in the image corresponding to a portion of a virtual object representation of a hand de-illuminates light emitters at one or more locations. The de-illuminated light emitters in the head-mounted display allow light from the user's environment to continue to be perceived by the user.
[0005] In one embodiment of the method according to the invention, the absence of light emanating from a light emitter at a specific location allows light from the user's environment to continue reaching the user at that specific location.
[0006] In one embodiment of the method according to the invention, the image including the user's hand may also include the user's environment as seen from the user's viewpoint.
[0007] In one embodiment of the method according to the invention, generating a virtual object representation of a user's hand based on an image may include: determining the position of the virtual object representation of the user's hand in a virtual environment based on the pose of the user's hand determined from an image including the user's hand.
[0008] In one embodiment of the method according to the invention, the texture represented by the virtual object of the user's hand may correspond to an instruction to prevent the light emitter from illuminating.
[0009] In one embodiment of the method according to the invention, a virtual object representation of a user's hand can be associated with a color that is also associated with the background of the virtual environment.
[0010] In one embodiment of the method according to the invention, rendering an image of a virtual environment from a user's viewpoint may include: determining whether a virtual object representation of the user's hand and at least one other virtual object in the virtual environment are visible from the user's viewpoint. In addition, the method may further include: determining that the virtual object representation of the hand is at least partially in front of one or more positions in the virtual environment by: projecting light rays into the virtual environment having a starting point and direction based on the user's viewpoint; and determining the point where the light rays intersect with the virtual object representation of the user's hand in the virtual environment, wherein the light rays intersect the virtual object representation before intersecting with another object in the virtual environment.
[0011] In one embodiment of the method according to the invention, generating a virtual object representation of a user's hand based on an image may include:
[0012] At least determine the hand posture based on the image;
[0013] At least a triangular mesh representing the virtual object of the hand is generated based on the image and pose;
[0014] At least determine the distance of the hand from the user's viewpoint based on the image; and
[0015] A height map is generated based on a triangular mesh representing the hand, indicating changes in one or more positions of the hand according to a determined distance from the hand. In addition, the method may further include: determining, based on the height map, one or more positions of a virtual object representation of the hand in front of another virtual object in the virtual environment by comparing the distance of the hand and the height map associated with the virtual object representation of the hand with the distance and the height map associated with another virtual object at a specific location; and determining, based on the comparison, that the virtual object representation of the hand is the object closest to the viewpoint.
[0016] In one embodiment of the method according to the invention, the instructions for displaying an image of a virtual environment may also cause a light emitter to illuminate at one or more locations, wherein, from the user's viewpoint, at least a portion of another virtual object is in front of a virtual object representation on the user's hand.
[0017] In one embodiment of the method according to the invention, one or more computing devices may be embodied in a head-mounted display, and the one or more computing devices are separate computing devices. Furthermore, the method may also include the step of allocating methods between the computing devices of the head-mounted display and the separate computing devices based on one or more metrics of available computing resources.
[0018] In one embodiment of the method according to the invention, the image may be generated by a first camera of a head-mounted display, and generating a virtual object representation of the user's hand may further include accessing a second image generated by a second camera of the head-mounted display; and locating the user's hand relative to the user's viewpoint based on the image and the second image. Furthermore, generating a virtual object representation of the user's hand may further include generating an array of positions corresponding to the positions of the user's hand, storing values of distances between one or more positions of the hand and the user's viewpoint in the array, and associating the array with the virtual object representation of the user's hand. Alternatively or otherwise, the second camera of the head-mounted display may be a depth-sensing camera.
[0019] In one aspect, the present invention relates to one or more computer-readable non-transitory storage media embodying software that, when executed, is operable to perform the methods described above, or:
[0020] Access images including the user's hands on the head-mounted display;
[0021] At least a virtual object representation of the user's hand is generated based on the image, and the virtual object representation of the hand is defined in the virtual environment;
[0022] Based on a virtual object representation of the hand and at least one other virtual object in the virtual environment, an image of the virtual environment is rendered from the user's viewpoint. The image includes a set of pixels corresponding to a portion of the virtual object representation of the hand visible from the user's viewpoint.
[0023] Instructions are provided to the light emitter set of the head-mounted display for displaying an image of a virtual environment, wherein the set of pixels in the image corresponding to a portion of a virtual object representation of a hand prevents the light emitters at one or more locations from being illuminated.
[0024] In one embodiment of a computer-readable non-transitory storage medium, the absence of light emanating from a light emitter at a particular location may allow light from the user's environment to continue illuminating the user at that location.
[0025] Embodiments of the present invention may include artificial reality systems or combinations thereof. Artificial reality is a form of reality that has been modulated in some way before being presented to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), mixed reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, and any of these may be presented in a single channel or multiple channels (such as stereoscopic video that produces a three-dimensional effect for the viewer). Furthermore, in certain embodiments, artificial reality may be associated with, for example, applications, products, accessories, services, or some combination thereof for creating content in and / or using in artificial reality (e.g., performing activities therein). Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.
[0026] A system according to the present invention may include one or more processors, and one or more computer-readable non-transitory storage media coupled to the processors and including instructions operable when executed by the one or more processors to cause the system to perform the methods described above, or:
[0027] Access images including the user's hands on the head-mounted display;
[0028] At least a virtual object representation of the user's hand is generated based on the image, and the virtual object representation of the hand is defined in the virtual environment;
[0029] Based on a virtual object representation of the hand and at least one other virtual object in the virtual environment, an image of the virtual environment is rendered from the user's viewpoint. The image includes a set of pixels corresponding to a portion of the virtual object representation of the hand visible from the user's viewpoint.
[0030] Instructions are provided to the light emitter set of the head-mounted display for displaying an image of a virtual environment, wherein the set of pixels in the image corresponding to a portion of a virtual object representation of a hand prevents the light emitters at one or more locations from being illuminated.
[0031] In one embodiment of the system according to the invention, if the light emitter at a specific location is not illuminated, light from the user's environment can continue to reach the user at that specific location.
[0032] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Specific embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments of the invention are disclosed in particular in the appended claims relating to methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from the intentional retrospection to any prior claim (particularly multiple dependencies) may also be claimed so that any combination of claims and their features is disclosed and can be claimed, regardless of the dependency chosen in the appended claims. Claimable subject matter includes not only combinations of features as described in the appended claims but also any other combination of features in the claims, wherein each feature mentioned in a claim may be combined with any other feature or combination of features in the claims. Furthermore, any embodiments and features described or depicted herein may be claimed in individual claims and / or in any combination with any embodiments or features described or depicted herein or with any features of the appended claims. Attached Figure Description
[0033] Figure 1A An example artificial reality system is shown.
[0034] Figure 1B An example eye display system of a head-mounted device system is shown.
[0035] Figure 2 A system diagram for displaying the engine is shown.
[0036] Figures 3A-3B Example images viewed through an artificial reality system are shown.
[0037] Figures 4A-4B Example images viewed through an artificial reality system are shown.
[0038] Figure 5 This illustrates the visual representation of a user's hand when detecting occlusion by a virtual object.
[0039] Figure 6 A visual representation of the image used to generate the virtual environment is shown.
[0040] Figure 7 A visual representation of the image used to generate the virtual environment is shown.
[0041] Figures 8A-8BAn example method for providing real-world object occlusion for virtual objects in augmented reality is shown.
[0042] Figure 9 An example computer system is shown. Detailed Implementation
[0043] In a particular embodiment, a method is performed by one or more computing systems of an artificial reality system. The computing system may be embodied in a head-mounted display or a less portable computing system. The method includes accessing an image of a user's hand, including the head-mounted display. The image may also include the user's environment. The image may be generated by one or more cameras of the head-mounted display. The method may include generating at least a virtual object representation of the user's hand based on the image, the virtual object representation of the hand being defined in a virtual environment. The virtual object representation of the user's hand may be generated based on a three-dimensional mesh representing the user's hand in the virtual environment, the three-dimensional mesh having been prepared based on detected poses of the user's hand in the environment. The method may include rendering an image of the virtual environment from the user's viewpoint to the virtual environment based on the virtual object representation of the hand and at least one other virtual object in the virtual environment. The viewpoint to the virtual environment may be determined based on a correspondence between the user's viewpoint to the real environment and the user's viewpoint to the virtual environment. The rendered image may include a set of pixels corresponding to a portion of the virtual object representation of the hand visible from the user's viewpoint. This set of pixels can be determined by determining that the virtual object representation of the user's hand is at least partially in front of other virtual objects in the virtual environment from the user's viewpoint. The method may include providing instructions to a set of light emitters in a head-mounted display for displaying an image of a virtual environment, wherein a set of pixels in the image corresponding to a portion of a virtual object representation of a hand de-illuminates light emitters at one or more locations. The de-illuminated light emitters in the head-mounted display allow light from the user's environment to continue to be perceived by the user.
[0044] In certain embodiments, this disclosure relates to the problem of associating physical objects with a virtual environment presented to a user. In augmented reality (AR), a virtual environment can be displayed to the user as an augmentation layer on top of a real environment. This can be done by creating a correspondence or mapping from viewpoints in the real environment to viewpoints in the virtual environment. Embodiments of this disclosure relate to the task of efficiently rendering the occlusion of virtual objects by physical objects (e.g., a user's hand) in augmented reality. When a user's hand is seen through an AR device (e.g., a head-mounted display (HMD)), the user naturally expects that portions of each hand in front of a virtual object in the user's line of sight should occlude the virtual object in front of the user's hand. Conversely, it is naturally expected that any portion of each hand behind a virtual object in the user's line of sight should be occluded by the virtual object. For example, when a hand is holding an object, a portion of the hand may be behind the object. The user may expect that the portion of the object behind the hand should not be displayed so that the user will be able to see her own physical hand in front of the occluded object. Presenting users with a view of their own physical hand in front of or interacting with a virtual object can help users feel comfortable in the augmented reality environment. For example, embodiments of this disclosure can help reduce motion sickness or susceptibility to motion sickness in users.
[0045] Current AR technologies cannot efficiently address these issues. In a common approach to presenting AR experiences, users view the environment through a standard screen (e.g., a smartphone screen). The virtual environment is overlaid on an image of the environment captured by a camera. This requires significant computational resources, as the environmental images must be captured and processed rapidly, quickly draining the mobile device's battery. Furthermore, this type of experience is not particularly immersive for the user, as they are limited to viewing the environment through a small screen. In related approaches, many current systems struggle to accurately detect the user's hands using available camera technology, allowing the user's hands to be used to manipulate virtual objects in the virtual environment. There is a lack of advanced techniques, such as those disclosed here, to accurately model the user's hands in the virtual environment and the effects caused by those hands within it, and to render the virtual environment based on those effects. As an additional example, current rendering methods for artificial reality systems cannot render most virtual environments at sufficiently high frame rates and quality levels to allow users to comfortably experience them for any considerable period. As discussed here, high frame rates can be particularly advantageous for mixed or augmented reality experiences, as the juxtaposition of virtual objects with the user's real environment allows the user to quickly identify any technical glitches in the rendering. The method described in this article solves all the technical problems, as well as many others.
[0046] In this disclosure, an example of a user's hand will be given; however, the methods described herein can be used for other types of objects. Such other objects include other parts of the user's body, objects held by the user (e.g., a pen or other pointer), specific objects designated for passage by the user (e.g., the user's child or pet), general objects designated for passage by the user or the provider of the AR headset (e.g., vehicles, other people), and many other objects. To allow real objects to pass, such as hand occlusion, one or more cameras on the AR headset (e.g., an HMD) capture images of the scene, including identifying the occluded object. A computing device (which, as described herein, may be embodied in the HMD or may communicate with the HMD via wired or wireless communication) performs a hand-tracking algorithm to detect the hand in the image. The positions of hand features (such as fingers and joints) in the image are then determined. A virtual object representation of the hand (e.g., a 3D mesh that looks like the user's hand) is generated based on the detected positions of the fingers and joints.
[0047] The computing device determines the distance from the user's hand to the user's viewpoint in the real environment. The computing device correlates this distance with the distance from a virtual representation of the hand to the user's viewpoint in the virtual environment. The computing device also creates a grid to store height information for different regions of the hand. Based on multiple images and the 3D grid, the computing device determines the height of points on the hand (e.g., the height of a specific point on the hand relative to the average or median height of the hand or a reference point). The determined height indicates the position of points on the hand relative to the rest of the hand. Combined with the determined distances, the height can be used to determine the precise position of various parts of the hand relative to the user.
[0048] When rendering and presenting a virtual environment to a user, the portion of the hand closer to the user than any virtual object should be visible, while portions of the hand behind at least one virtual object should be occluded. Through the HMD of the AR system, the user can see their actual physical hand. In a particular embodiment, the light-emitting components (e.g., LEDs) used to display virtual objects in the virtual environment in the HMD can be selectively disabled to allow light from the real environment to reach the user's eyes through the HMD. That is, the computing device creates a cut-out in the rendered image, in which the hand will appear. Thus, for example, by indicating the display position where the light emitter is not illuminating the position corresponding to the thumb, the user can see their actual physical thumb through the HMD. Since the light emitter is turned off, portions of any virtual objects behind it are not displayed.
[0049] To correctly render real-world object occlusion, the light emitter corresponding to a portion of an object (e.g., a portion of a user's hand) should be turned off or indicated as unlit when that portion of the hand is closer to the user in the real environment than any virtual object in the virtual environment, and turned on when a virtual object exists between the user and that portion of the hand. Portions of objects that are farther from the user in the virtual environment than in the real environment (e.g., fingers on the hand) are shown as being behind the virtual object by the light emitter. Comparing distances to real-world objects and distances to virtual objects is possible because, for example, the distances are determined by a hand tracking algorithm. The virtual object distances are known to the application or scene being performed on the AR system and are available to both the AR system and the HMD.
[0050] Given a virtual object representation of a user's hand and a known height of the hand's position, the HMD can make the portion of the hand visible to the user visible, as follows. Frames showing the virtual environment are rendered by the main rendering component at a first frame rate (such as 30fps) based on the user's current pose (e.g., position and orientation). As part of the rendered frame, two items are generated: (1) a two-dimensional opaque texture for the hand based on a 3D mesh, and (2) a height map for the hand. The two-dimensional texture is saved as a texture representing a planar object (also referred to as a "surface" throughout this disclosure). These operations can be performed by the HMD or by a separate computing device (e.g., a cloud computer, desktop computer, laptop computer, or mobile device) communicating with the HMD hardware. By using colors specifically assigned to the texture, it is easy to indicate that a light emitter is not illuminated. In AR, a default background that allows light from the real environment to pass through may be desired. That would greatly enhance the immersive experience of virtual objects appearing in a real environment. Therefore, the background can be associated with a color that is translated into an instruction for a light emitter not being illuminated. In certain embodiments, this color may be referred to as "opaque black," for example, indicating that no light passes through behind the object associated with the texture. Virtual object representations (e.g., planar objects) may be associated with such a color to turn off (i.e., not illuminated) the LEDs corresponding to pixels in a position visible to the hand.
[0051] The HMD then renders subframes at a second frame rate (such as 200fps) based on the previously generated frames (e.g., based on frames generated at 30fps). For each subframe, the AR system can perform a primary visibility test for each individual pixel or pixel tile in the virtual environment, based on the user's pose (which may differ from the pose used to generate the primary 30fps frame), by projecting one or more rays from the user's current viewpoint into the virtual environment. Rays are rays in the case of individual pixels, or can be conceptualized as cones in the case of tiles. In certain embodiments, the virtual environment may have a limited number of virtual objects. For example, the AR system may limit the number of discrete objects in the virtual environment (including objects created to represent the user's hand). In some embodiments, multiple objects expected to move in a similar manner for a series of frames can be grouped together. Virtual objects can be represented by planar objects with corresponding heightmap information. If a projected ray intersects a surface corresponding to the user's hand, it samples the associated texture to render the subframe. Since the texture is opaque black, indicating that the light emitter is not illuminated, the actual physical hand is visible through the unilluminated area of the HMD, which is transparent.
[0052] In certain embodiments, the AR system can support per-pixel height testing for each surface based on a heightmap. In this case, the portion of a hand in front of a virtual object can occlude the virtual object at pixel-level resolution, while pixels of the same hand behind the virtual object can be occluded by individual pixels of the virtual object on the display. In embodiments using heightmaps, per-pixel height testing is determined based on the depth of the real or virtual object from the user's viewpoint plus the difference recorded in the heightmap. Using per-pixel height for virtual objects significantly improves the visual appearance of hand occlusion.
[0053] In certain embodiments, per-pixel height information may be unavailable (e.g., due to limitations in available computing resources). In this case, the hand can be represented by a virtual object simply as a planar object. As mentioned above, the depth of the planar object (e.g., the distance from the planar object to the user's viewpoint in the virtual environment) is known. In certain embodiments, the depth can be determined using mathematical transformations for each location on a flat surface. The virtual object representations of real objects (e.g., planar objects) and virtual objects (which themselves can be represented by surfaces) form an ordered grouping of surfaces, where the order is based on the depth of the objects. Thus, the planar object corresponding to the end is located in front of or behind each plane.
[0054] As mentioned above, using heightmaps to determine the actual object occlusion of virtual objects can increase power consumption. For example, if the HMD generates the heightmap, it can consume significant resources. As another example, if the HMD is generated by another computing device, the heightmap may have to be transmitted to the HMD. The AR system can determine whether and how to use a heightmap based on available power. If sufficient power is available, a heightmap can be generated and transmitted. If insufficient power is available, using a plane without a heightmap may be sufficient to produce a reasonable appearance of hand occlusion, especially when the distance between the virtual object and the user's viewpoint is large.
[0055] In certain embodiments, the primary rendering device (e.g., a computing device other than the HMD) may have more processing power and power supply than the HMD. The primary rendering device can therefore perform certain parts of the techniques described herein. For example, the primary rendering device can perform hand tracking calculations, generating a 3D mesh, a 2D opaque texture, and a height map for the hand. For example, the primary rendering device can receive images created by the HMD's camera and perform the necessary processing using specially designed computing hardware to be more efficient or powerful on a particular task. However, if the HMD has sufficient processing power and power, it can perform one or more of these steps (e.g., using its own onboard computing components) to reduce latency. In certain embodiments, all steps described herein are performed on the HMD.
[0056] Figure 1AAn example artificial reality system 100 is illustrated. In a particular embodiment, the artificial reality system 100 may include a head-mounted device system 110 (which may be embodied in an HMD), a body-worn computing system 120, a cloud computing system 132 in a cloud computing environment 130, etc. In a particular embodiment, the head-mounted device system 110 may include a display engine 112 connected to two eye display systems 116A and 116B via a data bus 114. The head-mounted device system 110 may be a system including a head-mounted display (HMD) that can be mounted on a user's head to provide artificial or augmented reality to the user. The head-mounted device system 110 may be designed to be lightweight and highly portable. Therefore, the available power in the head-mounted device system's power source (e.g., battery) may be limited. The display engine 112 may provide display data to the eye display systems 116A and 116B via the data bus 114 at a relatively high data rate (e.g., suitable for supporting refresh rates of 200 Hz or higher). The display engine 112 may include one or more controller blocks, texel memories, transform blocks, pixel blocks, etc. The texels stored in the texel memory can be accessed by the pixel blocks and can be provided to the eye display systems 116A and 116B for display. Further information regarding the described display engine 112 can be found in U.S. Patent Application No. 16 / 657,820, filed October 1, 2019; U.S. Patent Application No. 16 / 586,590, filed September 27, 2019; and U.S. Patent Application No. 16 / 586,598, filed September 27, 2019, which are incorporated herein by reference.
[0057] In certain embodiments, the body-worn computing system 120 may be worn on a user's body. In certain embodiments, the body-worn computing system 120 may be a computing system not worn on the user's body (e.g., a laptop, desktop, or mobile computing system). The body-worn computing system 120 may include one or more GPUs, one or more intelligent video decoders, memory, a processor, and other modules. The body-worn computing system 120 may have more computing resources than the display engine 112, but in some embodiments, the power in its power supply (e.g., a battery) may still be limited. The body-worn computing system 120 may be coupled to the head-mounted device system 110 via a wireless connection 144. The cloud computing system 132 may include a high-performance computer (e.g., a server) and may communicate with the body-worn computing system 120 via the wireless connection 142. In some embodiments, the cloud computing system 132 may also communicate with the head-mounted device system 110 via a wireless connection (not shown). The body-worn computing system 120 may generate data for rendering at a standard data rate (e.g., suitable for supporting refresh rates of 30 Hz or higher). Display engine 112 can upsample data received from body wearable computing system 120 to generate frames to be displayed by eye display systems 116A and 116B at a higher frame rate (e.g., 200Hz or higher).
[0058] Figure 1B An example eye-mounted display system (e.g., 116A or 116B) is shown for a head-mounted device system 110. In a particular embodiment, the eye-mounted display system 116A may include a driver 154, a pupil display 156, etc. The display engine 112 may provide display data to the pupil display 156, the data bus 114, and the driver 154 at a high data rate (e.g., suitable for supporting refresh rates of 200 Hz or higher).
[0059] Figure 2 A system diagram of display engine 112 is shown. In a particular embodiment, display engine 112 may include control block 210, transformation blocks 220A and 220B, pixel blocks 230A and 230B, display blocks 240A and 240B, etc. One or more components of display engine 112 may be configured to communicate via a high-speed bus, shared memory, or any other suitable method. Figure 2 As shown, the control block 210 of the display engine 112 can be configured to communicate with the transformation blocks 220A and 220B, the pixel blocks 230A and 230B, and the display blocks 240A and 240B. This communication may include data, control signals, interrupts, and other instructions, as explained in further detail herein.
[0060] In a particular embodiment, control block 210 can be controlled from a wearable computing system (e.g., Figure 1A The control block 210 receives input and initializes the pipeline in the display engine 112 to complete rendering for display. In a particular embodiment, the control block 210 may receive data and control packets from the wearable computing system at a first data rate or frame rate. The data and control packets may include information such as one or more data structures, which include texture data and position data, as well as additional rendering instructions. In a particular embodiment, the data structure may include two-dimensional rendering information. The data structure may be referred to herein as a “surface.” The control block 210 may distribute data to one or more other blocks of the display engine 112 as needed. The control block 210 may initiate pipeline processing for one or more frames to be displayed. In a particular embodiment, each of the eye display systems 116A and 116B may include its own control block 210. In a particular embodiment, one or more of the eye display systems 116A and 116B may share the control block 210.
[0061] In a particular embodiment, transform blocks 220A and 220B can determine initial visibility information of surfaces to be displayed in an artificial reality scene. Typically, transform blocks 220A and 220B can project rays with originating points based on pixel locations in the image to be displayed and generate filtering commands (e.g., filtering based on bilinear or other types of interpolation techniques) to be sent to pixel blocks 230A and 230B. Transform blocks 220A and 220B can perform ray projection based on the user's current viewpoint to the user's real or virtual environment. The user's viewpoint can be determined using sensors from a head-mounted device, such as one or more cameras (e.g., monochrome, panchromatic, depth sensing), an inertial measurement unit, an eye tracker, and / or any suitable tracking / localization algorithm, such as Simultaneous Localization and Mapping (SLAM) of the environment and / or virtual scene, where surfaces are localized, and the results can be sent to pixel blocks 230A and 230B.
[0062] Typically, according to a particular embodiment, both transform blocks 220A and 220B may include a four-stage pipeline. The stages of transform blocks 220A or 220B may be performed as follows: A ray projector may emit ray beams corresponding to an array of one or more aligned pixels, referred to as tiles (e.g., each tile may include 16×16 aligned pixels). Before entering the artificial reality scene, the ray beams may be distorted according to one or more distortion meshes. The distortion meshes may be configured to correct for geometric distortion effects originating at least from the eye display systems 116A and 116B of the head-mounted device system 110. In a particular embodiment, transform blocks 220A and 220B may determine whether each ray beam intersects a surface in the scene by comparing the bounding box of each tile with the bounding box of a surface. If a ray beam does not intersect an object, it may be discarded. Tile surface intersections are detected, and the corresponding tile surface pairs are passed to pixel blocks 230A and 230B.
[0063] Typically, according to a particular embodiment, pixel blocks 230A and 230B can determine color values from tile surface pairs to generate pixel color values. The color value for each pixel can be sampled from the texture data of the surface received and stored by control block 210. Pixel blocks 230A and 230B can receive tile surface pairs from transform blocks 220A and 220B and can schedule bilinear filtering. For each tile surface pair, pixel blocks 230A and 230B can sample the color information for the pixel corresponding to the tile using the color value corresponding to the intersection of the projected tile and the surface. In a particular embodiment, pixel blocks 230A and 230B can process the red, green, and blue components separately for each pixel. In a particular embodiment, as described herein, the pixel blocks can employ one or more processing shortcuts based on the color associated with the surface (e.g., color and opacity). In a particular embodiment, the pixel block 230A of the display engine 112 of the first-eye display system 116A can operate independently and in parallel with the pixel block 230B of the display engine 112 of the second-eye display system 116B. The pixel block can then output its color determination to the display block.
[0064] Typically, display blocks 240A and 240B can receive pixel color values from pixel blocks 230A and 230B, convert the data format to be more suitable for the display (e.g., if the display requires a specific data format, such as in a scanline display), apply one or more brightness corrections to the pixel color values, and prepare the pixel color values for output to the display. Display blocks 240A and 240B can convert the tile-order pixel color values generated by pixel blocks 230A and 230B into scanline or line-order data that the physical display may require. Brightness correction may include any necessary brightness correction, gamma mapping, and dithering. Display blocks 240A and 240B can output the corrected pixel color values directly to the physical display (e.g., ...). Figure 1B The pupil display 156 in the display engine 112 can output pixel values to a block outside the display engine 112 via a driver 154, or in various formats. For example, eye display systems 116A and 116B or head-mounted device system 110 may include additional hardware or software to further customize back-end color processing to support wider display interfaces or optimize display speed or fidelity.
[0065] In a particular embodiment, controller block 210 may include microcontroller 212, texel memory 214, memory controller 216, data bus 217 for I / O communication, data bus 218 for input stream data 205, etc. Memory controller 216 and microcontroller 212 can be coupled via data bus 217 for I / O communication with other modules of the system. Microcontroller 212 can receive control packets such as position data and surface information via data bus 217. Input stream data 205 can be input to controller block 210 from the wearable computing system after being set by microcontroller 222. Input stream data 205 can be converted into the desired texel format by memory controller 216 and stored in texel memory 214. In a particular embodiment, texel memory 214 may be static random access memory (SRAM).
[0066] In a particular embodiment, the wearable computing system can send input stream data 205 to a memory controller 216, which can convert the input stream data into texels with the desired format and store the texels in a swizzle mode in a texel memory 214. The texel memory organized in these swizzle modes allows pixel blocks 230A and 230B to use a single read operation to retrieve the texels (e.g., in a 4×4 texel block) needed to determine at least one color component (e.g., red, green, and / or blue) of each pixel in all pixels associated with a tile (e.g., a "tile" refers to an aligned pixel block, such as a 16×16 pixel block). As a result, if the texel array is not stored in the appropriate mode, the display engine 112 can avoid the excessive multiplexing operations typically required for reading and assembling the texel array, and thus comprehensively reduce the computational resource requirements and power consumption of the display engine 112 and the head-mounted device system.
[0067] In a particular embodiment, pixel blocks 230A and 230B can generate pixel data for display based on texels retrieved from texel memory 212. Memory controller 216 can be coupled to pixel blocks 230A and 230B via two 256-bit data buses 204A and 204B, respectively. Pixel blocks 230A and 230B can receive tile / surface pairs 202A and 202B from corresponding transform blocks 220A and 220B, and can identify texels needed to determine at least one color component of all pixels associated with a tile. Pixel blocks 230A and 230B can retrieve the identified texels (e.g., a 4×4 texel array) in parallel from texel memory 214 via memory controller 216 and 256-bit data buses 204A and 204B. For example, the 4×4 texel array needed to determine at least one color component of all pixels associated with a tile can be stored in a memory block and retrieved using a memory read operation. Pixel blocks 230A and 230B can use multiple sample filter blocks (e.g., one for each color component) to perform interpolation on different groups of texels in parallel to determine the corresponding color component for the corresponding pixel. Pixel values 203A and 203B for each eye can be sent to display blocks 240A and 240B for further processing before being displayed by eye display systems 116A and 116B, respectively.
[0068] In certain embodiments, the artificial reality system 100, particularly the head-mounted device system 110, can be used to render an augmented reality environment to a user. The augmented reality environment may include elements of a virtual reality environment (e.g., virtual reality objects) rendered for the user, such that virtual elements appear on top of or a portion of the user's real environment. For example, the user may wear an HMD (e.g., head-mounted device system 110) embodying features of the techniques disclosed herein. The HMD may include a display that allows light from the user's environment to normally continue reaching the user's eyes. However, when the light-emitting components of the display (e.g., LEDs, OLEDs, microLEDs, etc.) are illuminated at specific locations on the display, the color of the LEDs may be superimposed on the user's environment. Thus, when the light emitter of the display is illuminated, virtual objects may appear in front of the user's real environment, while when the light-emitting components at a specific location are not illuminated, the user's environment may be visible through the display at that specific location. This can be done without re-rendering the environment using a camera. This process can increase the user's immersion and comfort when using the artificial reality system 100, while reducing the computational power and battery consumption required to render the virtual environment.
[0069] In certain embodiments, selectively illuminated displays can be used to facilitate user interaction with virtual objects in a virtual environment. For example, the position of a user's hand can be tracked. In previous systems, even hand-tracking systems, the only way to represent a user's hand in a virtual environment was to generate and render some virtual representation of the hand. This may be unsuitable in many use cases. For example, in a workplace environment, it may be desirable for a user's hand to be visible to the user as they interact with objects. For instance, an artist may want to see their own hand when they are holding a virtual paintbrush or other tool. Techniques for achieving such advanced displays are disclosed herein.
[0070] Figure 3A An example of a user's internal view from a head-mounted device is shown, illustrating an example of a user viewing their own hand through an artificial reality display system. The head-mounted device's internal view 300 shows a synthetic (e.g., mixed) reality display comprising a view of a virtual object 304 and a view of the user's physical hand 302. The user is attempting to press a button on the virtual object 304. Thus, the user's hand 302 is positioned closer to the user's viewpoint than a portion of the virtual object 304. This aligns with the user's common intuition, as their hand appears in front of the elevator button and panel portion when attempting to press a button (e.g., an elevator button).
[0071] Figure 3B Alternative views for the same scene are shown. Figure 3A An internal view 300 of the user's head-mounted device is shown. Figure 3BOnly pixels 310, displayed by the light-emitting components of the head-mounted device system, are shown. These pixels include a representation of the virtual object 304. However, since any illuminating light emitter in the display system will cause the virtual object 304 to be overlaid on the environment, the rendering of the virtual object 304 must be modified so that the user's hand can be displayed on the monitor (e.g., ...). Figure 3A (As shown). Therefore, the rendering of the virtual object 304 is modified to include a contour map of the shape closely matching the user's hand. When the virtual environment is displayed, the contour map portion is treated as an area of unlit pixels 312. In other words, the luminous components that normally display the color of the virtual object 304 are instead indicated as unlit. The effect is to allow ambient light to continue passing through, thereby producing... Figure 3A The resulting composite effect is shown.
[0072] Figure 4A Another example of an internal view of a user's head-mounted device is shown. Figure 4A In the view 400 inside the head-mounted device, a user interacts with the virtual object 404 by wrapping their fingers around it. A view of the user's physical hand 402 shows the virtual object 404 partially obscuring the user's hand (e.g., the palm near the user's hand) and partially obscured by the user's hand 402 (e.g., the user's fingers). Due to the partial rendering techniques described herein, users can more intuitively understand that they are holding the virtual object 404, even if they do not physically feel the object.
[0073] Figure 4B It shows Figure 4A An alternative view of the scene shown. (Compared to) Figure 3B Same, Figure 4B Only pixel 410, displayed by the light-emitting components of the head-mounted device system, is shown. The pixels include a representation of virtual object 404. Areas of virtual object 404 are not displayed (e.g., the virtual object includes unilluminated pixel areas 412) because these areas are obscured by the user's hand in both the virtual and real-world views within the head-mounted device. Note that... Figure 4B This demonstrates that merely determining the location of the user's hands within the scene and always displaying the user's hands is insufficient. Consider 4A and Figure 4B This is a counterexample to the situation shown. If the user's hand is always made to appear in front of any virtual object, the artificial reality system 100 will be unable to render as shown. Figure 4A The scenario shown depicts a user's hand partially occluded by a virtual object 404. Therefore, in order to accurately display the user's hand in a mixed reality environment where interaction with virtual objects is possible, it is necessary to track the depth of individual portions of the user's hand (e.g., the distance between a point and the user's viewpoint) and compare it to the depth of the virtual object.
[0074] Figure 5 This illustrates a graphical representation of an image used to process a user's hand when determining which parts of the user's hand are occluded or obscured by virtual objects. Figure 5 This can be understood as illustrating a method for generating an image of a user's hand used in this disclosure and the state of its virtual representation. Figure 5 The example shown is from Figures 3A to 3B Continuing with the example, one or more images 302 of the user's physical hand are captured by one or more cameras of the head-mounted device system 110. A computing system (which in some embodiments may include the head-mounted device system 110 or a body-wearable computing system 120) performs a hand-tracking algorithm to determine the user's hand pose based on the relative positions of various discrete locations on the user's hand. For example, the hand-tracking algorithm may detect the positions of the fingertips, joints, palm, and back of the hand, as well as other different parts of the user's hand. The computing system may also determine the depth of the user's hand (e.g., the distance between the hand and the user's viewpoint represented by the camera).
[0075] Based on the determined pose, the computing system generates a 3D mesh 502 for the hand. The 3D mesh may include a virtual object representation of the user's hand. In a particular embodiment, the 3D mesh may be similar to a 3D mesh that can be generated to render a representation of the user's hand in an immersive artificial reality scene. The computing system also generates a height mesh 504 overlay for the 3D mesh. The computing system can generate the height mesh based on the 3D mesh. For example, the computing system can determine a fixed plane relative to the hand (e.g., based on the location of a landmark such as the user's palm). The computing system can determine the changes relative to the fixed plane based on the 3D mesh so that the computing system can determine the depth deviation of specific locations (e.g., fingertips, joints, etc.) from the fixed plane. These changes at points can be stored in the height mesh. The computing system can also calculate the height at the 3D mesh location by interpolating known distances (e.g., tracking the position) to determine inferred distances.
[0076] As described above, the 3D mesh 502 and height mesh 504 can be generated by the body-wearable computing system 120 based on the available computing and power resources of the body-wearable computing system 120 and the head-mounted device system 110. The artificial reality system 100 can seek to balance considerations of rendering latency and graphics quality with the availability of resources of the computing systems involved in the artificial reality system 100, such as computing power, memory availability, and power availability. For example, the artificial reality system 100 can proactively manage which computing systems are handling tasks such as hand tracking, generating the 3D mesh 502, and generating the height mesh 504. The artificial reality system 100 can determine that the head-mounted device system has sufficient battery power and processor availability and instruct the head-mounted device system 110 to perform these steps. In a particular embodiment, it may be preferred to allow the head-mounted device system 110 to handle as many rendering processes as possible to reduce latency introduced by data transfer between the head-mounted device system 110 and the body-wearable computing system 120.
[0077] 3D mesh 502 and height mesh 504 can be passed to frame renderer 506. In a particular embodiment, the frame renderer may be embodied in body-worn computing system 120 or another computing device with more available computing resources than head-mounted device system 110. Frame renderer 506 can convert 3D mesh 502 into a surface representation that includes two-dimensional virtual object primitives based on the boundaries of 3D mesh 502. The surface may include a 2D texture 508 for a user's hand. In a particular embodiment, 2D texture 508 may be associated with a specific color that marks the surface to represent the user's hand. In a particular embodiment, the color of the texture may be specified as an opaque black, meaning that when the texture is "displayed," all light-emitting components should not be illuminated, and when rendering the virtual environment, light (e.g., from virtual objects) should not be allowed to pass through the surface. The surface may also be associated with height map 510, which maps the position of height mesh 504 and any interpolated positions to specific locations on the surface. In a particular embodiment, there may be a direct correspondence between the position of the 2D texture 508 and the height map 510 (e.g., at the per-pixel level), where the height grid 504 may have a coarser resolution due to processing requirements. Frame renderer 506 can perform these steps while generating a surface representation of virtual objects in the virtual environment. Frame renderer 506 can generate the surface representation based on the user's viewpoint of the virtual environment. For example, head-mounted device system 110 may include various sensors that allow head-mounted device system 110 to determine the user's orientation in the virtual environment. Frame renderer 506 can use this orientation information when generating surfaces, where appropriately positioned 2D surfaces can be used to represent 3D objects based on the user's viewpoint. The generated surfaces (including 2D texture 508 and height map 510) can be passed to subframe renderer 512.
[0078] Subframe renderer 512 may be responsible for performing primary visibility determination for the virtual object representation (e.g., a surface) of the user's hand. As described further in detail herein, primary visibility determination may include casting light rays into the virtual scene and determining whether, for any light ray, the surface representation of the user's hand is the first intersecting surface. This indicates that the user's hand occludes the remaining virtual objects in the virtual environment and should be displayed to the user. Subframe renderer 512 is so named because it can generate frames (e.g., images to be displayed) at a higher rate than frame renderer 506. For example, where frame renderer 506 may generate data at, for example, 60 frames per second, subframe renderer 512 may generate data at, for example, 200 frames per second. In performing its primary visibility determination, subframe renderer 512 may use the updated user viewpoint (e.g., updated since receiving data from frame renderer 506) to fine-tune the visibility and positioning of any virtual object (e.g., by modifying the appearance of the corresponding surface). In a particular embodiment, to further reduce latency, a subframe renderer 512 may be incorporated into the head-mounted device system 110 as close as possible to the display system that will eventually output the frame. This is to reduce the latency between user movement (e.g., movement of their head or eyes) and the image shown to the user (containing that movement).
[0079] Figure 6 The process of generating images of a virtual scene for a user according to the embodiments discussed herein is illustrated. Figure 6 Continue to build on Figures 3A to 3B and Figure 5 Above the example shown. First, the camera of the head-mounted device system 110 captures an image of the user's environment. This image includes objects, such as the user's hand, which will occlude virtual objects in the scene. Simultaneously, the head-mounted device system 110 determines a first user pose 600. The first user pose 600 can be determined from one or more captured images (e.g., using SLAM or other positioning techniques). The first user pose 600 can be determined based on one or more onboard sensors (e.g., inertial measurement units) of the head-mounted device system. The captured images can be used, for example, by the head-mounted device system 110 or the body-wearable computing system 120 to perform hand tracking and generate a 3D mesh 502 and a height mesh 504. The first user pose 600, the 3D mesh 502, and the height mesh 504 can be passed to a frame renderer 506.
[0080] Frame renderer 506 can generate a surface virtual object representation of the user's hand based on the 3D mesh 502 and height mesh 504 as described above for use in frame 602. The surface may include a 2D opaque texture 508, a height map 510, and other information needed to represent the user's hand in the virtual environment, such as the position of the user's hand in the environment, the boundaries of the surface, etc. Frame renderer 506 can also generate surfaces representing other virtual objects 604 in the virtual environment. Frame renderer 506 can perform the calculations required to generate this information to support the first frame rate (e.g., 60fps). The frame renderer can pass all this information to subframe renderer 512 (e.g., wirelessly if frame renderer 506 is embodied in a wearable computer).
[0081] Subframe renderer 512 can receive virtual object representations of the surfaces of virtual object 604 and the user's hand. Subframe renderer 512 can perform visibility determination of surfaces in virtual scene 606. Subframe renderer 512 can determine a second user pose 608. The second user pose 608 can differ from the first user pose 600 because, even though the first user pose 600 is updated with each generated frame, the user can move slightly. Failure to consider the user's updated pose may significantly increase user discomfort when using the artificial reality system 100. Visibility determination may include: performing ray casting into the virtual environment based on the second user pose, where the origin of each of several rays (e.g., rays 610a, 610b, 610c, and 610d) is based on its position in the display of the head-mounted device system 110 and the second user pose 608. In a particular embodiment, ray casting may be performed similarly to or together with ray casting performed by transform blocks 220A and 220B of display engine 112.
[0082] For each ray used for visibility determination, subframe renderer 512 can project the ray into the virtual environment and determine whether the ray intersects a surface in the virtual environment. In a particular embodiment, depth testing (e.g., determining which surface intersects first) can be performed at the per-surface level. That is, each surface can have a unique height or depth value that allows subframe renderer 512 (or transform blocks 220A and 220B) to quickly identify interactive surfaces. For example, subframe renderer 512 can project ray 610a into the virtual environment and determine that the ray first intersects the surface corresponding to virtual object 304. Subframe renderer 512 can project ray 610b into the virtual environment and determine that the ray intersects virtual object 304 at a point near surface 612 corresponding to the user's hand. Subframe renderer can project ray 610c into the virtual environment and determine that the ray first intersects surface 612 corresponding to the user's hand. Subframe renderer 512 can project ray 610d into the virtual environment and determine that the ray does not intersect any object in the virtual environment.
[0083] Each ray of light projected into the environment can correspond to one or more pixels of an image to be displayed to the user. The pixel corresponding to a ray can be assigned a color value based on the surface where the ray intersects. For example, the pixel associated with rays 610a and 610b can be assigned a color value based on virtual object 304 (e.g., by sampling a texture value associated with the surface). The pixel associated with ray 610c can be assigned a color value based on surface 612. In a particular embodiment, the color value can be specified as opaque (e.g., where no blending occurs, or where light can pass through) and dark or black to instruct the light-emitting components of the display that will ultimately display the rendered image. The pixel associated with ray 610d can be assigned a default color. In a particular embodiment, the default color can be similar to or the same as the value assigned to surface 612. This default color can be selected to allow the user's environment to be visible when no virtual object is to be displayed (e.g., if there is empty space).
[0084] Subframe renderer 512 can prepare an image for display using color value determination for each ray. In a particular embodiment, this may include appropriate steps performed by pixel blocks 230A and 230B of display engine 112 and display blocks 240A and 240B. Subframe renderer 512 can synthesize the determined pixel color values to prepare a subframe 614 for display to a user. The subframe may include a representation of a virtual object 616 that appears to include a contour map of the user's hand. When displayed by the display components of a head-mounted device system, the contour map allows the user's hand to actually appear in the position on surface 612. Therefore, the user will be able to perceive their actual hand interacting with the virtual object 616.
[0085] Figure 7 The process of generating images of a virtual scene for a user according to the embodiments discussed herein is illustrated. Figure 7 Continue to build on Figures 4A to 4B Above the example shown. First, the camera of the head-mounted device system 110 captures an image of the user's environment. This image includes objects, such as the user's hand, which will occlude virtual objects in the scene. Simultaneously, the head-mounted device system 110 determines a first user pose 700. For example, the head-mounted device system 110 or the body-wearable computing system 120 can use the captured image to perform hand tracking and generate a 3D mesh 701 and a corresponding height mesh. The first user pose 700, the 3D mesh 701, and the corresponding height mesh can be passed to the frame renderer 506.
[0086] Frame renderer 506 can generate a surface (virtual object representation) of the user's hand based on the 3D mesh 701 and height mesh as described above for use in frame 702. The surface may include a 2D opaque texture 706, a height map 708, and other information needed to represent the user's hand in the virtual environment, such as the position of the user's hand in the environment, the boundaries of the surface, etc. Frame renderer 506 can also generate surfaces representing other virtual objects 404 in the virtual environment. Frame renderer 506 can perform the calculations required to generate this information to support the first frame rate (e.g., 60fps). The frame renderer can pass all this information to subframe renderer 512 (e.g., wirelessly if frame renderer 506 is incorporated into a wearable computer).
[0087] Subframe renderer 512 may receive virtual object representations for the surfaces of virtual object 404 and the user's hand. Subframe renderer 512 may perform visibility determination of surfaces in virtual scene 716. Subframe renderer 512 may determine a second user pose 710. The second user pose 710 may differ from the first user pose 700 because, even though the first user pose 700 is updated with each generated frame, the user may move slightly. Failure to consider the user's updated pose may significantly increase user discomfort when using artificial reality system 100. Visibility determination may include casting light into the virtual environment (e.g., light rays 720a, 720b, 720c, and 720d) based on the position in the display of head-mounted device system 110 and the second user pose 720, with each of a plurality of light rays originating from the second user pose 710. In a particular embodiment, the light casting may be performed similarly to or together with the light casting performed by transform blocks 220A and 220B of display engine 112.
[0088] For each ray used for visibility determination, subframe renderer 512 can project the ray into the virtual environment and determine whether the ray intersects a surface in the virtual environment. In a particular embodiment, depth testing (e.g., determining which surface intersects first) can be performed at the per-pixel level. That is, each surface in the virtual environment can have a height map or depth map that allows subframe renderer 512 (or transform blocks 220A and 220B) to distinguish interactions and intersections between surfaces on an individual pixel basis. For example, subframe renderer 512 can project ray 720a into the virtual environment and determine that the ray first intersects the surface corresponding to virtual object 404. Subframe renderer 512 can project ray 720b into the virtual environment and determine that the ray does not intersect any object in the virtual environment. For example, subframe renderer 512 can project ray 720c into the virtual environment and determine that the ray first intersects the surface corresponding to virtual object 404. Note that ray 720c intersects at a point along the overlapping region of the surface of virtual object 724, where virtual object 724 overlaps with surface 722 corresponding to the user's hand. The subframe renderer can project a ray 720d into the virtual environment and determine that the ray first intersects with the surface 722 corresponding to the user's hand.
[0089] Each ray of light projected into the environment can correspond to one or more pixels of an image to be displayed to the user. The pixel corresponding to a ray can be assigned a color value based on the surface where the ray intersects. For example, the pixel associated with rays 720a and 720c can be assigned a color value based on virtual object 404 (e.g., by sampling the texture value associated with the surface). The pixel associated with ray 720d can be assigned a color value based on surface 722. In a particular embodiment, the color value can be specified as opaque (e.g., where no mixing occurs, or where light can pass through) and dark or black to instruct the light-emitting components of the display that will ultimately display the rendered image. The pixel associated with ray 720b can be assigned a default color. In a particular embodiment, the default color can be similar to or the same as the value assigned to surface 722. This default color can be selected to allow the user's environment to be visible when no virtual object is to be displayed (e.g., if there is empty space).
[0090] Subframe renderer 512 can prepare an image for display using color value determination for each ray. In a particular embodiment, this may include appropriate steps performed by pixel blocks 230A and 230B of display engine 112 and display blocks 240A and 240B. Subframe renderer 512 can synthesize the determined pixel color values to prepare a subframe 728 for display to a user. The subframe may include a representation of a virtual object 726 that appears to include a contour map of the user's hand. When displayed by the display components of a head-mounted device system, the contour map allows the user's hand to actually appear in the position of the surface 722. Thus, the user will be able to perceive their actual hand interacting with the virtual object 726 and the virtual object. The virtual object 726 appropriately occludes portions of their hand.
[0091] Figures 8A to 8B An example method 800 for providing real-world object occlusion for virtual objects in augmented reality is illustrated. The method may begin at step 810, where at least one camera (e.g., a head-mounted display) of a head-mounted device system 110 captures an image of the environment of the head-mounted device system 110. In a particular embodiment, the image may be a standard full-color image, a monochrome image, a depth-sensing image, or any other suitable type of image. The image may include objects that the artificial reality system 100 determines should be used to occlude certain virtual objects in a virtual or mixed reality view. In the following example, the object is a user's hand, but it can be various suitable objects, as described above.
[0092] In a particular embodiment, steps 815 through 840 may involve generating frames to be displayed to the user based on the position of the user's hand and virtual objects in the virtual environment. In a particular embodiment, one or more of steps 815 through 840 may be performed by the head-mounted device system 110, the body-wearable computing system 120, or another suitable computing system having greater computing resources than the head-mounted device system 120 and communicatively coupled to the head-mounted device system 120. In a particular embodiment, the work performed in each step may be distributed among qualified computing systems by the work controller of the artificial reality system 100.
[0093] The method may continue at step 815, where the computing system can determine the user's viewpoint of the head-mounted device system to the environment. In a particular embodiment, the viewpoint may be determined entirely from captured images of the environment using suitable positioning and / or mapping techniques. In a particular embodiment, the viewpoint may be determined using data retrieved from other sensors of the head-mounted device system.
[0094] In step 820, the computing system can detect the user's hand pose. To detect the position of the user's hand, the computing system can first identify portions of the captured image that include the user's hand. The computing system can execute one of several algorithms to identify the presence of the user's hand in the captured image. After confirming that the hand is actually present in the image, the computing system can perform hand tracking analysis to identify the presence and position of several discrete locations on the user's hand in the captured image. In a particular embodiment, a depth tracking camera can be used to facilitate hand tracking. In a particular embodiment, hand tracking can be performed without using standard depth tracking with deep learning and model-based tracking. Deep neural networks can be trained and used to predict the position of a person's hand as well as landmarks, such as the joints and fingertips of the hand. Landmarks can be used to reconstruct high-DOF poses of the hand and fingers (e.g., 26-DOF poses). This pose can provide the position of the hand relative to the user's viewpoint in the environment.
[0095] In step 825, the computing system can generate a 3D mesh for the user's hand based on the detected pose. Using the captured image, the computing system can prepare a model of the user's hand to illustrate the detected pose. In a particular embodiment, the 3D mesh can be a fully modeled virtual object of the user's hand, allowing the user's hand to be adequately represented in the virtual scene. From the 3D mesh, the computing system can generate a height map of the user's hand in the environment. For example, the computing system can determine the depth of a portion of the user's hand (e.g., the distance of the user's hand from the viewpoint and / or camera). To facilitate accurate occlusion of the virtual object, the position of the user's hand must be known on a more granular basis. Accurately calculating depth based solely on an image can be computationally very expensive. The 3D mesh can be used to fill gaps because the outline of the user's hand can be determined from the 3D mesh, and a height map or a data structure storing various heights of the user's hand can be generated.
[0096] In step 830, the computing system can determine the user's viewpoint to the virtual environment surrounding the user. The virtual environment can be controlled by an artificial reality application executed by the artificial reality system 100. The virtual environment may include multiple virtual objects within it. Each virtual object can be associated with information describing its size, shape, texture, absolute position (e.g., relative to a fixed point), and relative position (e.g., relative to the viewpoint). The viewpoint in the virtual environment can be determined based on a correspondence between the user's viewpoint in the real environment and the user's viewpoint in the virtual environment. This correspondence can be determined by the artificial reality system 100 or can be specified by the user (e.g., calibrated).
[0097] In step 835, the computing system can generate a virtual object representation of the user's hand based on a 3D mesh. The virtual object representation (also referred to as a surface throughout this disclosure) can be associated with a texture and an associated height map. In a particular embodiment, the virtual object representation can be a 2D representation of a 3D mesh of the user's hand as observed from a viewpoint in a virtual environment. A 2D representation can be sufficient to represent a 3D mesh because the computing system can generate frames at a rate fast enough that the user will not be able to detect the generation of only a 2D representation (e.g., 60 Hz or higher).
[0098] In step 840, the computing system can generate representations of other virtual objects in the virtual environment based on models of virtual objects and the virtual environment, as well as the user's viewpoint on the virtual environment. Similar to the virtual object representation of the user's hand, these virtual object representations can be associated with the textures and height maps of the virtual objects, which can be used to quickly simulate 3D models, for example, when determining intersections between virtual objects. In summary, the virtual object representations constitute the data for frames used to render the virtual environment. In this context, a frame refers to the rate at which virtual object representations are created. This rate can be determined based on the probability of a particular degree of movement of the user's viewpoint.
[0099] This method can proceed to Figure 8B Step 850 is shown. Steps 850 through 890 may involve generating and providing subframes for displaying a virtual environment. Subframes may use the generated and produced data as frames of the virtual environment, updating the data as needed, and actually determining the image to be displayed to the user by the display components (e.g., an array of light-emitting components) of the head-mounted device system 110. Subframes may be prepared and displayed at a rate much higher than that of the prepared frames (e.g., 200 Hz or higher). Due to the high rate of preparing subframes, it may be necessary to limit inter-system communication. Therefore, the computations required for steps 850 through 890 may preferably be performed by the head-mounted device system 110 itself, rather than by the body-wearable computing system 120 or other computing systems communicatively coupled to the head-mounted device system 110.
[0100] In step 850, the computing system can determine the user's updated viewpoint to the environment and the virtual environment around the user. An updated viewpoint may be needed to allow for the preparation of subframes that take into account the user's minute movements (e.g., head movements, eye movements, etc.). The updated viewpoint can be determined in the same manner as described above.
[0101] In step 855, the computing system can perform a primary visibility determination regarding the virtual object representation of the user's hand and the user's viewpoint of the virtual object representation to the virtual environment. Specifically, the visibility determination can be performed using ray casting techniques. The computing system can organize several rays to be cast into the virtual environment. Each ray can correspond to one or more pixels of an image to be displayed by the head-mounted device system 110. Each pixel can correspond to one or more display positions of the light emitter of the head-mounted device system 110. The origin of each ray can be based on the user's viewpoint to the virtual environment and the position of the corresponding pixel. The direction of the ray can be determined similarly based on the user's viewpoint to the virtual environment. Casting rays into the virtual environment can be performed to simulate the behavior of light in the virtual environment. Generally, casting rays constitutes determining whether the ray intersects with a virtual object in the virtual environment and assigning a color to the pixel corresponding to the intersecting ray.
[0102] The process for each ray has been partially described for steps 860 through 880. In step 860, the computational system determines whether each ray intersects with a virtual object in the virtual environment. In a particular embodiment, the computational system can check for intersection by simulating the path of the ray entering the virtual environment over a fixed distance or time period. At each step of the path, the ray can compare its position with the position of various virtual object representations already generated for a particular frame. The position of the virtual object is known based on the location specified for the representation, including depth and a height map that can be associated with the representation. Using this ray casting technique, it is assumed that the correct color value for the pixel corresponding to the ray is the color value associated with the first virtual object it intersects with. If the computational system determines that the ray does not intersect with a virtual object, the method can proceed directly to step 875, where, as further described below, the color value associated with the ray (and subsequently the corresponding pixel) is set to a dedicated pass-through color, where the emitter remains unilluminated to allow light to pass directly from the environment to the user. If the computational system determines that the ray does indeed intersect with a virtual object, the method can proceed to step 865.
[0103] In step 865, the computing system can identify intersecting virtual object representations and sample the color of the corresponding texture at the intersection points. To sample the corresponding color, the computing system can first retrieve the texture associated with the intersecting virtual object representations. The computing system can then determine the points on the texture corresponding to the intersection. This may include converting the intersection points from global or view-based coordinates to texture-based coordinates. Once the appropriate coordinates are determined, the computing system can identify the color value stored at the appropriate location in the texture. In a particular embodiment, the artificial reality system 100 may use shortcuts to reduce the memory access time required to perform this operation. Such shortcuts may include tagging specific virtual object representations.
[0104] In step 870, the calculation system can determine the value of the sampled color. If the sampled color is opaque black (or specified as another color to indicate that the corresponding light emitter should remain unilluminated), the method can proceed to step 875. Otherwise, the method can proceed to step 880.
[0105] In step 875, the calculation system can set the color value of the pixel corresponding to the light source to a pass-through color. The pass-through color can be used to generate or otherwise associate with an instruction to the light emitter indicating that the light emitter should not illuminate when the final image is displayed to the user as a subframe. Example pass-through colors could include black (indicating the corresponding light emitter should be dark) or transparent (indicating that light should not be generated by the light emitter).
[0106] In step 880, the computing system may set a color value based on the color sampled from the texture at the intersection. In a particular embodiment, this may include scheduling color interpolation at multiple locations (e.g., when the intersection is between discrete texel positions of the texture). In a particular embodiment, setting the color may include performing color adjustment and brightness correction (as described above with respect to display blocks 240A and 240B of display engine 112).
[0107] In step 885, the computing system can generate an image to be displayed as a subframe based on the color values and corresponding pixel positions determined for each ray. This image can be called a composite image because it combines determined color values for certain locations (e.g., the image positions of virtual objects visible from the user's viewpoint of the virtual environment) and indicates that light should absolutely not be generated at other locations (e.g., locations where rays do not intersect with virtual object representations, or where rays first intersect with virtual object representations of the user's hand) (e.g., light emitters should not be illuminating). Therefore, by viewing only the composite image, the viewer will see rendered virtual objects and completely blank areas.
[0108] In step 890, the computing system can provide subframes for display. The computing system can generate instructions based on a composite image from the light emitters of the display for the head-mounted device system 110. The instructions can include color values, color brightness or intensity, and any variables that affect the display of the subframes. The instructions can also include instructions for certain light emitters to remain unlit or turned off when needed. As described throughout, the desired effect is that the user can see portions of the virtual environment (e.g., virtual objects in the virtual environment) from a determined viewpoint of the virtual environment, which are blended with portions of the real environment including the user's hands (which may be interacting with virtual objects in the virtual environment).
[0109] Where appropriate, certain embodiments may be repeated. Figures 8A to 8BOne or more steps of the method. Although this disclosure will Figures 8A to 8B The specific steps of the method are described and shown as occurring in a specific order, but this disclosure contemplates... Figures 8A to 8B Any suitable steps of the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates example methods for providing real-world object occlusion for virtual objects in augmented reality, including... Figures 8A to 8B This disclosure considers specific steps of the method, but also any suitable method for providing real-world object occlusion of virtual objects in augmented reality, including any suitable steps that may, where appropriate, include... Figures 8A to 8B The method may involve all, some, or none of the steps. Furthermore, although this disclosure describes and illustrates specific components, devices, or systems performing... Figures 8A to 8B This disclosure describes specific steps of the method, but it considers any suitable combination of any suitable components, devices, or systems to perform the procedure. Figures 8A to 8B Any appropriate steps of the method.
[0110] Figure 9 An example computer system 900 is illustrated. In a particular embodiment, one or more computer systems 900 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 900 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 900 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular embodiments include one or more portions of one or more computer systems 900. Herein, references to computer systems may cover computing devices and vice versa, where appropriate. Furthermore, references to computer systems may cover one or more computer systems, where appropriate.
[0111] This disclosure contemplates any suitable number of computer systems 900. This disclosure contemplates computer systems 900 in any suitable physical form. By way of example and not limitation, computer system 900 may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system grid, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of the foregoing. Where appropriate, computer system 900 may include one or more computer systems 900; may be single or distributed; may span multiple locations; may span multiple machines; may span multiple data centers; or may reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 900 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 900 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 900 may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.
[0112] In a particular embodiment, computer system 900 includes processor 902, memory 904, storage device 906, input / output (I / O) interface 908, communication interface 910, and bus 912. While this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of components in any suitable arrangement.
[0113] In a particular embodiment, processor 902 includes hardware for executing instructions, such as those constituting a computer program. By way of example, and not limitation, to execute instructions, processor 902 may retrieve (or fetch) instructions from internal registers, internal caches, memory 904, or storage device 906; decode and execute the instructions; and then write one or more results to internal registers, internal caches, memory 904, or storage device 906. In a particular embodiment, processor 902 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates that processor 902 may include any suitable number of suitable internal caches. By way of example, and not limitation, processor 902 may include one or more instruction caches, one or more data caches, and one or more translation back buffers (TLBs). Instructions in the instruction cache may be copies of instructions in memory 904 or storage device 906, and the instruction cache may accelerate the retrieval of those instructions by processor 902. The data in the data cache may be a copy of the data in memory 904 or storage device 906 for instructions executed at processor 902 to perform operations; the result of a subsequent instruction executed at processor 902 for writing to memory 904 or storage device 906; or other suitable data. The data cache can accelerate read or write operations of processor 902. The TLB can accelerate virtual address translation of processor 902. In a particular embodiment, processor 902 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates that processor 902 may include any suitable number of suitable internal registers. Where appropriate, processor 902 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 902. Although this disclosure describes and illustrates particular processors, this disclosure contemplates any suitable processor.
[0114] In a particular embodiment, memory 904 includes main memory for storing instructions to be executed by processor 902 or data to be operated by processor 902. By way of example and not limitation, computer system 900 may load instructions from storage device 906 or another source (e.g., another computer system 900) into memory 904. Processor 902 may then load instructions from memory 904 into internal registers or internal caches. To execute instructions, processor 902 may retrieve instructions from internal registers or internal caches and decode them. During or after instruction execution, processor 902 may write one or more results (which may be intermediate or final results) to internal registers or internal caches. Processor 902 may then write one or more of these results to memory 904. In a particular embodiment, processor 902 executes only the instructions in one or more internal registers or internal caches or memory 904 (as opposed to storage device 906 or elsewhere) and operates only on the data in one or more internal registers or internal caches or memory 904 (as opposed to storage device 906 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may couple processor 902 to memory 904. Bus 912 may include one or more memory buses, as described below. In a particular embodiment, one or more memory management units (MMUs) reside between processor 902 and memory 904 and facilitate access to memory 904 requested by processor 902. In a particular embodiment, memory 904 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 904 may include one or more memories 904. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memory.
[0115] In a particular embodiment, storage device 906 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 906 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage device 906 may include removable or non-removable (or fixed) media. Where appropriate, storage device 906 may be internal or external to computer system 900. In a particular embodiment, storage device 906 is a non-volatile solid-state memory. In a particular embodiment, storage device 906 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically changeable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage device 906 in any suitable physical form. Where appropriate, storage device 906 may include one or more storage control units that facilitate communication between processor 902 and storage device 906. Where appropriate, storage device 906 may include one or more storage devices 906. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.
[0116] In a particular embodiment, I / O interface 908 includes hardware, software, or both for providing one or more interfaces for communication between computer system 900 and one or more I / O devices. Where appropriate, computer system 900 may include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 900. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, display, mouse, printer, scanner, speaker, camera, stylus, tablet computer, touchscreen, trackball, video camera, other suitable I / O devices, or combinations of two or more of these. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 908 for the I / O device. Where appropriate, I / O interface 908 may include one or more device or software drivers to enable processor 902 to drive one or more of these I / O devices. Where appropriate, I / O interface 908 may include one or more I / O interfaces 908. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure contemplates any suitable I / O interface.
[0117] In a particular embodiment, communication interface 910 includes hardware, software, or both for providing one or more interfaces for communication (injection, packet-based communication) between computer system 900 and one or more other computer systems 900 or one or more networks. By way of example, and not limitation, communication interface 910 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks, such as Wi-Fi networks. This disclosure contemplates any suitable network and any suitable communication interface 910 for use therewith. By way of example, and not limitation, computer system 900 may communicate with one or more portions of an ad hoc network, personal area network (PAN), local area network (LAN), wide area network (WAN), metropolitan area network (MAN), or the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 900 may communicate with a wireless PAN (WPAN) (e.g., BLUETOOTH WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), or other suitable wireless networks, or a combination of two or more of these. Where appropriate, computer system 900 may include any suitable communication interface 910 for any of these networks. Where appropriate, communication interface 910 may include one or more communication interfaces 910. Although specific communication interfaces are described and shown in this disclosure, any suitable communication interface is contemplated in this disclosure.
[0118] In a particular embodiment, bus 912 includes hardware, software, or both for coupling components of computer system 900 to each other. By way of example and not limitation, bus 912 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 912 may include one or more buses 912. Although this disclosure describes and illustrates specific buses, this disclosure contemplates any suitable bus or interconnect.
[0119] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0120] In this document, "or" is inclusive rather than exclusive, unless otherwise expressly stated or the context otherwise indicates. Therefore, in this document, "A or B" means "A, B, or both," unless otherwise expressly stated or the context otherwise indicates. Furthermore, "and" is both joint and multiple, unless otherwise expressly stated or the context otherwise indicates. Therefore, in this document, "A and B" means "A and B, jointly or separately," unless otherwise expressly stated or the context otherwise indicates.
[0121] The scope of this disclosure includes all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates various embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any components, elements, features, functions, operations, or steps described or illustrated anywhere herein that will be understood by those skilled in the art. Furthermore, references in the appended claims to a means or system or a component of a means or system adapted to, arranged to, capable of, configured to, enabled, operable to, or operated to perform a particular function cover that means, system, or component, whether or not it or the particular function is activated, turned on, or unlocked, provided that the means, system, or component is so adapted, arranged, capable of, configured to, enabled, operable to, or operated. Moreover, although this disclosure describes or illustrates specific embodiments to provide particular advantages, specific embodiments may not provide, provide some, or all of these advantages.
Claims
1. A method for generating graphics for an artificial reality environment, comprising by one or more computing devices: accessing an image comprising a hand of a user of a head-mounted display; generating a planar virtual object representation of the hand of the user from at least the image comprising the hand of the user, the planar virtual object representation of the hand being defined in a virtual environment comprising at least one other virtual object; determining that a portion of the planar virtual object representation of the hand is less than a distance of the other virtual object from a viewpoint of the user; rendering a rendered image of the virtual environment from the viewpoint of the user based on the planar virtual object representation of the hand and the at least one other virtual object in the virtual environment, the rendered image comprising a set of pixels corresponding to the portion of the planar virtual object representation of the hand that is less than the distance of the other virtual object from the viewpoint of the user; and providing instructions to a set of light emitters of the head-mounted display for displaying the rendered image of the virtual environment, wherein the set of pixels in the rendered image corresponding to the portion of the planar virtual object representation of the hand causes light emitters at one or more locations to be unilluminated.
2. The method of claim 1, wherein a light emitter at a particular location is unilluminated, causing light from the user's environment to continue to the user at the particular location, or wherein the image comprising the hand of the user further comprises the user's environment from the viewpoint of the user.
3. The method of claim 1, wherein generating a planar virtual object representation of the hand of the user from at least the image comprises: determining a position of the planar virtual object representation of the hand of the user in the virtual environment based on a pose of the hand of the user determined from the image comprising the hand of the user.
4. The method of claim 1, wherein a texture of the planar virtual object representation of the hand of the user corresponds to instructions to cause light emitters to be unilluminated, or wherein the planar virtual object representation of the hand of the user is associated with a color that is further associated with a background of the virtual environment.
5. The method of claim 1, wherein rendering the image of the virtual environment from the viewpoint of the user comprises: determining whether the planar virtual object representation of the hand of the user and at least one other virtual object in the virtual environment are visible from the viewpoint of the user, and optionally, wherein the method further comprises determining that the planar virtual object representation of the hand is at least partially in front of the at least one other virtual object in the virtual environment by: casting a ray into the virtual environment having an origin and a direction based on the viewpoint of the user; and determining a point of intersection of the ray with the planar virtual object representation of the hand of the user in the virtual environment, wherein the ray intersects the planar virtual object representation before intersecting another object in the virtual environment.
6. The method of claim 1, wherein generating a planar virtual object representation of the hand of the user from at least the image comprises: determining a pose of the hand from at least the image; generating a triangular mesh corresponding to the hand from at least the image and the pose; determining a distance of the hand from a viewpoint of the user from at least the image; and generating a height map based on the triangular mesh corresponding to the hand, the height map indicating a change in one or more positions of the hand as a function of the determined distance of the hand.
7. The method of claim 6, wherein determining that the distance of the portion of the planar virtual object representation of the hand from the viewpoint of the user is less than the distance of the other virtual object from the viewpoint of the user comprises: at a particular position, comparing the distance of the hand and the height map associated with the planar virtual object representation of the hand to a distance and height map associated with the other virtual object; and determining that the planar virtual object representation of the hand is the object closest to the viewpoint based on the comparison.
8. The method of claim 1, wherein the instructions for displaying the image of the virtual environment further cause a light emitter to emit light at one or more positions at which a portion of the at least one other virtual object is in front of the planar virtual object representation of the hand of the user from the viewpoint of the user.
9. The method of claim 1, wherein one or more of the computing devices are embodied in the head-mounted display and one or more of the computing devices are separate computing devices, and optionally wherein the method further comprises: allocating steps of the method between the computing devices of the head-mounted display and the separate computing devices based on one or more metrics of available computing resources.
10. The method of claim 1, wherein the image is generated by a first camera of the head-mounted display; and wherein generating the planar virtual object representation of the hand of the user comprises: accessing a second image generated by a second camera of the head-mounted display; and positioning the hand of the user relative to the viewpoint of the user based on the image and the second image.
11. The method of claim 10, wherein generating the planar virtual object representation of the hand of the user further comprises: generating an array of positions having positions corresponding to positions of the hand of the user; storing values of distances between one or more positions of the hand and the viewpoint of the user in the array; and associating the array with the planar virtual object representation of the hand of the user, or wherein the second camera of the head-mounted display is a depth-sensing camera.
12. One or more computer-readable non-transitory storage media embodying software that is operable when executed to perform a method according to any of claims 1-11, or: accessing an image comprising a hand of a user of a head-mounted display; generating a planar virtual object representation of the hand of the user from at least the image comprising the hand of the user, the planar virtual object representation of the hand being defined in a virtual environment comprising at least one other virtual object; rendering a rendered image of the virtual environment from a viewpoint of the user based on the planar virtual object representation of the hand and the at least one other virtual object in the virtual environment, the rendered image comprising a set of pixels corresponding to the portion of the planar virtual object representation of the hand that is less than the distance of the other virtual object from the viewpoint of the user; and providing instructions to a set of light emitters of the head-mounted display for displaying the rendered image of the virtual environment, wherein the set of pixels in the rendered image corresponding to the portion of the planar virtual object representation of the hand causes the light emitters at one or more locations to be unilluminated.
13. The computer-readable non-transitory storage media of claim 12, wherein a light emitter at a particular location is unilluminated such that light from the user's environment continues to the user at the particular location.
14. A system for generating graphics for an artificial reality environment, comprising: one or more processors; and one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions that are operable when executed by one or more of the processors to cause the system to perform a method according to any of claims 1-11, or: accessing an image comprising a hand of a user of a head-mounted display; generating a planar virtual object representation of the hand of the user from at least the image comprising the hand of the user, the planar virtual object representation of the hand being defined in a virtual environment comprising at least one other virtual object; determining that a portion of the planar virtual object representation of the hand is less than a distance of the other virtual object from a viewpoint of the user; rendering a rendered image of the virtual environment from a viewpoint of the user based on the planar virtual object representation of the hand and the at least one other virtual object in the virtual environment, the rendered image comprising a set of pixels corresponding to the portion of the planar virtual object representation of the hand that is less than the distance of the other virtual object from the viewpoint of the user; and providing instructions to a set of light emitters of the head-mounted display for displaying the rendered image of the virtual environment, wherein the set of pixels in the rendered image corresponding to the portion of the planar virtual object representation of the hand causes the light emitters at one or more locations to be unilluminated. providing instructions to a set of light emitters of the head-mounted display to display the rendered image of the virtual environment, wherein the set of pixels of the rendered image corresponding to the portion of the planar virtual object representation of the hand causes one or more light emitters at a particular location to not emit light.
15. The system of claim 14, wherein a light emitter at a particular location does not emit light, causing light from the user’s environment to continue to the user at the particular location.
Citation Information
Patent Citations
Interpolation optimizations for a display engine for post-rendering processing
US11138747B1
Display engine for post-rendering processing
US11403810B2
Generating and Modifying Representations of Objects in an Augmented-Reality or Virtual-Reality Scene
US20200134923A1
Realistic occlusion for a head mounted augmented reality display
CN103472909A
Apparatus and method for estimating hand position utilizing head mounted color depth camera, and bare hand interaction system using same
US20170140552A1