Fast Hand Meshing for Dynamic Occlusion
By detecting and calculating the hand model, the hand grid is updated in real time, the problem of inaccurate hand occlusion processing in the existing technology is solved, and the fast and accurate dynamic occlusion effect in the cross-realistic system is achieved, which improves the authenticity of the user experience.
Patent Information
- Application Number
- CN202080046371.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-06-25
AI Technical Summary
When existing cross-reality systems deal with dynamic occlusion, especially hand occlusion, it is difficult to achieve fast and accurate meshing, resulting in unreal rendering of virtual objects.
By receiving data queries related to the hands in the scene, the sensors in the device worn by the user obtain the scene information, detect and calculate the hand model, and based on the model masking depth information, the hand grid is updated in real time, and provided to the application to render the virtual object part that is not blocked by the hands.
It realizes fast and accurate hand grid division, improves dynamic occlusion processing capabilities in cross-reality systems, and enhances the authenticity of user experience.
Smart Images

Figure CN114026606B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to cross-reality systems that use 3D world reconstruction to render scenes. Background Art
[0002] A computer can control a human user interface to create an X Reality (XR) environment in which some or all of the XR environment perceived by the user is generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, and some or all of these XR environments can be generated by the computer using in part data that describes the environment. For example, the data can describe virtual objects that can be rendered in a manner that the user senses or perceives as part of the physical world and with which the user can interact. Since the data is rendered and presented via a user interface device such as, for example, a head-mounted display device, the user can experience these virtual objects. The data can be shown to the user, or it can control audio that is played to the user, or it can control a haptic (or tactile) interface so that the user can experience the sensation of touching a virtual object that the user senses or perceives as being felt.
[0003] XR systems can be used for many applications across the fields of scientific visualization, medical training, engineering design and prototyping, telemanipulation and telepresence, and personal entertainment. Compared to VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems and also opens the door to various applications for presenting information about how to change the reality of the physical world in an easy-to-understand manner.
[0004] An XR system can represent the physical world around the user of the system as a "mesh". The mesh can be represented by a plurality of interconnected triangles. Each triangle has sides that connect points on the surface of an object within the physical world such that each triangle represents a portion of that surface. Information about that portion of the surface, such as color, texture, or other properties, can be stored in the triangle in an associated manner. In operation, the XR system can process image information to detect points and surfaces and thereby create or update the mesh. Summary of the Invention
[0005] Aspects of this application relate to methods and devices for fast hand meshing for dynamic occlusion. The techniques described herein can be used together, used separately, or used in any suitable combination.
[0006] Some embodiments relate to a method of operating a computing system to reconstruct a hand for dynamically occluding virtual objects. The method includes: receiving, from an application that renders virtual objects in a scene, a query for data related to a hand in the scene; obtaining information about the scene from a device worn by a user, the device including one or more sensors, the information about the scene including depth information indicative of a distance between the device worn by the user and a physical object in the scene; detecting whether the physical object in the scene includes a hand; when the hand is detected, calculating a model of the hand based at least in part on the information about the scene; masking the depth information indicative of the distance between the device worn by the user and the physical object in the scene with the model of the hand; calculating a hand mesh based on the depth information masked to the model of the hand, the calculating including: updating the hand mesh in real time as a relative position between the device and the hand changes; and providing the hand mesh to the application such that the application renders portions of the virtual object not occluded by the hand mesh.
[0007] In some embodiments, the model of the hand includes a plurality of key points of the hand indicative of points on segments of the hand.
[0008] In some embodiments, at least a portion of the plurality of key points of the hand corresponds to joints of the hand and fingertips of the hand.
[0009] In some embodiments, the method further includes: determining a contour of the hand based on the plurality of key points; and masking the depth information indicative of the distance between the device worn by the user and the physical object in the scene with the model of the hand. Masking the depth information includes: filtering out the depth information outside the contour of the model of the hand; and generating, at least in part based on the filtered depth information, a depth image of the hand, the depth image including a plurality of pixels, each pixel indicative of a distance to a point on the hand.
[0010] In some embodiments, filtering out the depth information outside the contour of the model of the hand includes: removing depth information associated with the physical object in the scene.
[0011] In some embodiments, masking the depth information indicative of the distance between the device worn by the user and the physical object in the scene with the model of the hand includes: associating portions of the depth image to hand segments; and updating the hand mesh in real time includes: selectively updating portions of the hand mesh representing a suitable subset of the hand segments.
[0012] In some embodiments, the method further includes: filling holes in the depth image before calculating the hand mesh.
[0013] In some embodiments, filling holes in the depth image includes: generating stereo depth information from a stereo camera of the device, the stereo depth information corresponding to a region of the holes in the depth image.
[0014] In some embodiments, filling holes in the depth image includes: accessing surface information from a 3D model of the hand, the surface information corresponding to a region of the holes in the depth image.
[0015] In some embodiments, calculating the hand mesh based on the depth information of the model masked to the hand includes: predicting a latency n according to a query for data related to the hand in the scene received from the application that renders the virtual object in the scene at time t; predicting a hand pose at a time of the query time t plus the latency n; and deforming the hand mesh with the predicted pose at the time of the query time t plus the latency n.
[0016] In some embodiments, the depth information indicating a distance between the device worn by the user and the physical object in the scene includes a sequence of depth images at a frame rate of at least 30 frames per second.
[0017] Some embodiments relate to an electronic system that can be carried by a user. The electronic system includes a device worn by the user. The device includes a display configured to render a virtual object and includes one or more sensors configured to capture a head pose of the user wearing the device and information of a scene including one or more physical objects, the information of the scene including depth information indicating a distance between the device and the one or more physical objects. The electronic system includes: a hand meshing component configured to execute computer-executable instructions to detect a hand in the scene, calculate a hand mesh of the detected hand, and update the hand mesh in real time as the head pose changes and / or the hand moves; and an application configured to execute computer-executable instructions to render the virtual object in the scene, wherein the application receives the hand mesh and a portion of the virtual object occluded by the hand from the hand meshing component.
[0018] In some embodiments, the hand meshing component is configured to calculate a hand mesh by: identifying key points on the hand; calculating segments between the key points; selecting information from the depth information based on proximity to one or more of the calculated segments; and calculating a mesh representing at least a portion of the hand mesh based on the selected depth information.
[0019] In some embodiments, the depth information includes a plurality of pixels, each pixel in the plurality of pixels representing a distance to an object in the scene. Calculating the mesh includes: grouping adjacent pixels representing a distance difference less than a threshold.
[0020] Some embodiments relate to a method of operating an AR system to render a virtual object in a scene including a physical object. The AR system includes at least one sensor and at least one processor. The method includes: capturing information of the scene with the at least one sensor, the information of the scene including depth information indicating a distance to a physical object in the scene; processing, with at least one processor, the captured information to detect a hand in the scene and calculate points on the hand; selecting a subset of the depth information based on proximity to the calculated points on the hand; and calculating a representation of the hand based on the selected depth information, wherein the representation of the hand indicates a surface of the hand.
[0021] In some embodiments, the method further includes: storing the calculated representation of the hand; and continuously processing the captured information to update the stored representation of the hand.
[0022] In some embodiments, calculating the representation of the hand includes: calculating one or more motion parameters of the hand based on the captured information; extrapolating a position of the hand at a future time based on the one or more motion parameters, the future time being determined based on a latency associated with rendering the virtual object using the calculated representation of the hand; and deforming the calculated representation of the hand to represent the hand at the extrapolated position.
[0023] In some embodiments, the method further includes: rendering a selected portion of the virtual object based on the representation of the hand, wherein the selected portion represents a portion of the virtual object not occluded by the hand.
[0024] In some embodiments, the depth information includes a depth map including a plurality of pixels, each pixel representing a distance. Calculating the representation of the hand based on the selected depth information includes: identifying groups of pixels representing surface segments.
[0025] In some embodiments, calculating the representation of the hand includes: defining a mesh representing the hand based on the identified groups of pixels.
[0026] In some embodiments, defining the mesh includes: identifying triangular regions corresponding to the identified surface segments.
[0027] The foregoing summary is provided by way of illustration and not by way of limitation. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings are not necessarily to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For clarity, not every component may be labeled in every drawing. In the drawings:
[0029] Figure 1 is a schematic diagram showing an example of a simplified augmented reality (AR) scene according to some embodiments.
[0030] Figure 2 is a schematic diagram of an exemplary simplified AR scene according to some embodiments, showing an exemplary world reconstruction use case, including visual occlusion, physics-based interaction, and environmental reasoning.
[0031] Figure 3 is a schematic diagram showing the data flow in an AR system configured to provide an experience of interacting with the physical world with AR content according to some embodiments.
[0032] Figure 4 is a schematic diagram showing an example of an AR display system according to some embodiments.
[0033] Figure 5A is a schematic diagram showing that when a user wears an AR display system, the AR display system renders AR content as the user moves through a physical world environment according to some embodiments.
[0034] Figure 5B is a schematic diagram showing a viewing optical component and accessory components according to some embodiments.
[0035] Figure 6 is a schematic diagram showing an AR system using a world reconstruction system according to some embodiments.
[0036] Figure 7 is a schematic diagram showing an AR system configured to generate a hand mesh for dynamic occlusion in real time according to some embodiments.
[0037] Figure 8 is a flowchart showing a method for generating a hand mesh for dynamic occlusion in real time according to some embodiments.
[0038] Figure 9A is an exemplary image captured by a sensor corresponding to one eye according to some embodiments.
[0039] Figure 9B are two exemplary images captured by two sensors corresponding to the left eye and the right eye according to some embodiments.
[0040] Figure 9C is an exemplary depth image according to some embodiments, which can be at least partially obtained from Figure 9A the image of Figure 9B or the image of
[0041] Figure 9D is an exemplary image showing the contour of a model of a Figure 8 hand according to some embodiments.
[0042] Figure 9E is a schematic diagram showing an exemplary eight-key-point model of a hand according to some embodiments.
[0043] Figure 9F is a schematic diagram showing an exemplary twenty-two-key-point model of a hand according to some embodiments.
[0044] Figure 9G is a schematic diagram showing a dense hand mesh according to some embodiments.
[0045] Figure 10 is a flowchart showing the details of masking depth information using a Figure 8 hand model according to some embodiments.
[0046] Figure 11 is a flowchart showing the details of real-time computing a hand mesh based on masked depth information of a Figure 8 hand segmentation according to some embodiments. DETAILED DESCRIPTION
[0047] Methods and apparatuses for fast hand meshing for dynamic occlusion in an X Reality (XR) system are described herein. The XR system can create and use three-dimensional (3D) world reconstruction. To provide a realistic XR experience to a user, the XR system must know the user's physical environment in order to correctly associate the positions of virtual objects relative to real objects. World reconstruction can be built from images and depth information about those physical environments, which are collected using sensors that are part of the XR system. The world reconstruction can then be used by any of the multiple components of such a system. For example, the world reconstruction can be used by components that perform visual occlusion processing, compute physics-based interactions, or perform environmental reasoning.
[0048] Occlusion handling identifies portions of virtual objects that should not be rendered and / or displayed to the user because there are objects in the physical world blocking the user's view of where the virtual objects would be perceived by the user. Physics-based interactions are computed to determine where or how the virtual objects appear to the user. For example, virtual objects can be rendered to appear to rest on physical objects, move through empty space, or collide with the surfaces of physical objects. World reconstruction provides a model from which information about the objects in the physical world can be obtained for such computations.
[0049] There are significant challenges in providing such a system. A large amount of processing may be required to compute world reconstruction and occlusion information. Additionally, the XR system must correctly know how to position virtual objects relative to the user's head, body, etc. As the user's position changes relative to the physical environment, the relevant parts of the physical world also change, which may require further processing. Moreover, as objects move in the physical world (e.g., a cup moves on a table), the 3D reconstruction data typically needs to be updated. Updates to the data representing the environment the user is experiencing must be performed quickly without using a large amount of computing resources of the computer generating the XR environment because other functions cannot be performed while world reconstruction is being executed. Additionally, components that "consume" data exacerbate the demand for computer resources in processing the reconstruction data.
[0050] Dynamic occlusion handling identifies portions of virtual objects that should not be rendered and / or shown to the user because there are physical objects blocking the user's view of where the virtual objects would be perceived by the user, and the relative positions between the physical objects and the virtual objects change over time. Occlusion handling for the user's hand is particularly important for providing an ideal XR experience. However, the inventors have recognized and realized that improved occlusion handling specifically for the hand can provide a more realistic XR experience for the user. For example, an XR system may generate a mesh of an object for use in occlusion handling based on graphic images captured at a frame rate of about 5 frames per second (fps). However, due to hand movement and / or head movement, such as higher than 15 fps, higher than 30 fps, or higher than 45 fps, this rate may not be sufficient to keep up with the speed of position changes between the hand and the virtual objects behind the hand.
[0051] Users of XR devices can interact with the device by making hand gestures. The hands are crucial for latency because they are directly used for user interaction. During interaction with the device, the user's hands can move quickly, e.g., faster than the user moves to scan the physical environment to reconstruct the world. Further, the user's hands are closer to the XR device worn by the user. Thus, the relative position between the user's hands and the virtual objects behind the hands is also sensitive to head movement. If the representation of the hands for occlusion processing does not update fast enough to keep up with these relative motion sources, the occlusion processing will not be based on the position of the hands and the occlusion processing will be inaccurate. If the virtual objects behind the hands are not correctly rendered as occluded by the hands during hand movement and / or head movement, the XR scene will appear unrealistic to the user. The virtual objects may appear on top of the hands as if the hands were transparent. Otherwise, the virtual objects may not appear in their expected positions. The hands may appear to have the color pattern of the virtual objects, or other artifacts may occur. Thus, the movement of the hands disrupts the user's immersion in the XR experience.
[0052] The inventors have recognized and realized that when the object occluding the virtual object is the user's hand, particularly high computational requirements may be needed. However, this computational burden can be alleviated by techniques for generating hand occlusion data with low computational resources and at high rates. The hand occlusion data can be generated by calculating a hand mesh based on live depth sensor data, which is acquired at a higher frequency than the graphical images. In some embodiments, the live depth sensor data can be acquired at a frame rate of at least 30 fps. To be able to process this data quickly, a small amount of data can be processed to create a hand model for use in occlusion processing by masking the live depth data using a model in which the hand is simply represented by multiple segments identified from key points. Additionally, to increase the accuracy of the occlusion processing, hand occlusion data can be generated by predicting the change in hand pose between the time the depth data is captured and the time the hand mesh will be used for occlusion processing. The hand mesh may be distorted to represent the hand in the predicted pose.
[0053] The techniques described herein can be used together or separately with many types of devices and for many types of scenarios, including wearable or portable devices with limited computational resources that provide cross-reality scenarios. In some embodiments, the techniques can be implemented by a service that forms part of an XR system.
[0054] Figure 1 and Figure 2 Such scenarios are shown. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3 - 6An exemplary AR system is shown that includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.
[0055] Reference Figure 1 depicts an outdoor AR scene 4, where a user of AR technology sees a park-like setting 6 of the physical world, characterized by people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the user of AR technology also perceives that they "see" a robotic statue 10 standing on the concrete platform 8 of the physical world, and a flying, cartoon-like avatar character 2 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 2 and robotic statue 10) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to produce AR technology that promotes a comfortable, natural feeling and rich virtual image element presentation among other virtual or physical world image elements.
[0056] Such an AR scene can be implemented using a system that includes a world reconstruction component, which can build and update a representation of the physical world surface around the user. This representation can be used for occlusion rendering, placing virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations that use information about the physical world. Figure 2 depicts another example of an indoor AR scene 200 according to some embodiments, showing exemplary world reconstruction use cases, including visual occlusion 202, physics-based interaction 204, and environmental reasoning 206.
[0057] The exemplary scene 200 is a living room with walls, a bookshelf on one side of the wall, a floor lamp at the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology can also perceive virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking out from the bookshelf, and an ornament in the form of a windmill placed on the coffee table. For the image on the wall, AR technology requires not only information about the wall surface but also information about objects and surfaces in the room (such as the shape of the lamp), which will occlude the image to correctly render the virtual object. For the flying bird, AR technology requires information about all objects and surfaces around the room in order to render the bird with realistic physical effects, to avoid objects and surfaces or to bounce when the bird collides. For the deer, AR technology requires information about the surface (such as the floor or the coffee table) to calculate the placement position of the deer. For the windmill, the system can identify that it is an object separate from the table and can infer that it is movable, while the corner of the shelf or the corner of the wall can be inferred to be stationary. This distinction can be used to infer which parts of the scene are used or updated in each of various operations.
[0058] A scene can be presented to a user via a system including a plurality of components, the plurality of components including a user interface that can stimulate one or more of the user's senses, including vision, sound, and / or touch. Additionally, the system can include one or more sensors that can measure parameters of a physical portion of the scene, including the position and / or movement of the user within the physical portion of the scene. Further, the system can include one or more computing devices, as well as associated computer hardware such as memory. These components can be integrated into a single device or more distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.
[0059] Figure 3 An AR system 302 is depicted in accordance with some embodiments, which is configured to provide an experience of AR content interacting with the physical world 306. The AR system 302 can include a display 308. In the illustrated embodiment, the display 308 can be worn by the user as part of a head-mounted headset such that the user can wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display can be transparent such that the user can observe see-through reality 310. The see-through reality 310 can correspond to the portion of the physical world 306 that is within the current viewpoint of the AR system 302, which can correspond to the user's viewpoint in the case where the user wears a head-mounted headset incorporating the display and sensors of the AR system to obtain information about the physical world.
[0060] AR content can also be presented on the display 308, overlaying the see-through reality 310. To provide accurate interaction between the AR content and the see-through reality 310 on the display 308, the AR system 302 can include sensors 322 configured to capture information about the physical world 306.
[0061] The sensors 322 can include one or more depth sensors that output depth maps 312. Each depth map 312 can have a plurality of pixels, each of which can represent the distance to a surface in the physical world 306 in a particular direction relative to the depth sensor. Raw depth data can be received from the depth sensors to create the depth maps. The depth maps can be updated as fast as the depth sensors can form new images, which can be hundreds or thousands of times per second. However, the data can be noisy and incomplete and have holes shown as black pixels in the illustrated depth maps.
[0062] The system may include other sensors, such as an image sensor. The image sensor may acquire information that can be processed to represent the physical world in other ways. For example, an image may be processed in the world reconstruction component 316 to create a mesh that represents the connected parts of objects in the physical world. Metadata about such objects (including, for example, color and surface texture) may similarly be acquired using sensors and stored as part of the world reconstruction.
[0063] The system may also acquire information about the user's head pose relative to the physical world. In some embodiments, the sensors may include an inertial measurement unit (IMU) that can be used to calculate and / or determine the head pose 314. The head pose 314 for the depth map may indicate, for example, the current viewing point of the sensor that captures the depth map in six degrees of freedom (6DoF), but the head pose 314 can be used for other purposes, such as associating image information with a particular part of the physical world or associating the position of a display worn on the user's head with the physical world. In some embodiments, the head pose information may be derived in other ways different from the IMU (such as analyzing objects in the image).
[0064] The world reconstruction component 316 may receive the depth map 312, the head pose 314, and any other data from the sensors and integrate the data into the reconstruction 318, which may at least appear to be a single combined reconstruction. The reconstruction 318 may be more complete and less noisy than the sensor data. The world reconstruction component 316 may update the reconstruction 318 using spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0065] The reconstruction 318 may include a representation of the physical world in one or more data formats (including, for example, voxels, meshes, planes, etc.). Different formats may represent alternative representations of the same part of the physical world or may represent different parts of the physical world. In the example shown, on the left side of the reconstruction 318, a part of the physical world is presented as a global surface; on the right side of the reconstruction 318, a part of the physical world is presented as a mesh.
[0066] The reconstruction 318 may be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or objects in the real world change. Aspects of the reconstruction 318 may be used, for example, by a component 320 that generates a changing global surface representation in world coordinates, which may be used by other components.
[0067] Based on this information, AR content can be generated, such as through the AR application 304. The AR application 304 can be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental reasoning. It can perform these functions by querying data in different formats from the reconstruction 318 generated by the world reconstruction component 316. In some embodiments, the component 320 can be configured to output an update when the representation in the region of interest of the physical world changes. For example, the region of interest can be set to approximate a part of the physical world near the user of the system, such as the part within the user's field of view, or projected (predicted / determined) to enter the user's field of view.
[0068] The AR application 304 can use this information to generate and update AR content. The virtual part of the AR content can be presented on the display 308 in combination with the see-through reality 310, thus creating a realistic user experience.
[0069] In some embodiments, an AR experience can be provided to the user through a wearable display system. Figure 4 An example of a wearable display system 80 (hereinafter referred to as "system 80") is shown. The system 80 includes a head-mounted display device 62 (hereinafter referred to as "display device 62"), and various mechanical and electronic modules and systems that support the functions of the display device 62. The display device 62 can be coupled to a frame 64, which can be worn by the user or viewer 60 of the display system (hereinafter referred to as "user 60") and is configured to position the display device 62 in front of the eyes of the user 60. According to various embodiments, the display device 62 can display sequentially. The display device 62 can be monocular or binocular. In some embodiments, the display device 62 can be Figure 3 an example of the display 308 in
[0070] In some embodiments, a speaker 66 is coupled to the frame 64 and positioned near the ear canal of the user 60. In some embodiments, another speaker (not shown) is positioned near the other ear canal of the user 60 to provide stereo / plastic sound control. The display device 62 is operatively coupled to a local data processing module 70, such as by a wired wire or a wireless connection 68, and the local data processing module 70 can be installed in various configurations, such as fixedly attached to the frame 64, fixedly attached to a helmet or hat worn by the user 60, embedded in headphones, or otherwise removably attached to the user 60 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0071] The local data processing module 70 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which may be used to assist in the processing, caching, and storage of data. The data includes: a) data captured from sensors (e.g., that may be operatively coupled to the frame 64) or otherwise attached to the user 60, such as image capture devices (such as cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, radio devices, and / or gyroscopes; and / or b) data that is obtained and / or processed using the remote processing module 72 and / or the remote data repository 74 and that may be passed to the display device 62 after such processing or acquisition. The local data processing module 70 may be operatively coupled to the remote processing module 72 and the remote data repository 74 via communication links 76, 78, such as via wired or wireless communication links, respectively, such that these remote modules 72 are operatively coupled to each other and may be used as resources for the local processing and data modules. In some embodiments, Figure 3 the world reconstruction component 316 in Figure 3 may be implemented at least in part in the local data processing module 70. For example, the local data processing module 70 may be configured to execute computer-executable instructions to generate a physical world representation based at least in part on at least a portion of the data.
[0072] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 70 may include a single processor (e.g., a single-core or multi-core ARM processor), which would limit the computational budget of the module but enable a smaller device. In some embodiments, the world reconstruction component 316 may use less than the computational budget of a single ARM core to generate a physical world representation in real time over a non-predefined space such that the remaining computational budget of a single ARM core may be accessed for other uses, such as, for example, extracting a mesh.
[0073] In some embodiments, the remote data repository 74 may include a digital data storage facility that may be made available via the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, thereby allowing for fully autonomous use from the remote modules. For example, the world reconstruction may be stored in whole or in part in the repository.
[0074] In some embodiments, the local data processing module 70 is operatively coupled to a battery 82. In some embodiments, the battery 82 is a removable power source, such as above a counter battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during non-operating times of the system 80 and a removable battery, such that the user 60 can operate the system 80 for longer periods of time without having to connect to a power source to charge the lithium-ion battery or without having to shut down the system 80 to replace the battery.
[0075] Figure 5A Shown is a user 30 wearing an AR display system that renders AR content as the user 30 moves through a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at location 34, and the AR display system records environmental information of the traversable world relative to location 34 (e.g., digital representations of real objects in the physical world, which can be stored and updated as the real objects change in the physical world), such as pose relationships to mapped features or directional audio inputs. Location 34 is aggregated into data input 36 and processed at least by a traversable world module 38, which can be implemented, for example, by processing on a remote processing module 72 Figure 4 In some embodiments, the traversable world module 38 can include a world reconstruction component 316.
[0076] The traversable world module 38 determines the location and manner in which AR content 40, as determined from data input 36, can be placed in the physical world. The AR content is "placed" in the physical world by presenting both the physical world rendering and the AR content via a user interface, with the AR content rendered as if interacting with objects in the physical world and the objects in the physical world presented as if the AR content occludes the user's view of those objects when appropriate. In some embodiments, the shape and location of the AR content 40 can be determined by appropriately selecting portions of a fixed element 42 (e.g., a table) from the reconstruction (e.g., reconstruction 318) to place the AR content. As an example, the fixed element can be a table, and the virtual content can be positioned such that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in the field of view 44, which can be the current field of view or an estimated future field of view. In some embodiments, the AR content can be placed relative to a mapped mesh model 46 of the physical world.
[0077] As depicted, the fixed element 42 serves as a proxy for any fixed element within the physical world that can be stored in the traversable world module 38, such that the user 30 can perceive the content on the fixed element 42 without the system having to map build to the fixed element 42 every time the user 30 sees the fixed element 42. Thus, the fixed element 42 can be a mapped grid model from a previous modeling session or can be determined by a separate user but still stored on the traversable world module 38 for future reference by multiple users. Thus, the traversable world module 38 can identify the environment 32 from a previously map-built environment and display AR content without the user 30's device first having to map build the environment 32, thereby saving computational processes and cycles and avoiding latency of any rendered AR content.
[0078] A mapped grid model 46 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 40 can be mapped and stored in the traversable world module 38 for future access by the user 30 or other users without remapping or modeling. In some embodiments, the data input 36 is an input such as a geographical location, user identification, and current activity to indicate to the traversable world module 38 which fixed element 42 among one or more fixed elements is available, which AR content 40 was last placed on the fixed element 42, and whether to display the same content (such AR content is "persistent" content regardless of how the user views a particular traversable world model).
[0079] Even in embodiments where an object is considered fixed, the traversable world module 38 can be updated from time to time to account for the possibility of changes in the physical world. The model of the fixed object may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed. To render an AR scene with a sense of realism, the AR system can update the positions of these non-fixed objects at a much higher frequency than that used to update the fixed objects. To be able to accurately track all objects in the physical world, the AR system can obtain information from multiple sensors, including one or more image sensors.
[0080] Figure 5BIs a schematic view of the viewing optical assembly 48 and its attached components. In some embodiments, two eye-tracking cameras 50 that point to the user's eyes 49 detect metrics of the user's eyes 49, such as the eye shape, eyelid occlusion, pupil direction, and blink on the user's eyes 49. In some embodiments, one of the sensors can be a depth sensor 51, such as a time-of-flight sensor, which emits signals into the world and detects the reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor can, for example, quickly determine whether an object has entered the user's field of view due to the movement of those objects or changes in the user's posture. However, information about the position of an object in the user's field of view can alternatively or additionally be collected by other sensors. Depth information can, for example, be obtained from a stereoscopic vision image sensor or a plenoptic sensor.
[0081] In some embodiments, the world camera 52 records a view larger than the periphery to map the environment 32 and detect inputs that can affect the AR content. In some embodiments, the world camera 52 and / or the camera 53 can be a grayscale and / or color image sensor, which can output grayscale and / or color image frames at fixed time intervals. The camera 53 can further capture an image of the physical world within the user's field of view at a specific time. Even if the values of the pixels of a frame-based image sensor do not change, sampling of its pixels can be repeated. Each of the world camera 52, the camera 53, and the depth sensor 51 has a corresponding field of view 54, 55, and 56 to collect data from and record a physical world scene of a physical world environment 32 such as Figure 5A depicted in the physical world environment 32.
[0082] The inertial measurement unit 57 can determine the movement and orientation of the viewing optical assembly 48. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 51 is operatively coupled to the eye-tracking camera 50 to confirm the measured accommodation relative to the actual distance at which the user's eyes 49 are gazing.
[0083] It should be understood that the viewing optical assembly 48 can include Figure 5B some of the components shown in, and can include components instead of or in addition to the components shown. For example, in some embodiments, the viewing optical assembly 48 can include two world cameras 52 instead of four. Alternatively or additionally, the cameras 52 and 53 do not need to capture visible light images of their entire fields of view. The viewing optical assembly 48 can include other types of components. In some embodiments, the viewing optical assembly 48 can include one or more dynamic vision sensors (DVSs), the pixels of which can respond asynchronously to relative changes in light intensity above a threshold.
[0084] In some embodiments, based on time-of-flight information, the viewing optical component 48 may not include the depth sensor 51. For example, in some embodiments, the viewing optical component 48 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and the angle of incident light, whereby the depth information may be determined. For example, a plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or additionally, a plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Instead of or in addition to the depth sensor 51, such sensors may be used as a source of depth information.
[0085] It should also be understood that Figure 5B the configuration of the components in is shown as an example. The viewing optical component 48 may include components having any suitable configuration, which may be set to provide the user with the maximum field of view practical for a particular set of components. For example, if the viewing optical component 48 has a world camera 52, the world camera may be placed in the central region of the viewing optical component rather than on the side.
[0086] The information from the sensors in the viewing optical component 48 may be coupled to one or more processors in the system. The processor may generate data that can be rendered to enable the user to perceive interacting with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, physical and virtual content may be depicted in a scene by modulating the opacity of a display device that the user views in the physical world. The opacity may be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, when viewed through a user interface, the image data may include only virtual content, which may be modified such that the virtual content is perceived by the user as interacting realistically with the physical world (e.g., clipping the content to account for occlusion). Regardless of how the content is presented to the user, a model of the physical world is required such that the characteristics of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, position, motion, and visibility of the virtual objects. In some embodiments, the model may include a reconstruction of the physical world, such as the reconstruction 318.
[0087] The model may be created based on data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected from multiple users, which may be aggregated in a computing device remote from all users (and the data may be "in the cloud").
[0088] The model may be created at least in part by a world reconstruction system, for example, Figure 6more particularly depicted in Figure 3 world reconstruction component 316. The world reconstruction component 316 can include a perception module 160 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 160 can represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel can correspond to a 3D cube of a predetermined volume in the physical world and include surface information that indicates whether there is a surface within the volume represented by the voxel. A value can be assigned to the voxel that indicates whether its corresponding volume has been determined to include a surface of a physical object, is determined to be empty, or has not been measured by the sensor and thus its value is unknown. It should be understood that it is not necessary to explicitly store the values of voxels determined to be empty or unknown, as the values of voxels can be stored in computer memory in any suitable manner, including not storing information for voxels determined to be empty or unknown.
[0089] In addition to generating information for the persistent world representation, the perception module 160 can also identify and output an indication of a change in the area around the user of the AR system. This indication of change can trigger an update to the volumetric data stored as part of the persistent world or trigger other functions, such as triggering component 304 that generates AR content to update the AR content.
[0090] In some embodiments, the perception module 160 can identify changes based on a signed distance function (SDF) model. The perception module 160 can be configured to receive sensor data such as, for example, depth map 160a and head pose 160b and then fuse the sensor data into the SDF model 160c. The depth map 160a can directly provide SDF information, and the image can be processed to obtain SDF information. The SDF information represents the distance from the sensor used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and thus from the perspective of the user. The head pose 160b can enable the SDF information to be related to the voxels in the physical world.
[0091] In some embodiments, the perception module 160 can generate, update, and store a representation of a portion of the physical world within the perception range. The perception range can be determined at least in part based on the reconstruction range of the sensor, which can be determined at least in part based on the limitations of the observation range of the sensor. As a specific example, an active depth sensor operating with active IR pulses can operate reliably within a certain distance range, creating an observation range of the sensor that can range from a few centimeters or tens of centimeters to several meters.
[0092] The world reconstruction component 316 may include additional modules that can interact with the perception module 160. In some embodiments, the persistent world module 162 may receive a representation of the physical world based on data acquired by the perception module 160. The persistent world module 162 may also include representations of the physical world in various formats. For example, volumetric metadata 162b such as voxels, as well as meshes 162c and planes 162d, may be stored. In some embodiments, other information such as depth maps may be saved.
[0093] In some embodiments, the perception module 160 may include modules that generate representations of the physical world in various formats, including for example meshes 160d, planes, and semantics 160e. These modules may generate the representations based on data within the perception range of one or more sensors at the time of generating the representation, as well as data captured at a previous time and information in the persistent world 162. In some embodiments, these components may operate with respect to depth information captured by a depth sensor. However, the AR system may include visual sensors and may generate such representations by analyzing monocular or binocular visual information.
[0094] In some embodiments, these modules may operate on regions of the physical world. When the perception module 160 detects a change in the physical world in a sub-region of the physical world, those modules may be triggered to update the sub-region of the physical world. For example, such a change may be detected by detecting a new surface in the SDF model 160c or other criteria (such as changing the values of a sufficient number of voxels representing the sub-region).
[0095] The world reconstruction component 316 may include components 164 that can receive a representation of the physical world from the perception module 160. Information about the physical world may be extracted by these components according to, for example, a usage request from an application. In some embodiments, the information may be pushed to the usage components, such as via an indication of a change in a pre-identified region or a change in the representation of the physical world within the perception range. The components 164 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.
[0096] In response to a query from the components 164, the perception module 160 may send a representation of the physical world in one or more formats. For example, when the components 164 indicate that the usage is for visual occlusion or physics-based interaction, the perception module 160 may send a representation of the surface. When the components 164 indicate that the usage is for environmental reasoning, the perception module 160 may send the mesh, plane, and semantics of the physical world.
[0097] In some embodiments, the perception module 160 may include components that format information for providing to component 164. Examples of such components may be the ray casting component 160f. Using a component (e.g., component 164) may query information about the physical world from a particular viewpoint. The ray casting component 160f may be selected from one or more representations of physical world data within the field of view from that viewpoint.
[0098] Information about the physical world can also be used for occlusion processing. This information may be used by the visual occlusion component 164a, which may be part of the world reconstruction component 316. For example, the visual occlusion component 164a may provide the application with information indicating which parts of a visual object are occluded by physical objects. Alternatively or additionally, the visual occlusion component 164a may provide the application with information about physical objects, which the application can use for occlusion processing. As described above, accurate information about the hand position is important for occlusion processing. In the examples described herein, the visual occlusion component 164a may maintain a hand model in response to a request from the application and provide the model to the application when requested. Figure 7 An example of such processing is shown, which may be performed across Figure 6 one or more of the components illustrated in, or, in some embodiments, by different or additional components.
[0099] Figure 7 An AR system 700 configured to generate a hand mesh for dynamic occlusion processing in real time according to some embodiments is depicted. The AR system 700 may be implemented on an AR device. The AR system 700 may include a data collection unit 702 configured to use sensors on the AR device to capture the pose (e.g., head pose, hand pose, etc.) of a user wearing the AR device (e.g., display device 62) and scene information. The scene information may include depth information that indicates the distance between the AR device and physical objects in the scene.
[0100] The data collection unit 702 includes a hand tracking component. The hand tracking component may process sensor data, such as depth information and image information, to detect one or more hands in the scene. Other sensor data may be processed to detect one or more hands in the scene. When detected, one or more hands may be represented in a sparse manner, e.g., by a set of key points. For example, the key points may represent joints, fingertips, or other boundaries of hand segments. The information collected or generated by the data collection unit 702 may be passed to the hand meshing unit 704 for generating a richer model of one or more hands, e.g., a mesh, based on the sparse representation.
[0101] The hand meshing unit 704 is configured to calculate a hand mesh of one or more detected hands and update the hand mesh in real time as the pose changes and / or the hand moves.
[0102] The AR system 700 may include an application 706 that is configured to receive the hand mesh from the hand meshing unit 704 and render one or more virtual objects in the scene. In some embodiments, the application 706 may receive occlusion data from the hand meshing unit 704. In some embodiments, the occlusion data may indicate the portions of the virtual objects that are occluded by one or more hands. In some embodiments, such as the illustrated embodiment, the occlusion data may be a model of one or more hands, and the application 706 may calculate the occlusion data based on the model. As a specific example, the occlusion data may be the hand mesh of one or more hands received from the hand meshing unit 704.
[0103] Figure 8 is a flowchart showing a method 800 for real-time generating a hand mesh for dynamic occlusion according to some embodiments. In some embodiments, the method 800 may be executed by one or more processors within the AR system 700. The method 800 may begin when the hand meshing unit 704 of the AR system 700 receives (act 802) a query for data related to one or more hands in the scene from the application 706 of the AR system 700. The method 800 may include detecting (act 804) one or more hands in the scene based on information about the scene captured by the data collection unit 702 of the AR system 700.
[0104] When one or more hands are detected, the method 800 may include calculating (act 806) one or more models of the one or more hands based on the information about the scene. The one or more models of the one or more hands may be sparse, indicating the positions of key points on the hand rather than the surface. Those key points may represent the joints or end portions of hand segments. The key points may be identified from sensor data about the one or more hands, including, for example, stereo images of the one or more hands. Depth information and, in some cases, monocular images of the one or more hands may alternatively or additionally be used to identify the key points.
[0105] U.S. Provisional Patent Application No. 62 / 850,542, entitled "Hand Pose Estimation," describes exemplary methods and apparatuses for obtaining information about the position and pose of a hand and modeling the hand based on the obtained information. A copy of the filed version of U.S. Application No. 62 / 850,542 is attached hereto as an appendix, and the entire contents thereof are incorporated herein by reference for all purposes. The techniques described in that application may be used to construct a sparse model of the hand.
[0106] In some embodiments, one or more models of one or more hands can be calculated based on scene information captured by sensors of an AR device. Examples of scene information include an example image captured by one sensor corresponding to a single eye, two example images captured by two sensors corresponding to the left and right eyes, and an exemplary depth image that can be obtained at least in part from an image of Figure 9A or an image of Figure 9B . Figure 9A of Figure 9A , two example images captured by two sensors corresponding to the left and right eyes, and Figure 9B of Figure 9B , and Figure 9C an exemplary depth image of Figure 9C that can be obtained at least in part from an image of Figure 9A or an image of Figure 9B . Figure 9A or Figure 9B of Figure 9B .
[0107] In some embodiments, one or more models of one or more hands can include a plurality of key points of the hand, which can indicate points on segments of the hand. Some key points can correspond to joints of the hand and fingertips of the hand. Figure 9E and Figure 9F depict schematic diagrams showing an exemplary eight-key point model of a hand and an exemplary twenty-two-key point model of a hand, respectively.
[0108] In some embodiments, the key points of the hand model can be used to determine the contour of the hand. Figure 9D depicts an exemplary contour of a hand that can be determined based on the key points of the hand model. For example, adjacent key points can be connected by lines, as schematically shown in Figure 9E and Figure 9F , and the contour of the hand can be represented as a distance from the lines. The distance from the lines can be determined based on an image of the hand, information about human anatomy, and / or other information. It should be understood that once the key points of the hand are identified, the model of the hand can be updated later using previously acquired information about the hand. For example, the length of the lines connecting the key points may not change. Figure 9E and Figure 9F As schematically shown, and the contour of the hand can be represented as a distance from the lines. The distance from the lines can be determined based on an image of the hand, information about human anatomy, and / or other information. It should be understood that once the key points of the hand are identified, the model of the hand can be updated later using previously acquired information about the hand. For example, the length of the lines connecting the key points may not change.
[0109] A sparse hand model can be used to select a limited amount of data from which a richer hand model, such as including surface information, can be constructed. In some embodiments, the selection can be performed by using the contour of the hand to mask additional data (such as depth data). Thus, method 800 can include masking (act 808) depth information indicating the distance between the AR device and physical objects in the scene using one or more models of one or more hands.
[0110] Figure 10Depicts a flowchart showing details of masking (action 808) depth information with one or more models of one or more hands. Action 808 may include filtering out (action 1002) depth information outside the contours of one or more models of one or more hands. This action results in removing depth information associated with physical objects other than the hands in the scene. Action 808 may include generating (action 1004) a depth image of one or more hands based on the filtered depth information.
[0111] The depth image may include pixels, each of which may indicate the distance to a point on one or more hands. In some embodiments, depth information may be captured in a way that does not capture depth information for all surfaces of one or more hands. For example, depth information may be captured using an active IR sensor. For example, if a user wears a ring with a dark gemstone, the IR may not be reflected by the dark ring, resulting in holes in the insufficient information collected. Action 808 may include filling (action 1006) the holes in the depth image. In some embodiments, the holes may be filled by identifying the holes in the depth image and generating stereoscopic depth information corresponding to the identified regions from the stereoscopic cameras of the augmented reality device. In some embodiments, the holes may be filled by identifying the holes in the depth image and accessing surface information corresponding to the identified regions from one or more 3D models of one or more hands.
[0112] Optionally, one or more hand meshes may be calculated as a plurality of sub-meshes, each sub-mesh representing a segment of one or more hands. The segments may correspond to segments bounded by key points. Since many of these segments are bounded by joints, these segments correspond to parts of the hand that can move independently of at least some other segments of the hand, and thus the hand mesh can be updated quickly by updating the sub-meshes associated with the segments that have moved since the last hand mesh calculation. In such an embodiment, action 808 may include identifying (action 1008) key points in the depth image corresponding to key points of one or more models of one or more hands. The hand segments separated by the key points may be calculated. Action 808 may include associating (action 1010) portions of the depth image to the hand segments separated by the key points identified in the depth image.
[0113] Method 800 may include calculating (action 810) one or more hand meshes based on the masked depth information to one or more models of one or more hands. One or more hand meshes may be a representation of one or more hands, indicating the surfaces of one or more hands. Figure 9G Depicts a calculated dense hand mesh according to some embodiments. However, the present application is not limited to calculating dense hand meshes. In some embodiments, sparse hand meshes may be sufficient for dynamic occlusion.
[0114] In some embodiments, one or more hand meshes may be calculated based on depth information. For example, a mesh may be a collection of regions, typically represented as triangles, representing a portion of a surface. Such regions may be identified by grouping adjacent pixels in a depth image that have a distance difference less than a threshold, indicating that these pixels are likely on the same surface. One or more triangles defining such pixel regions may be identified and added to the mesh. However, other techniques for forming a mesh based on depth information may be used.
[0115] Calculating one or more hand meshes may include updating one or more hand meshes in real time as the relative position between the AR device and one or more hands changes. There may be a latency between the time when scene information for calculating one or more hand meshes is captured and the time when one or more calculated hand meshes are used by an application (e.g., to render content). In some embodiments, the movement of segments of one or more hands may be tracked so that the future positions of those segments of one or more hands can be projected / predicted. When virtual objects processed using one or more hand meshes are to be rendered, one or more hand meshes may be deformed to conform to the speculated positions of the segments.
[0116] Figure 11 A flowchart is depicted showing details of real-time calculation (act 810) of one or more hand meshes based on depth information masked to hand segments according to some embodiments. Act 810 may include determining (act 1102) the time t at which the hand mesh division unit 704 receives a query for data related to one or more hands from the application 706.
[0117] Act 810 may include predicting (act 1104) a latency n based on the query received from the application 706 at time t. Act 810 may include predicting (act 1106) a pose (e.g., a hand pose) at a time that is the query time t plus the latency n. Act 810 may include deforming one or more hand meshes using the predicted pose at a time that is the query time t plus the latency n (act 1108). In some embodiments, predicting a hand pose may include predicting the movement of key points of one or more hands in a predicted depth image at a time that is the query time t plus the latency n. Such a prediction may be made based on tracking the positions of the key points over time. This tracking enables determination of motion parameters such as speed or acceleration. Assuming that the determined motion parameters remain the same, extrapolation from a previous position to a future position may be used to project the position. Alternatively or additionally, a Kalman filter or similar projection technique may be used to determine the projection.
[0118] In some embodiments, deforming one or more hand meshes using the predicted pose may include deforming portions of the one or more hand meshes corresponding to a subset of hand segments that are between keypoints predicted to change at a query time t plus a delay n.
[0119] In deforming a previously calculated mesh to represent one or more hands at time t plus a delay n, a variety of factors may be considered. For example, the value of n may reflect the time required to deform the one or more meshes and the application to use the one or more meshes in a rendered object. The value may be estimated or measured from the structure or testing in the operation of the software. As another example, the value may be dynamically determined based on measuring the delay in use or adjusting the previously established delay based on the processing load when making a request for the one or more meshes.
[0120] When deforming one or more meshes, the amount of deformation can be based on the time when data is captured to form one or more hand meshes and the delay until the one or more meshes will be used. In some embodiments, one or more hand meshes can be created in addition to responding to a request from an application. For example, once an application indicates (e.g., through an API call) that it is configured for occlusion processing, the AR system can periodically calculate one or more updated hand meshes. Alternatively, the hand tracking process can run continuously using a certain amount of the system's computing resources. In any case, one or more hand meshes can be updated relatively frequently, such as at least 30 times per second. Nevertheless, there may be a delay between capturing data to make a mesh and receiving a mesh request, and this delay can also be taken into account in deforming the hand mesh.
[0121] Method 800 may include providing (act 812) the one or more hand meshes to application 706 so that the application renders portions of the virtual object that are not occluded by the one or more hand meshes.
[0122] Having thus described several aspects of some embodiments, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.
[0123] As an example, embodiments are described in conjunction with an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied in an MR environment or more generally in other XR environments and VR environments.
[0124] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via any suitable combination of a network (such as the cloud), a discrete application, and / or a device, network, and a discrete application.
[0125] Such changes, modifications, and improvements are intended to be part of this disclosure and to be within the spirit and scope of this disclosure. Additionally, although the advantages of this disclosure are indicated, it should be understood that not every embodiment of this disclosure will include every described advantage. In some cases, some embodiments may not implement any of the features described herein as advantageous. Accordingly, the foregoing description and drawings are provided by way of example only.
[0126] The above embodiments of this disclosure can be implemented in any of a variety of ways. For example, an embodiment can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. Such a processor can be implemented as an integrated circuit having one or more processors in the integrated circuit components, including commercially available integrated circuit components known in the art, having names such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor can be implemented in a custom circuit (such as an ASIC) or in a semi-custom circuit created by configuring a programmable logic device. As another alternative, the processor can be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores such that one or a subset of these cores can constitute the processor. However, the processor can be implemented using any suitable format of circuitry.
[0127] Furthermore, it should be understood that a computer can be embodied in any of a variety of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer can be embedded in a device that is not typically considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smart phone, or any other suitable portable or stationary electronic device.
[0128] In addition, the computer may have one or more input and output devices. These devices can be used in particular to present a user interface. Examples of output devices that can be used to provide a user interface include printers or displays for visual presentation of output, and speakers or other sound generating devices for auditory presentation of output. Examples of input devices that can be used for a user interface include keyboards and pointing devices such as mice, touchpads, and digitizing tablet computers. As another example, the computer can receive input information by voice recognition or other audible formats. In the illustrated embodiment, the input / output devices are shown physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated into the same unit as the processor or other elements of the computing device. For example, a keyboard may be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated via a wireless connection.
[0129] Such computers can be interconnected via one or more networks in any suitable form, including as a local area network or a wide area network such as a corporate network or the Internet. Such networks can be based on any suitable technology, and can operate according to any suitable protocol, and can include wireless networks, wired networks, or fiber optic networks.
[0130] In addition, the various methods or processes outlined herein can be encoded as software executable on one or more processors employing any of a variety of operating systems or platforms. Additionally, such software can be written using any of a variety of suitable programming languages and / or programming or scripting tools, and can also be compiled into executable machine language code or intermediate code that is executed on a framework or virtual machine.
[0131] In this regard, the present disclosure may be embodied as one or more program-encoded computer-readable storage media (or multiple computer-readable media) (e.g., computer memory, one or more floppy disks, compact discs (CDs), optical discs, digital video discs (DVDs), magnetic tapes, flash memories, field programmable gate arrays or other circuitry in semiconductor devices or other tangible computer storage media), which, when executed on one or more computers or other processors, will perform methods implementing the various embodiments of the present disclosure discussed above. As is apparent from the foregoing examples, the computer-readable storage media can retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such one or more computer-readable storage media can be removable, such that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage media" encompasses only computer-readable media that can be considered a manufacture (i.e., a manufactured article) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage media, such as a propagated signal.
[0132] In a general sense, the terms "program" or "software" are used herein to refer to a set of computer code or computer-executable instructions that can be used to program a computer or other processor to implement the various aspects of the present disclosure as described above. Additionally, it should be understood that, according to one aspect of this embodiment, one or more computer programs that execute the methods of the present disclosure when executed need not reside on a single computer or processor, but can be distributed in a modular fashion among multiple different computers or processors to implement the various aspects of the present disclosure.
[0133] Computer-executable instructions can take many forms, such as program modules executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Generally, in various embodiments, the functions of program modules can be combined or distributed as needed.
[0134] Furthermore, data structures can be stored in a computer-readable medium in any suitable form. For simplicity of illustration, a data structure may be shown to have fields that are related by their positions in the data structure. Similarly, such relationships can be implemented by allocating storage for fields based on their positions in the computer-readable medium that convey the relationships between the fields. However, any suitable mechanism can be used to establish relationships between the information in the fields of a data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.
[0135] Aspects of the present disclosure may be used alone, in combination, or in a variety of arrangements not specifically discussed in the foregoing embodiments, and thus, in their application, are not limited to the details and arrangements of the components set forth in the foregoing description or shown in the drawings. For example, aspects described in one embodiment may be combined with aspects described in other embodiments in any manner.
[0136] In addition, the present disclosure may be embodied as a method, and an example of the method has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which the acts are performed in an order different from that shown, and some acts may be performed simultaneously, even though shown as sequential acts in the illustrative embodiments.
[0137] The use of ordinal terms such as "first", "second", "third", etc. in the claims to modify a claim element itself does not denote any precedence, priority, or order of one claim element with respect to another in the order of performing method acts or a temporal order, but is merely used as a label to distinguish one claim element having a certain name from another having the same name (but for which an ordinal term is used) to distinguish claim elements.
[0138] Also, the words and terms used herein are for the purpose of description and should not be regarded as limiting. The use of "including", "comprising", or "having", "contains", "involves", and variations thereof herein is intended to cover the items listed thereafter and their equivalents as well as other items.
Claims
1. A method of operating a computing system to reconstruct a hand for dynamically occluding a virtual object, the method comprising: Receive a query for data related to a hand in the scene from an application that renders a virtual object in the scene; Obtain information about the scene from a device worn by the user, the device including one or more sensors, the information about the scene including depth information indicating a distance between the device worn by the user and a physical object in the scene; Detect whether the physical object in the scene includes a hand; When the hand is detected, calculate a model of the hand at least in part based on the information about the scene; Mask the depth information indicating the distance between the device worn by the user and the physical object in the scene with the model of the hand; Based on the depth information masked to the model of the hand, calculate a hand mesh, the calculation including: updating the hand mesh in real time as the relative position between the device and the hand changes; And Provide the hand mesh to the application so that the application renders the portions of the virtual object not occluded by the hand mesh.
2. The method according to claim 1, wherein: The model of the hand includes a plurality of key points of the hand indicating points on segments of the hand.
3. The method according to claim 2, wherein: At least a portion of the plurality of key points of the hand corresponds to joints of the hand and fingertips of the hand.
4. The method according to claim 2, wherein: The method further includes: determining a contour of the hand based on the plurality of key points; and Masking the depth information includes: Filtering out the depth information outside the contour of the model of the hand; and Generating a depth image of the hand at least in part based on the filtered depth information, the depth image including a plurality of pixels, each pixel indicating a distance to a point on the hand.
5. The method according to claim 4, wherein, Filtering out the depth information outside the contour of the model of the hand includes: Removing the depth information associated with the physical object in the scene.
6. The method according to claim 2, wherein: Masking the depth information indicating the distance between the device worn by the user and the physical object in the scene with the model of the hand includes: associating portions of the depth image with hand segments; and Updating the hand mesh in real time includes: selectively updating portions of the hand mesh representing a suitable subset of the hand segments.
7. The method according to claim 6, further comprising: Before calculating the hand mesh, fill holes in the depth image.
8. The method according to claim 7, wherein, Filling holes in the depth image includes: Generating stereo depth information from a stereo camera of the device, the stereo depth information corresponding to a region of the holes in the depth image.
9. The method according to claim 7, wherein, Filling holes in the depth image includes: Accessing surface information from a 3D model of the hand, the surface information corresponding to a region of the holes in the depth image.
10. The method according to claim 1, wherein, Calculating the hand mesh based on the depth information masked to the model of the hand includes: Predicting a latency n according to a query for the data related to the hand in the scene received from the application that renders the virtual object in the scene at time t; Predicting a hand pose at a time of query time t plus the latency n; and Deform the hand mesh using the predicted pose at the time of the query time t plus the delay n.
11. The method according to claim 1, wherein, The depth information indicating the distance between the device worn by the user and the physical object in the scene includes a sequence of depth images at a frame rate of at least 30 frames per second.
12. An electronic system that can be carried by a user, comprising: A device worn by the user, wherein the device includes a display configured to render virtual objects and includes one or more sensors configured to capture the head pose of the user wearing the device and information about a scene including one or more physical objects, the information about the scene including depth information indicating the distance between the device and the one or more physical objects; A hand mesh division component configured to execute computer-executable instructions to detect a hand in the scene, calculate a hand mesh of the detected hand, and update the hand mesh in real time as the head pose changes and / or the hand moves; and An application configured to execute computer-executable instructions to render the virtual object in the scene, wherein the application receives the hand mesh and the portion of the virtual object occluded by the hand from the hand mesh division component.
13. The electronic system according to claim 12, wherein: The hand mesh division component is configured to calculate the hand mesh by: Identifying key points on the hand; Calculating segments between the key points; Selecting information from the depth information based on proximity to one or more of the calculated segments; And Based on the selected depth information, calculating a mesh representing at least a portion of the hand mesh.
14. The electronic system according to claim 13, wherein: The depth information includes a plurality of pixels, each of the plurality of pixels representing a distance to an object in the scene; And Calculating the mesh includes: grouping adjacent pixels representing a distance difference less than a threshold.
Citation Information
Patent Citations
Realistic occlusion for a head mounted augmented reality display
CN103472909A
Virtual-real occlusion interaction method and system under AR environment
CN109471521A