Fast hand meshing for dynamic occlusion
The method of fast hand meshing in XR systems addresses latency and computational challenges by generating hand meshes from depth sensor data with predicted hand poses, ensuring accurate and immersive occlusion of virtual objects with the user's hand.
Patent Information
- Application Number
- JP2025075562
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing XR systems face challenges in dynamically occluding virtual objects with the user's hand due to high computational requirements and latency issues, especially when hands move rapidly, leading to inaccurate rendering and breaking user immersion.
A method for fast hand meshing that involves generating a hand mesh from live depth sensor data at high frame rates, using a hand model represented by feature points, and predicting hand pose changes to ensure accurate occlusion processing with low computational resources.
Enables realistic and immersive XR experiences by accurately occluding virtual objects with the user's hand in real-time, maintaining high frame rates and reducing computational burden.
Smart Images

Figure 2025107314000001_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to a cross-reality system that uses 3D world reconstruction to render scenes.
Background Art
[0002] A computer can create an X Reality (XR or cross-reality) environment that controls a human user interface and in which part or all of the XR environment is generated by the computer as perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can describe, for example, virtual objects that can be rendered so that a user can perceive or sense them as part of the physical world and interact with the virtual objects. The user can experience these virtual objects as a result of data being rendered and presented through a user interface device such as a head-mounted display device. The data can be controlled to display for the user to see, or to reproduce audio for the user to hear, or to control a haptic (or tactile) interface to enable the user to experience a touch sensation that the user perceives or senses as the user feels the virtual object.
[0003] XR systems can be useful for many applications spanning the fields of scientific visualization, medical training, engineering design, and prototyping, teleoperation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more objects in relation to real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment when using XR systems and also expands the possibilities for various applications by presenting realistic and easily understandable information about how the physical world can be transformed.
[0004] The XR system can represent the physical world around the user of the system as a "mesh". The mesh can be represented by a plurality of interconnected triangles. Each triangle has edges that join points on the surface of an object within the physical world such that each triangle represents a part of the surface. Information about parts of the surface, such as color, texture, or other properties, can be stored associated within the triangle. In operation, the XR system processes image information to detect points and surfaces so as to create or update the mesh.
Summary of the Invention
Means for Solving the Problems
[0005] Aspects of the present application relate to methods and apparatuses for fast hand meshing for dynamic occlusion. The techniques described herein may be used together, separately, or in any suitable combination.
[0006] Some embodiments relate to a method of operating a computing system and reconstructing a hand to dynamically occlude virtual objects. The method includes receiving, from an application that renders virtual objects within a scene, a query regarding data related to a hand within the scene; capturing information of the scene from a device worn by a user, the device comprising one or more sensors, the information of the scene including depth information indicative of a distance between the device worn by the user and a physical object within the scene; detecting whether the physical object within the scene includes the hand; when the hand is detected, calculating a hand model at least in part based on the information of the scene; masking depth information indicative of a distance between the device worn by the user and the physical object within the scene using the hand model; calculating a hand mesh based on the depth information masked with respect to the hand model, the calculating step including updating the hand mesh in real time as a relative location between the device and a change of the hand; and supplying the hand mesh to the application such that the application renders portions of the virtual objects not occluded by the hand mesh.
[0007] In some embodiments, the hand model includes a plurality of feature points of the hand indicating points on segments of the hand.
[0008] In some embodiments, at least some of the plurality of feature points of the hand correspond to joints of the hand and fingertips of the hand.
[0009] In some embodiments, the method further includes determining a hand contour based on a plurality of feature points, and masking depth information indicating a distance between a device worn by a user and a physical object in a scene using a hand model. The step of masking depth information includes filtering and removing depth information outside the contour of the hand model, and generating a depth image of the hand based at least in part on the filtered depth information, the depth image including a plurality of pixels, each pixel indicating a distance to a point on the hand.
[0010] In some embodiments, the step of filtering and removing depth information outside the contour of the hand model includes removing depth information associated with a physical object in the scene.
[0011] In some embodiments, the step of masking depth information indicating a distance between a device worn by a user and a physical object in a scene using a hand model includes associating a portion of the depth image with a hand segment, and the step of updating the hand mesh in real time includes selectively updating a portion of the hand mesh representing a suitable subset of the hand segment.
[0012] In some embodiments, the method further includes filling holes in the depth image before calculating the hand mesh.
[0013] In some embodiments, the step of filling holes in the depth image includes generating stereo depth information from a stereo camera of the device, the stereo depth information corresponding to a region of the holes in the depth image.
[0014] In some embodiments, the step of filling holes in the depth image includes accessing surface information from a 3D model of the hand, the surface information corresponding to a region of the holes in the depth image.
[0015] In some embodiments, the step of calculating a hand mesh based on depth information masked for a hand model includes predicting a latency n from a query received at time t regarding data related to the hand in the scene from an application that renders virtual objects in the scene, predicting a hand pose at time t + latency n, and distorting the hand mesh using the pose predicted at time t + latency n.
[0016] In some embodiments, the depth information indicating the distance between a device worn by a user and a physical object in the scene includes a sequence of depth images at a frame rate of at least 30 frames per second.
[0017] Some embodiments relate to an electronic system portable by a user. The electronic system includes a device worn by the user. The device includes a display configured to render virtual objects and one or more sensors configured to capture information about the head pose of the user wearing the device and a scene including one or more physical objects. The information about the scene includes depth information indicating the distance between the device and the one or more physical objects. The electronic system includes a hand meshing component configured to execute computer-executable instructions to detect a hand in the scene, calculate a hand mesh of the detected hand, and update the hand mesh in real time as the head pose changes and / or the hand moves, and an application configured to execute computer-executable instructions to render virtual objects in the scene and receive, from the hand meshing component, the hand mesh and the portion of the virtual object occluded by the hand.
[0018] In some embodiments, the hand meshing component is configured to calculate a hand mesh by identifying feature points on the hand, calculating segments between the feature points, selecting information based on proximity to one or more of the calculated segments from depth information, and calculating a mesh representing at least a portion of the hand mesh based on the selected depth information.
[0019] In some embodiments, the depth information includes a plurality of pixels, each of which represents the distance to an object in the scene. The step of calculating the mesh includes grouping adjacent pixels that represent differences at distances less than a threshold.
[0020] Some embodiments relate to a method of operating an AR system to render virtual objects within a scene that includes physical objects. The AR system includes at least one sensor and at least one processor. The method includes capturing information about the scene using at least one sensor, the information about the scene including depth information indicating the distance to physical objects in the scene; processing the captured information using at least one processor to detect a hand in the scene and calculate points on the hand; selecting a subset of the depth information based on proximity to the calculated points on the hand; and calculating a representation of the hand based on the selected depth information, the representation of the hand indicating the surface of the hand.
[0021] In some embodiments, the method further includes storing the calculated representation of the hand and continuously processing the captured information to update the stored representation of the hand.
[0022] In some embodiments, the step of calculating a hand representation includes calculating one or more parameters of the hand movement based on the captured information, projecting the hand position at a future time based on a latency associated with the step of rendering a virtual object using the calculated hand representation based on the one or more parameters of the movement, and morphing the calculated hand representation to represent the hand at the projected position.
[0023] In some embodiments, the method further includes rendering a selected portion of the virtual object based on the hand representation, the selected portion representing a portion of the virtual object that is not occluded by the hand.
[0024] In some embodiments, the depth information includes a depth map including a plurality of pixels each representing a distance. Based on the selected depth information, the step of calculating a hand representation includes identifying a group of pixels representing a surface segment.
[0025] In some embodiments, the step of calculating a hand representation includes defining a mesh representing the hand based on the identified group of pixels.
[0026] In some embodiments, the step of defining a mesh includes identifying triangular regions corresponding to the identified surface segments.
[0027] The foregoing summary is provided by way of illustration and not by way of limitation. The present invention provides, for example, the following. (Item 1) A method of operating a computing system to dynamically occlude a virtual object and reconstruct a hand, the method comprising: receiving, from an application that renders a virtual object in a scene, a query regarding data related to a hand in the scene; Capturing information of the scene from a device worn by a user, the device comprising one or more sensors, the information of the scene including depth information indicating a distance between the device worn by the user and a physical object in the scene; Detecting whether a physical object in the scene includes a hand; When the hand is detected, calculating a model of the hand, at least partially based on the information of the scene; Using the model of the hand to mask depth information indicating a distance between the device worn by the user and a physical object in the scene; Calculating a mesh of the hand based on the depth information masked with respect to the model of the hand, the calculating including updating the mesh of the hand in real time as a relative location between the device and a change of the hand; Supplying the mesh of the hand to an application so that the application renders a portion of the virtual object not occluded by the mesh of the hand; A method comprising. (Item 2) The method according to item 1, wherein the model of the hand includes a plurality of feature points of the hand indicating points on segments of the hand. (Item 3) The method according to item 2, wherein at least some of the plurality of feature points of the hand correspond to joints of the hand and fingertips of the hand. (Item 4) The method further comprises Determining an outline of the hand based on the plurality of feature points; Using the model of the hand to mask depth information indicating a distance between the device worn by the user and a physical object in the scene; Including, wherein masking the depth information Filtering and removing depth information outside an outline of the model of the hand; At least partially generating the depth image of the hand based on the filtered depth information, wherein the depth image includes a plurality of pixels, and each pixel indicates the distance to a point on the hand The method according to item 2, including (Item 5) Filtering and removing the depth information outside the contour of the hand model includes The method according to item 4, including removing the depth information associated with the physical object in the scene (Item 6) Masking the depth information indicating the distance between the device worn by the user and the physical object in the scene using the hand model includes associating a portion of the depth image with the hand segment Updating the hand mesh in real time includes selectively updating the portion of the hand mesh representing a suitable subset of the hand segment The method according to item 2 (Item 7) The method according to item 6, further including filling the holes in the depth image before calculating the hand mesh (Item 8) Filling the holes in the depth image includes Generating stereo depth information from the stereo camera of the device, wherein the stereo depth information corresponds to the area of the holes in the depth image. The method according to item 7 (Item 9) Filling the holes in the depth image includes Accessing the surface information from the 3D model of the hand, wherein the surface information corresponds to the area of the holes in the depth image. The method according to item 7 (Item 10) Calculating the hand mesh based on the depth information masked for the hand model includes Predicting the waiting time n from the query received at time t regarding the data related to the hand in the scene from the application that renders the virtual object in the scene Predicting the hand posture at the time of the query time t + the waiting time n; Distorting the hand mesh using the posture predicted at the time of the query time t + the waiting time n; The method according to item 1, comprising: (Item 11) The depth information indicating the distance between the device worn by the user and the physical object in the scene includes a sequence of depth images at a frame rate of at least 30 frames per second, according to the method of item 1. (Item 12) An electronic system portable by a user, A device worn by the user, the device comprising a display configured to render a virtual object, and one or more sensors configured to capture information about the head posture of the user wearing the device and a scene including one or more physical objects, the information of the scene including depth information indicating the distance between the device and the one or more physical objects, a device, and A hand meshing component, the hand meshing component executing computer-executable instructions to detect a hand in the scene as the head posture changes and / or the hand moves, calculate a hand mesh of the detected hand, and update the hand mesh in real time; a hand meshing component, An application, the application executing computer-executable instructions to be configured to render the virtual object in the scene, the application receiving, from the hand meshing component, the hand mesh and the portion of the virtual object occluded by the hand; an application An electronic system comprising: (Item 13) The hand meshing component, Identifying feature points on the hand; Calculating segments between the feature points; selecting information based on the proximity to one or more of the calculated segments from the depth information; calculating a mesh representing at least a part of the hand mesh based on the selected depth information A portable electronic system according to item 12, configured to calculate a hand mesh by performing the above. (Item 14) The depth information includes a plurality of pixels, and each of the plurality of pixels represents the distance to an object in the scene. Calculating the mesh includes grouping adjacent pixels representing differences at distances less than a threshold. A portable electronic system according to item 13. (Item 15) A method of operating an AR system to render virtual objects in a scene including physical objects, the AR system comprising at least one sensor and at least one processor, the method comprising: capturing scene information using the at least one sensor, the scene information including depth information indicating the distance to physical objects in the scene; using the at least one processor, processing the captured information to detect a hand in the scene and calculate points on the hand; selecting a subset of the depth information based on the proximity to the calculated points on the hand; calculating a representation of the hand based on the selected depth information, the representation of the hand indicating the surface of the hand; A method including the above. (Item 16) storing the calculated hand representation; continuously processing the captured information and updating the stored hand representation The method according to item 15, further including the above. (Item 17) Calculating the representation of the hand includes calculating one or more parameters of the movement of the hand based on the captured information, projecting the position of the hand at a future time determined based on a latency associated with rendering a virtual object using the calculated representation of the hand based on the one or more parameters of the movement, and performing a morphing process on the calculated representation of the hand to represent the hand at the projected position, The method according to item 15, comprising. (Item 18) The method according to item 15, further comprising rendering a selected portion of the virtual object based on the representation of the hand, wherein the selected portion represents a portion of the virtual object not occluded by the hand. (Item 19) The depth information includes a depth map including a plurality of pixels each representing a distance, Calculating the representation of the hand based on the selected depth information includes identifying a group of pixels representing a surface segment, The method according to item 15. (Item 20) Calculating the representation of the hand includes defining a mesh representing the hand based on the identified group of pixels, the method according to item 19. (Item 21) Defining the mesh includes identifying triangular regions corresponding to the identified surface segments, the method according to item 20.
Brief Description of the Drawings
[0028] The accompanying drawings are not intended to be drawn to scale. In the drawings, each same or substantially same component illustrated in various figures is represented by a like number. For the purpose of clarity, not all components are labeled in all the drawings.
[0029]
Figure 1
[0030]
Figure 2
[0031]
Figure 3
[0032]
Figure 4
[0033]
Figure 5A
[0034]
Figure 5B
[0035]
Figure 6
[0036]
Figure 7
[0037]
Figure 8
[0038]
Figure 9
[0039]
Figure 10
[0040]
Figure 11
[0041] What is described in this specification is a method and apparatus for fast hand meshing for dynamic occlusion in an XR system. The XR system may create and use a 3D reconstruction. To provide a realistic XR experience to the user, the XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects with real objects. World reconstruction may be built from images and depth information about those physical surroundings collected using sensors that are part of the XR system. World reconstruction may then be used by any of a plurality of components of such a system. For example, world reconstruction may be used by a component that performs visual occlusion processing, calculates physics-based interactions, or performs environment inference.
[0042] Occlusion processing identifies portions of the virtual object that should not be rendered and / or displayed to the user because objects in the physical world block the user's view of where the virtual object should be perceived by the user. Physics-based interactions are calculated to determine where or how a virtual object appears to the user. For example, a virtual object may be rendered to appear to be resting on a physical object, moving through the air, or colliding with the surface of a physical object. World reconstruction provides a model from which information about objects in the physical world can be obtained for such calculations.
[0043] When providing such a system, there are significant challenges. Substantial processing may be required to perform world reconstruction and calculate occlusion information. Additionally, the XR system must correctly understand how to position virtual objects in relation to the user's head, body, etc. As the user's position in relation to the physical environment changes, the relevant parts of the physical world may also change, which may require further processing. Furthermore, 3D reconstruction data is often required to be updated as objects move within the physical world (e.g., a cup moves on a table). Updates to the data representing the environment the user is experiencing must be performed quickly without using too many computing resources of the computer generating the XR environment, as it is impossible to perform other functions while performing world reconstruction. Additionally, the processing of the reconstruction data by the components that "consume" that data can exacerbate the demand for computer resources.
[0044] Dynamic occlusion processing identifies portions of virtual objects that should not be rendered and / or displayed to the user because there are physical objects that block the user's view of where the virtual objects should be perceived by the user, and the relative positions between the physical objects and the virtual objects change over time. Occlusion processing for considering the user's hand can be particularly important for providing a desirable XR experience. However, the inventors recognized and appreciated the true value that occlusion processing specifically improved for the hand can provide a more realistic XR experience for the user. The XR system can generate a mesh for an object used in occlusion processing based on, for example, a graphic image obtained at a frame rate of about 5 frames per second (fps). However, that rate may not meet the speed of location changes between the hand and virtual objects behind the hand due to hand movement and / or head movement, e.g., exceeding 15 fps, 30 fps, or 45 fps.
[0045] Users of XR devices may interact with the device by performing gestures using their hands. Since the hands are used directly in user interactions, latency is important. The user's hands can move at high speed during interaction with the device, for example, faster than a user moves to scan the physical environment for world reconstruction. Further, the user's hands are closer to the XR device worn by the user. Thus, the relative position between the user's hand and the virtual object behind the hand is also sensitive to head movement. If the hand representation used for occlusion processing is not updated fast enough to keep up with these sources of relative motion, the occlusion processing will not be based on the hand's location and the occlusion processing will be inaccurate. If the virtual object behind the hand is not rendered correctly to appear occluded by the hand during hand movement and / or head movement, the XR scene will appear unrealistic to the user. The virtual object may appear on top of the hand as if the hand were transparent. Otherwise, the virtual object may not appear to be in its intended location. The hand may appear to have the color pattern of the virtual object, or other artifacts may appear. As a result, hand movement will break the user's immersion in the XR experience.
[0046] The inventors recognized and appreciated the true value that when an object occludes a virtual object, especially high computational requirements can be demanded when the object is the user's hand. However, the computational burden can be reduced by techniques that generate hand occlusion data at a high rate using low computational resources. The hand occlusion data may be generated by calculating a hand mesh from live depth sensor data obtained at a higher frequency than a graphic image. In some embodiments, the live depth sensor data may be obtained at a frame rate of at least 30 fps. To enable high-speed processing of the data, a small amount of data is processed to create a hand model for use in occlusion processing by masking the live depth data using a model in which the hand is represented by a plurality of segments simply identified from feature points. Further, to increase the accuracy of the occlusion processing, the hand occlusion data may be generated by predicting a change in the hand pose between the time of capture of the depth data and the time when the hand mesh will be used for occlusion processing. The hand mesh may be distorted to represent the hand in the predicted pose.
[0047] Techniques as described herein may be used with or separately from many types of devices, including wearable or portable devices with limited computational resources, for many types of scenarios, including providing cross-reality scenarios. In some embodiments, the techniques may be implemented by a service that forms part of an XR system.
[0048] Figures 1-2 illustrate such a scenario. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3-6 illustrate an exemplary AR system that may operate in accordance with the techniques described herein and includes one or more processors, memory, sensors, and a user interface.
[0049] Referring to FIG. 1, an outdoor AR scene 4 is depicted, and to the user of AR technology, a physical-world park-like setting 6 is visible, featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the user of AR technology also "sees" and perceives a robotic image 10 standing on the concrete platform 8 of the physical world and a flying comic-like avatar character 2 that appears to be an anthropomorphic bumblebee, although these elements (e.g., avatar character 2 and robotic image 10) do not exist within the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is difficult to produce AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements among other virtual or physical-world image elements.
[0050] Such an AR scene can be achieved using a system that includes a world reconstruction component that can construct and update a representation of the physical-world surface surrounding the user. This representation may be used for purposes such as occluding rendering, placing virtual objects in a physics-based interaction state, and for virtual character path planning and navigation, or for other operations for which information about the physical world is used. FIG. 2 depicts another example of an indoor AR scene 200 that shows an exemplary world reconstruction use case that includes visual occlusion 202, physics-based interaction 204, and environment inference 206, according to some embodiments.
[0051] The exemplary scene 200 is a living room having a wall, a bookshelf on one side of the wall, a floor lamp at the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology may also perceive virtual objects such as an image on the wall behind the sofa, a bird flying in through the door, a deer peeking out from the bookshelf, and a figurine in the form of a windmill placed on the coffee table. With respect to the image on the wall, AR technology requests information about objects and surfaces in the room such as lamp shapes, not only on the surface of the wall, which occludes the image and correctly renders the virtual object. With respect to the flying bird, AR technology renders the bird with realistic physics and requests information about all objects and surfaces around the room to avoid objects and surfaces or bounces from them in case of a collision with the bird. With respect to the deer, AR technology requests information about surfaces such as the floor or the coffee table and calculates where to place the deer. With respect to the windmill, the system may identify that it is an object separate from the table and may infer that it is movable, while the corner of the shelf or the wall may be inferred to be stationary. Such specificities may be used in inferences regarding parts of the scene that are used or updated in each of the various operations.
[0052] The scene may be presented to the user via a system that includes a plurality of components including a user interface that may stimulate one or more user senses including vision, sound, and / or touch. In addition, the system may include one or more sensors that may measure parameters of the physical part of the scene including the position and / or movement of the user within the physical part of the scene. Further, the system may include one or more computing devices with associated computer hardware such as memory. These components may be integrated within a single device or may be distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated within a wearable device.
[0053] Figure 3 depicts an AR system 302 configured to provide an experience of AR content that interacts with the physical world 306, according to some embodiments. The AR system 302 may include a display 308. In the illustrated embodiment, the display 308 may be worn by a user as part of a headset such that the user can wear the display across their eyes like a pair of goggles or glasses. At least a portion of the display may be transparent such that the user can observe the see-through reality 310. The see-through reality 310 may correspond to a portion of the physical world 306 within the current viewing perspective of the AR system 302 that the user can capture information about the physical world when wearing a headset that incorporates both the display and sensors of the AR system and that may correspond to the user's viewing perspective.
[0054] The AR content may also be presented on the display 308, overlaid on the see-through reality 310. To provide an accurate interaction between the AR content and the see-through reality 310 on the display 308, the AR system 302 may include sensors 322 configured to capture information about the physical world 306.
[0055] The sensors 322 may include one or more depth sensors that output a depth map 312. Each depth map 312 may have a plurality of pixels, each representing the distance to a surface within the physical world 306 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensors and may be used to create the depth map. Such depth maps may be updated as fast as the depth sensors can form new images, which may be hundreds or thousands of times per second. However, the data may be noisy, incomplete, and may have holes, shown as black pixels on the depth map.
[0056] The system may include other sensors such as an image sensor. The image sensor may obtain information that can be processed to represent the physical world in other ways. For example, the image may be processed in the world reconstruction component 316 to create a mesh representing the connected parts of the objects in the physical world. For example, metadata about such objects, including color and surface texture, may also be obtained using the sensor and stored as part of the world reconstruction.
[0057] The system may also obtain information about the user's head pose relative to the physical world. In some embodiments, the sensor 310 may include an inertial measurement unit (IMU) that can be used to calculate and / or determine the head pose 314. The head pose 314 for the depth map may indicate, for example, the current viewing point of the sensor that captures the depth map, with six degrees of freedom (6DoF), but the head pose 314 may be used for other purposes, such as relating the image information to a specific part of the physical world or relating the position of a display worn on the user's head to the physical world. In some embodiments, the head pose information may be derived in ways other than from the IMU, such as by analyzing objects in the image.
[0058] The world reconstruction component 316 may receive the depth map 312, the head pose 314, and any other data from the sensors and integrate that data into the reconstruction 318, which may appear as at least a single combined reconstruction. The reconstruction 318 may be more complete and less noisy than the sensor data. The world reconstruction component 316 may update the reconstruction 318 using a spatial and temporal average of sensor data over time from multiple viewpoints.
[0059] The reconstruction 318 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same part of the physical world or different parts of the physical world. In the illustrated embodiment, on the left side of the reconstruction 318, a part of the physical world is presented as a global surface, and on the right side of the reconstruction 318, a part of the physical world is presented as a mesh.
[0060] The reconstruction 318 may be used for AR functions, such as producing a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation may change as the user moves or as objects within the physical world change. The side of the reconstruction 318 may be used, for example, by a component 320 that produces a changing global surface representation within world coordinates, which may also be used by other components.
[0061] AR content may be generated by an AR application 304 or the like based on this information. The AR application 304 may be, for example, a game program that implements one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. These functions may be implemented by querying the reconstruction 318 produced by the world reconstruction component 316 for data in different formats. In some embodiments, the component 320 may be configured to output an update when the representation within the area of interest of the physical world changes. The area of interest may be set to approximate a part of the physical world in the vicinity of the user of the system, such as a part within the user's field of view, or projected (predicted / determined) to occur within the user's field of view.
[0062] The AR application 304 may use this information to generate and update AR content. The virtual portion of the AR content may be presented on the display 308 in combination with see-through reality 310 to create an immersive user experience.
[0063] In some embodiments, the AR experience may be provided to the user through a wearable display system. FIG. 4 illustrates an example of a wearable display system 80 (hereinafter referred to as "system 80"). System 80 includes a head-mounted display device 62 (hereinafter referred to as "display device 62") and various mechanical and electronic modules and systems to support the functions of display device 62. Display device 62 may be coupled to a frame 64, which is wearable by a display system user or viewer 60 (hereinafter referred to as "user 60") and is configured to position display device 62 in front of the eyes of user 60. According to various embodiments, display device 62 may be a sequential display. Display device 62 may be monocular or binocular. In some embodiments, display device 62 may be an example of display 308 in FIG. 3.
[0064] In some embodiments, a speaker 66 is coupled to the frame 64 and positioned proximate to the outer ear canal of user 60. In some embodiments, another speaker (not shown) is positioned adjacent to the other outer ear canal of user 60 to provide stereo / adjustable sound control. Display device 62 is operably coupled to a local data processing module 70 by means such as a wired conductor or wireless connectivity 68, which may be mounted in various configurations, such as fixed to the frame 64, fixed to a helmet or cap worn by user 60, built into headphones, or alternatively removably attached to user 60 (e.g., in a backpack configuration, in a belt attachment configuration).
[0065] The local data processing module 70 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which may be utilized to assist in data processing, caching, and storage. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope (e.g., operatively coupled to frame 64 or otherwise attachable to user 60), and / or b) data that may potentially be obtained and / or processed using the remote processing module 72 and / or the remote data repository 74 for passage to the display device 62 after processing or reading. The local data processing module 70 may be operatively coupled to the remote processing module 72 and the remote data repository 74, respectively, via communication links 76, 78, such as a wired or wireless communication link, such that these remote modules 72, 74 are operatively coupled to each other and available as resources to the local processing and data module 70. In some embodiments, the world reconstruction component 316 in FIG. 3 may be implemented, at least in part, within the local data processing module 70. For example, the local data processing module 70 may be configured to execute computer-executable instructions to generate a physical world representation, at least in part, based on at least a portion of the data.
[0066] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 70 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computing budget of the module 70 but will enable a smaller device. In some embodiments, the world reconstruction component 316 may use a computing budget less than a single ARM core so that the remaining computing budget of the single ARM core can be accessed for other uses, such as extracting a mesh, etc., and generate a physical world representation in real time on an ad hoc space.
[0067] In some embodiments, the remote data repository 74 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed within the local data processing module 70, enabling completely autonomous use from the remote module. World reconstruction may be stored, for example, in whole or in part within this repository 74.
[0068] In some embodiments, the local data processing module 70 is operably coupled to the battery 82. In some embodiments, the battery 82 is a removable power source, such as a commercially available battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during the non-operating time of the system 80 and a removable battery, so that the user 60 can operate the system 80 for a longer time period without having to connect to a power source to charge the lithium-ion battery or shut down the system 80 to replace the battery.
[0069] Figure 5A illustrates a user 30 wearing an AR display system that renders AR content as the user 30 moves through the physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at position 34, and the AR display system records surrounding information of a passable world (e.g., a digital representation of real objects in the physical world that can be memorized and updated as the real objects in the physical world change) with respect to the pose relationship to the mapped features or the position 34 such as the directional audio input. The position 34 is aggregated into data input 36 and processed by a passable world module 38 that can be implemented, at least, by processing on the remote processing module 72 of FIG. 4, for example. In some embodiments, the passable world module 38 may include a world reconstruction component 316.
[0070] The passable world module 38 determines the location and manner in which the AR content 40 can be placed within the physical world such that the AR content 40 is determined from the data input 36. The AR content is "placed" within the physical world by presenting both the physical world and the representation of the AR content via the user interface, and the AR content is rendered as if it were interacting with the objects in the physical world, and the objects in the physical world are presented as if the AR content were obscuring the user's view of those objects when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 42 (e.g., a table) from the reconstruction (e.g., reconstruction 318) and determining the shape and position of the AR content 40. As an example, the fixed element may be a table, and the virtual content may be positioned to appear as if it were on that table. In some embodiments, the AR content may be placed within the structures in the field of view 44, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be placed with respect to the mapped mesh model 46 of the physical world.
[0071] As depicted, the fixed element 42 serves as a substitute for any fixed element within the physical world, which may be stored within the passable world module 38 such that the user 30 can perceive the content on the fixed element 42 without the system having to map it to the fixed element 42 each time it is visible to the user 30. The fixed element 42 may thus be a mapped mesh model from a previous modeling session or determined from a separate user, provided that it may be stored on the passable world module 38 for future reference by multiple users. Thus, the passable world module 38 enables the user 30's device to recognize the environment 32 from a previously mapped environment and display AR content without first mapping the environment 32, saving calculation processes and cycles and avoiding latency for any rendered AR content.
[0072] The mapped mesh model 46 of the physical world may be created by the AR display system, interacts with the AR content 40, and the appropriate surfaces and metrics for displaying it can be mapped and stored within the passable world module 38 for future retrieval by the user 30 or other users without the need for remapping or modeling. In some embodiments, the data input 36 is an input such as a geographical location, user identification, and current activity, and provides the passable world module 38 with the fixed element 42 of one or more available fixed elements, the AR content 40 last placed on the fixed element 42, and whether to display that same content (such AR content being "persistent" content regardless of whether the user is viewing a particular passable world model).
[0073] Even in embodiments where an object is considered to be fixed, the passable world module 38 may be updated at any time to account for the variability of the physical world. The model of the fixed object may be updated very infrequently. Other objects within the physical world may be considered to be moving or otherwise not fixed. To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects at a much higher frequency than the frequency at which the fixed objects are updated. To enable accurate tracking of all objects within the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0074] FIG. 5B is a schematic diagram of the viewing optics assembly 48 and associated components. In some embodiments, two eye tracking cameras 50 directed towards the user's eyes 49 detect metrics of the user's eyes 49, such as eye shape, eyelid occlusion, pupil direction, and glints on the user's eyes 49. In some embodiments, one of the sensors is a depth sensor 51, such as a time-of-flight sensor, that emits signals into the world and detects the reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor may, for example, quickly determine whether an object is entering the user's field of view as a result of either the movement of those objects or a change in the user's pose. However, information about the position of an object within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may, for example, be obtained from a stereoscopic image sensor or a plenoptic sensor.
[0075] In some embodiments, the world camera 52 records a view that exceeds the peripheral field of view, maps the environment 32, and detects inputs that can affect the AR content. In some embodiments, the world camera 52 and / or the camera 53 may be grayscale and / or color image sensors, which may output grayscale and / or color image frames at fixed time intervals. The camera 53 may further capture a physical world image within the user's field of view at a specific time. The pixels of the frame-based image sensor may be sampled iteratively even if their values are invariant. The world camera 52, the camera 53, and the depth sensor 51 each have individual fields of view of 54, 55, and 56, respectively, and collect and record data from a physical world scene such as the physical world environment 32 depicted in FIG. 5A.
[0076] The inertial measurement unit 57 may determine the movement and orientation of the visual optical system assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 51 is operably coupled to the eye-tracking camera 50 as a confirmation of the focusing adjustment measured with respect to the actual distance that the user's eye 49 is looking at.
[0077] The visual optical system assembly 48 may include some of the components illustrated in FIG. 5B, and it should be understood that it may include components instead of or in addition to the illustrated components. In some embodiments, for example, the visual optical system assembly 48 may include two world cameras 52 instead of four. Alternatively, or in addition, the cameras 52 and 53 do not need to capture visible light images of their full fields of view. The visual optical system assembly 48 may include other types of components. In some embodiments, the visual optical system assembly 48 may include one or more dynamic vision sensors (DVSs), and its pixels may respond asynchronously to relative changes in light intensity that exceed a threshold.
[0078] In some embodiments, the visual optical system assembly 48 may not include the time-of-flight based depth sensor 51. In some embodiments, for example, the visual optical system assembly 48 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of the incident light, from which depth information can be determined. For example, a plenoptic camera may include an image sensor overlaid with a transmissive diffractive mask (TDM). Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle sensing pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a source of depth information instead of, or in addition to, the depth sensor 51.
[0079] Also, it should be understood that the component configuration in FIG. 5B is illustrated as an example. The visual optical system assembly 48 may include components with any suitable configuration, which may be set to provide the user with a practical maximum field of view for a particular set of components. For example, if the visual optical system assembly 48 has one world camera 52, the world camera may be installed within the central region of the visual optical system assembly instead of on the side.
[0080] Information from sensors within the visual optics assembly 48 may be coupled to one or more of the processors within the system. The processor may generate data that can be rendered to cause a user to perceive virtual content that interacts with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in a single scene by modulating the opacity of a display device that the user views through the physical world. The opacity may be controlled to create the appearance of the virtual object and to block from view objects in the physical world that are occluded by the virtual object. In some embodiments, the image data may be modified to be perceived by the user as if the virtual content is realistically interacting with the physical world when viewed through the user interface (e.g., clipping the content to account for occlusion), and may include only the virtual content. Regardless of how the content is presented to the user, the model of the physical world is required to accurately calculate the characteristics of the virtual objects that can be affected by the physical objects, including the shape, position, motion, and visibility of the virtual objects. In some embodiments, the model may include a reconstruction of the physical world, such as the reconstruction 318.
[0081] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected by multiple users, which may be aggregated within a computing device remote from all users (and may be "in the cloud").
[0082] The model may be created, at least in part, by a world reconstruction system, such as the world reconstruction component 316 of FIG. 3, described in more detail in FIG. 6. The world reconstruction component 316 may include a perception module 160 that can generate, update, and store a representation for a portion of the physical world. In some embodiments, the perception module 160 may represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world and may include surface information indicating whether a surface exists within the volume represented by the voxel. The voxels may be assigned a value indicating whether the corresponding volume has been determined to be empty, has been determined to include the surface of a physical object, or has not yet been measured using the sensor and thus its value is unknown. It should be understood that the value may be stored in computer memory in any suitable manner, including indicating that voxels determined to be empty or unknown need not be explicitly stored and that the voxel values do not store information about voxels determined to be empty or unknown.
[0083] In addition to generating information for the persistent world representation, the perception module 160 may identify and output an indication of a change in the area surrounding the user of the AR system. Such an indication of a change may trigger an update to the volumetric data stored as part of the persistent world, or may trigger other functions such as triggering component 304 to generate and update AR content.
[0084] In some embodiments, the perception module 160 may identify changes based on a signed distance function (SDF) model. The perception module 160 may be configured to receive sensor data such as, for example, a depth map 160a and a head pose 160b, and then fuse the sensor data into the SDF model 160c. The depth map 160a may directly provide SDF information, and the image may be processed to become SDF information. The SDF information represents the distance from the sensor used to capture that information. Since those sensors may be part of the wearable unit, the SDF information may represent the physical world from the line of sight of the wearable unit and thus the user's line of sight. The head pose 160b may enable the SDF information to be associated with voxels within the physical world.
[0085] In some embodiments, the perception module 160 may generate, update, and store a representation for a portion of the physical world within the perception range. The perception range may be determined at least in part based on the reconstruction range of the sensor, which may be determined at least in part based on the limits of the observation range of the sensor. As a specific example, an active depth sensor operating using active IR pulses may reliably operate over a certain range of distances and create the observation range of the sensor, which may be several centimeters or several tens of centimeters to several meters.
[0086] The world reconstruction component 316 may include additional modules that may interact with the perception module 160. In some embodiments, the persistent world module 162 may receive a representation for the physical world based on data obtained by the perception module 160. The persistent world module 162 may also include various formats of the representation of the physical world. For example, volumetric metadata 162b such as voxels may be stored along with a mesh 162c and a plane 162d. In some embodiments, other information such as a depth map may also be stored.
[0087] In some embodiments, the perception module 160 may include modules that generate a representation for the physical world in various formats, such as, for example, mesh 160d, plane, and semantics 160e. These modules may generate the representation based on data within the perception range of one or more sensors at the time the representation is generated, data captured at previous times, and information within the persistent world 162. In some embodiments, these components may act on depth information captured using a depth sensor. However, the AR system may include a vision sensor and generate such a representation by analyzing monocular or binocular vision information.
[0088] In some embodiments, these modules may act on regions of the physical world. Those modules may be triggered to update a sub-region of the physical world when the perception module 160 detects a change in the physical world within that sub-region. Such changes may be detected, for example, by detecting a new surface within the SDF model 160c or by other criteria such as changing the values of a sufficient number of voxels representing the sub-region.
[0089] The world reconstruction component 316 may include a component 164 that can receive a representation of the physical world from the perception module 160. Information about the physical world may be pulled by these components, for example, in accordance with usage requests from an application. In some embodiments, the information may be pushed to the usage components via an indication of a change in a pre-identified region or a change in the physical world representation within the perception range. The component 164 may include, for example, a game program and other components that perform processing for visual occlusion, physics-based interactions, and environmental inference.
[0090] In response to a query from component 164, the perception module 160 may transmit a representation for the physical world in one or more formats. For example, when component 164 indicates that the use is for visual occlusion or physics-based interaction, the perception module 160 may transmit a surface representation. When component 164 indicates that the use is for environmental inference, the perception module 160 may transmit a mesh, plane, and semantics of the physical world.
[0091] In some embodiments, the perception module 160 may include a component that formats information and provides it to component 164. An example of such a component may be a raycasting component 160f. The using component (e.g., component 164) may query for information about the physical world from a particular perspective, for example. The raycasting component 160f may select from one or more representations of the physical world data within the field of view from that perspective.
[0092] Information about the physical world may also be used for occlusion processing. That information may be used by a visual occlusion component 164a, which may be part of a world reconstruction component 316. The visual occlusion component 164a may supply, for example, an application with information indicating portions of a visual object that are occluded by a physical object. Alternatively, or in addition, the visual occlusion component 164a may provide an application with information about a physical object, which may use that information for occlusion processing. As described above, accurate information about the hand's position is important for occlusion processing. In one embodiment, as described herein, the visual occlusion component 164a may maintain a hand model and, in response to a request by an application, provide that model to the application. FIG. 7 illustrates an example of such processing, which may be implemented across one or more of the components illustrated in FIG. 6 or, in some embodiments, by different or additional components.
[0093] FIG. 7 depicts an AR system 700 configured to generate a hand mesh in real time for dynamic occlusion processing according to some embodiments. The AR system 700 may be implemented on an AR device. The AR system 700 may include a data collection portion 702 configured to use sensors on the AR device to capture information about the pose of a user (e.g., head pose, hand pose, and the like) wearing the AR device (e.g., a display device 62) and the scene. The information about the scene may include depth information indicating the distance between the AR device and physical objects in the scene.
[0094] The data collection part 702 includes a hand tracking component. The hand tracking component may process sensor data such as depth and image information, and detect one or both hands in the scene. Other sensor data may also be processed to detect one or both hands in the scene. When detected, one or both hands can be represented in a sparse way by a set of feature points or the like. The feature points can represent, for example, joints, fingertips, or other boundaries of hand segments. The information collected or generated by the data collection part 702 may be passed to the hand meshing part 704 to be used to generate a richer model of one or both hands, such as a mesh, based on the sparse representation.
[0095] The hand meshing part 704 is configured to calculate the hand mesh of the detected one or both hands and update the hand mesh in real time as the posture changes and / or the hand moves.
[0096] The AR system 700 may include an application 706 configured to receive the hand mesh from the hand meshing part 704 and render one or more virtual objects in the scene. In some embodiments, the application 706 may receive occlusion data from the hand meshing part 704. In some embodiments, the occlusion data may indicate the portions of the virtual objects occluded by one or both hands. In some embodiments, for example, in the illustrated embodiment, the occlusion data may be a model of one or both hands, from which the application 706 may calculate the occlusion data. As a specific example, the occlusion data may be the hand mesh of one or both hands received from the hand meshing part 704.
[0097] Figure 8 is a flowchart illustrating a method 800 for generating a hand mesh in real time for dynamic occlusion, according to some embodiments. In some embodiments, method 800 may be implemented by one or more processors within AR system 700. Method 800 may begin when the hand meshing component 704 of AR system 700 receives a query regarding data related to one or both hands in the scene from application 706 of AR system 700 (act 802). Method 800 may include a step (act 804) of detecting one or both hands in the scene based on information about the scene captured by data collection portion 702 of AR system 700.
[0098] Once one or both hands are detected, method 800 may include a step (act 806) of calculating one or more models of one or both hands based on the information about the scene. The one or more models of one or both hands may be sparse and may indicate the positions of feature points on the hand rather than the surface. Those feature points may represent joints or the end portions of hand segments. The feature points may be recognized from sensor data about one or both hands, including, for example, stereo images of one or both hands. Depth information, and in some instances, monocular images of one or both hands, may be used alternatively or in addition to identify the feature points.
[0099] U.S. Provisional Patent Application No. 62 / 850,542, entitled “Hand Pose Estimation,” describes exemplary methods and apparatuses for obtaining information about the position and pose of a hand and modeling the hand based on the obtained information. A copy of the application version of U.S. Application No. 62 / 850,542 is attached hereto and is hereby incorporated by reference in its entirety for all purposes. Techniques as described in that application may be used to construct a sparse model of the hand.
[0100] In some embodiments, one or more models of one or both hands may be calculated based on information about the scene captured by the sensors of the AR device. Examples of information about the scene include the exemplary image of FIG. 9A captured by one sensor corresponding to one eye, the two exemplary images of FIG. 9B captured by two sensors corresponding to the left and right eyes, and the exemplary depth image of FIG. 9C that can be obtained at least in part from the image of FIG. 9A or the image of FIG. 9B.
[0101] In some embodiments, one or more models of one or both hands may include a plurality of feature points of the hand, which may indicate points on the segment of the hand. Some of the feature points may correspond to the joints of the hand and the fingertips of the hand. FIGS. 9E and 9F depict schematic diagrams illustrating an exemplary 8 - feature - point model of the hand and an exemplary 22 - feature - point model of the hand, respectively.
[0102] In some embodiments, the feature points of the hand model may be used to determine the contour of the hand. FIG. 9D depicts an exemplary contour of the hand, which may be determined based on the feature points of the hand model. For example, adjacent feature points may be connected by lines, as schematically illustrated in FIGS. 9E and 9F, and the contour of the hand may be shown as the distance from the line. The distance from the line may be determined from the image of the hand, information about the human anatomical structure, and / or other information. It should be understood that once the feature points of the hand are identified, the hand model may be updated at a later time using information previously obtained about the hand. For example, the length of the line connecting the feature points may not change.
[0103] The sparse hand model may be used to select a limited amount of data from which a richer model of the hand, including surface information, may be constructed, for example. In some embodiments, the selection may be made by using the hand contour to mask additional data such as depth data. Thus, method 800 may include masking depth information (act 808) indicative of the distance between the AR device and physical objects in the scene using one or more models of one or both hands.
[0104] FIG. 10 depicts a flowchart illustrating details of the step of masking depth information (act 808) using one or more models of one or both hands, according to some embodiments. Act 808 may include filtering out depth information outside the contour of one or more models of one or both hands (act 1002). This act results in removing depth information associated with physical objects other than the hand in the scene. Act 808 may include generating a depth image of one or both hands (act 1004) based on the filtered depth information.
[0105] The depth image may include pixels that may indicate the distance to points on one or both hands, respectively. In some embodiments, the depth information may be captured such that the depth information is not captured for all surfaces of one or both hands. For example, the depth information may be captured using an active IR sensor. If the user is wearing a ring with a dark stone, for example, the IR may not reflect from the dark ring such that there will be holes in the collected insufficient information. Act 808 may include filling holes in the depth image (act 1006). In some embodiments, the holes may be filled by identifying the holes in the depth image and generating stereo depth information corresponding to the identified regions from the stereo camera of the AR device. In some embodiments, the holes may be filled by identifying the holes in the depth image and accessing surface information corresponding to the identified regions from one or more 3D models of one or both hands.
[0106] Optionally, the mesh of one or both hands may be calculated as a plurality of sub-meshes, and each sub-mesh represents a segment of one or both hands. The segments may correspond to segments bounded by feature points. Since many of those segments are bounded by joints, the segments correspond to parts of the hand that can move independently of at least other segments of the hand so that the hand mesh can be updated quickly by updating the sub-meshes associated with the segments that have moved since the last hand mesh calculation. In such an embodiment, act 808 may include identifying feature points in the depth image that correspond to feature points of one or more models of one or both hands (act 1008). Segments of the hand separated by feature points may be calculated. Act 808 may include associating a portion of the depth image with a segment of the hand separated by the identified feature points in the depth image (act 1010).
[0107] Method 800 may include calculating a mesh of one or both hands based on depth information masked for one or more models of one or both hands (act 810). The mesh of one or both hands may be a representation of one or both hands showing the surface of one or both hands. FIG. 9G depicts a calculated dense hand mesh according to some embodiments. However, the present application is not limited to the step of calculating a dense hand mesh. In some embodiments, a sparse hand mesh may be sufficient for dynamic occlusion.
[0108] In some embodiments, the one-handed or two-handed mesh may be calculated from depth information. For example, the mesh may often be a set of regions represented as triangles that represent a part of the surface. Such regions may be identified by grouping adjacent pixels in a depth image where the difference in distance with respect to them is less than a threshold, indicating that the pixels are likely to be on the same surface. One or more triangles bounding such regions of pixels may be identified and added to the mesh. However, other techniques for forming the mesh from depth information may be used.
[0109] The step of calculating the one-handed or two-handed mesh may include the step of updating the one-handed or two-handed mesh in real time as the relative location between the AR device and the one hand or two hands changes. There may be a latency between the time when scene information used to calculate the one-handed or two-handed mesh is captured and the time when one or more calculated hand meshes are used, for example, by an application, for example, to render content. In some embodiments, the movement of the one-handed or two-handed segments may be tracked such that the future positions of those segments of the one hand or two hands can be projected / predicted. The one-handed or two-handed mesh may be distorted to conform to the projected location of the segment at the time when the virtual object processed using the one-handed or two-handed mesh will be rendered.
[0110] FIG. 11 depicts a flowchart illustrating the details of the step (act 810) of calculating in real time a one-handed or two-handed mesh based on depth information masked with respect to hand segmentation, according to some embodiments. Act 810 may include the step (act 1102) of determining a time t when the hand meshing portion 704 receives a query from the application 706 regarding data related to one hand or two hands.
[0111] Act 810 may include the step of predicting the waiting time n from the query received from application 706 at time t (Act 1104). Act 810 may include the step of predicting the posture (e.g., hand posture) at the time of query time t + waiting time n (Act 1106). Act 810 may include the step of distorting the mesh of one or both hands using the posture predicted at the time of query time t + waiting time n (Act 1108). In some embodiments, the step of predicting the hand posture may include the step of predicting the movement of feature points of one or both hands in the depth image at the time of query time t + waiting time n. Such prediction may be performed based on the step of tracking the position of the feature points over time. Such tracking enables the determination of motion parameters such as speed or acceleration. The projection of the position may be performed based on the extrapolation from the previous position to the future position assuming that the determined motion parameters remain the same. Alternatively, or in addition, the projection may be determined using a Kalman filter or a similar projection technique.
[0112] In some embodiments, the step of distorting the mesh of one or both hands using the predicted posture may include the step of distorting the portion of the mesh of one or both hands corresponding to a subset of the hand segments that are between the feature points predicted to change at the query time t + waiting time n.
[0113] When distorting the previously calculated mesh to represent one or both hands at time t + waiting time n, multiple factors may be considered. The value of n may reflect, for example, the time required for the process of distorting one or more meshes and for using one or more meshes when the application renders the object. The value may be estimated or measured from the operation, structure, or test of the software. As another example, the value may be dynamically determined based on the step of adjusting the previously established waiting time based on measuring the waiting time during use or based on the processing load at the time when the request for one or more meshes is made.
[0114] When distorting one or more meshes, the amount of distortion may be based on the time when the data used to form the mesh of one or both hands was captured and the waiting time until one or more meshes will be used. In some embodiments, the mesh of one or both hands may be created other than in response to a request from an application. For example, once the application indicates that it is configured for occlusion processing, such as by making a call through an API, the AR system may periodically calculate one or more updated hand meshes. Alternatively, the hand tracking process may be continuously activated using a certain amount of the system's computational resources. In any case, the mesh of one or both hands may be updated relatively frequently, such as at least 30 times per second. Note that there may be a delay between when the data is captured to create the mesh and when a request for the mesh is received, and this delay may also be considered when distorting the hand mesh.
[0115] Method 800 may include supplying (act 812) the mesh of one or both hands to application 706 such that the application renders portions of the virtual object that are not occluded by the mesh of one or both hands.
[0116] Although some aspects of some embodiments have been described so far, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art.
[0117] As an example, embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applicable within an MR environment, or more generally, within other XR environments and VR environments.
[0118] As another example, the embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0119] Such variations, modifications, and improvements are intended to be part of the present disclosure and are intended to be within the spirit and scope of the present disclosure. Further, while the advantages of the present disclosure are shown, it should be understood that not all embodiments of the present disclosure include all of the advantages described. Some embodiments may not implement any of the features described as advantageous herein and in some instances. Therefore, the foregoing description and drawings are merely examples.
[0120] The foregoing embodiments of the present disclosure can be implemented in any of a number of ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided within a single computer or distributed among multiple computers. Such processors can include, by way of example, one or more processors within an integrated circuit component, including commercially available integrated circuit components known in the art such as a CPU chip, a GPU chip, a microprocessor, a microcontroller, or a coprocessor, etc. In some embodiments, the processor may be implemented within a custom circuit such as an ASIC or within a semi-custom circuit resulting from configuring a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores such that a subset of one or those cores can constitute the processor. However, the processor may be implemented using circuitry in any suitable format.
[0121] Furthermore, it should be understood that the computer can be embodied in any of several forms such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. In addition, the computer may be embodied in a device that is generally not considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or fixed electronic device.
[0122] In addition, the computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound generating device for audible presentation of output. Examples of input devices that can be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, the computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output device is shown physically separate from the computing device. However, in some embodiments, the input and / or output device may be physically integrated within the same unit as the processor or other elements of the computing device. For example, a keyboard can be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output device may be completely disconnected from the computing device and functionally integrated through a wireless connection.
[0123] Such computers may be interconnected by one or more networks in any suitable form, including forms such as a local area network or wide area network, such as a corporate network or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0124] In addition, the various methods and processes outlined herein may be encoded as software that is executable on one or more processors employing any one of a variety of operating systems or platforms. Additionally, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may also be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine.
[0125] In this aspect, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy (registered trademark) disks, compact disk (CD), optical disk, digital video disk (DVD), magnetic tape, flash memory, circuit configurations within a field programmable gate array or other semiconductor device, or other tangible computer storage media) encoded with one or more programs that, when executed on one or more computers or other processors, perform a method of implementing the various embodiments of the present disclosure discussed above. As will be apparent from the foregoing examples, a computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such a computer-readable storage medium or media can be transportable such that the one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" includes only computer-readable media that can be regarded as a manufactured (i.e., a manufactured article) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.
[0126] As described above, the terms "program" or "software" are used herein in a general sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure. Additionally, according to one aspect of the present embodiment, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion among several different computers or processors so as to implement various aspects of the present disclosure.
[0127] Computer-executable instructions may take many forms, such as program modules, which are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0128] Also, data structures may be stored on a computer-readable medium in any suitable form. For simplicity of illustration, a data structure may be shown to have fields that are related through locations within the data structure. Such relationships may likewise be achieved by allocating storage for fields that accompany locations within the computer-readable medium that convey the relationships between fields. However, any suitable mechanism, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements, may be used to establish relationships between the information within the fields of a data structure.
[0129] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and thus, its use is not limited to the details and arrangements of the components described in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined with aspects described in other embodiments in any manner.
[0130] Also, the present disclosure may be embodied as a method, for which examples are provided. The acts performed as part of the method may be ordered in any suitable manner. Thus, in the illustrative embodiments, although shown as consecutive acts, embodiments may be constructed in which acts are performed in a different order than illustrated, including performing some acts simultaneously.
[0131] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not by itself imply any priority, precedence, or order of one claim element over another, or the temporal order in which acts of a method are performed, but the ordinal terms are used only as labels to distinguish one claim element having a certain name from another element having the same name (for the purpose of using the ordinal terms).
[0132] Also, the phrases and terminology used herein are for the purpose of explanation and should not be regarded as limiting. The use of "comprising," "including," "having," "containing," "involving," and variations thereof herein means including the items recited thereafter and their equivalents and additional items.
Claims
【Claim 1】 The invention described in this specification.