Fast hand meshing for dynamic occlusion
The method of fast hand meshing in XR systems addresses computational and latency issues by dynamically occluding virtual objects with hands, enhancing the realism and immersion of XR experiences.
Patent Information
- Application Number
- JP2021576046
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-06-25
AI Technical Summary
Existing XR systems face challenges in dynamically occluding virtual objects with hands due to high computational demands and latency issues, leading to unrealistic XR experiences when hand movements are not accurately accounted for in occlusion processing.
A method for fast hand meshing that involves capturing depth information from sensors, detecting hands, calculating a hand model, masking depth information using the model, and updating a hand mesh in real time to accurately occlude virtual objects, using techniques that reduce computational burden and enhance processing speed.
Enables a more realistic XR experience by ensuring accurate rendering of virtual objects behind hands, reducing latency and computational demands, and maintaining user immersion during hand interactions.
Smart Images

Figure 0007711002000001 
Figure 0007711002000002 
Figure 0007711002000003
Abstract
Description
Technical Field
[0001] This application generally relates to a cross-reality system that uses 3D world reconstruction to render scenes.
Background Art
[0002] A computer can create an X Reality (XR or cross-reality) environment that controls a human user interface and in which part or all of the XR environment is generated by the computer as perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can describe virtual objects that can be rendered, for example, so that a user can perceive or sense them as part of the physical world and interact with the virtual objects. The user can experience these virtual objects as a result of data being rendered and presented through a user interface device such as a head-mounted display device. The data can be controlled to be displayed for the user to see, or reproduced for the user to hear, or to control an audio that can be reproduced for the user to hear, or to control a haptic (or tactile) interface, enabling the user to experience a touch sensation that the user perceives or senses as the user feels the virtual object.
[0003] XR systems can be useful for many applications spanning the fields of scientific visualization, medical training, engineering design, and prototyping, teleoperation and telepresence, and personal entertainment. AR and MR, in contrast to VR, include one or more objects in relation to real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment when using XR systems and also expands the possibilities for various applications by presenting realistic and easily understandable information about how the physical world can be altered.
[0004] The XR system can represent the physical world around the user of the system as a "mesh". The mesh can be represented by a plurality of interconnected triangles. Each triangle has edges that join points on the surface of an object within the physical world such that each triangle represents a part of the surface. Information about parts of the surface, such as color, texture, or other properties, can be stored associated within the triangle. In operation, the XR system processes image information to detect points and surfaces so as to create or update the mesh.
Summary of the Invention
Means for Solving the Problems
[0005] Aspects of the present application relate to methods and apparatus for fast hand meshing for dynamic occlusion. The techniques described herein may be used together, separately, or in any suitable combination.
[0006] Some embodiments relate to a method of operating a computing system and reconstructing a hand to dynamically occlude virtual objects. The method includes receiving, from an application that renders a virtual object within a scene, a query regarding data related to a hand within the scene; capturing, from a device worn by a user, information of the scene, the device comprising one or more sensors, the information of the scene including depth information indicating a distance between the device worn by the user and a physical object within the scene; detecting whether the physical object within the scene includes the hand; when the hand is detected, calculating a hand model, at least in part, based on the information of the scene; masking the depth information indicating a distance between the device worn by the user and the physical object within the scene, using the hand model; calculating a hand mesh based on the depth information masked with respect to the hand model, the calculating step including updating the hand mesh in real time as a relative location between the device and a change of the hand; and supplying the hand mesh to the application such that the application renders a portion of the virtual object not occluded by the hand mesh.
[0007] In some embodiments, the hand model includes a plurality of feature points of the hand indicating points on segments of the hand.
[0008] In some embodiments, at least some of the plurality of feature points of the hand correspond to hand joints and fingertips of the hand.
[0009] In some embodiments, the method further includes determining a hand contour based on a plurality of feature points, and masking depth information indicating a distance between a device worn by a user and a physical object in a scene using a hand model. The step of masking the depth information includes filtering and removing depth information outside the contour of the hand model, and generating a depth image of the hand based at least in part on the filtered depth information, the depth image including a plurality of pixels, each pixel indicating a distance to a point on the hand.
[0010] In some embodiments, the step of filtering and removing depth information outside the contour of the hand model includes removing depth information associated with a physical object in the scene.
[0011] In some embodiments, the step of masking depth information indicating a distance between a device worn by a user and a physical object in a scene using a hand model includes associating a portion of the depth image with a hand segment, and the step of updating the hand mesh in real time includes selectively updating a portion of the hand mesh representing a suitable subset of the hand segment.
[0012] In some embodiments, the method further includes filling holes in the depth image before calculating the hand mesh.
[0013] In some embodiments, the step of filling holes in the depth image includes generating stereo depth information from a stereo camera of the device, the stereo depth information corresponding to a region of the holes in the depth image.
[0014] In some embodiments, the step of filling holes in the depth image includes accessing surface information from a 3D model of the hand, the surface information corresponding to a region of the holes in the depth image.
[0015] In some embodiments, the step of calculating a hand mesh based on depth information masked for a hand model includes predicting a latency n from a query received at time t regarding data related to a hand in a scene from an application that renders virtual objects in the scene, predicting a hand pose at time t + latency n, and distorting the hand mesh using the pose predicted at time t + latency n.
[0016] In some embodiments, the depth information indicating the distance between a device worn by a user and physical objects in a scene includes a sequence of depth images at a frame rate of at least 30 frames per second.
[0017] Some embodiments relate to an electronic system portable by a user. The electronic system includes a device worn by the user. The device includes a display configured to render virtual objects and one or more sensors configured to capture information about the head pose of the user wearing the device and a scene including one or more physical objects. The scene information includes depth information indicating the distance between the device and the one or more physical objects. The electronic system includes a hand meshing component configured to execute computer-executable instructions to detect a hand in the scene, calculate a hand mesh of the detected hand, and update the hand mesh in real time as the head pose changes and / or the hand moves, and an application configured to execute computer-executable instructions to render virtual objects in the scene and receive, from the hand meshing component, the hand mesh and portions of virtual objects occluded by the hand.
[0018] In some embodiments, the hand meshing component is configured to calculate a hand mesh by identifying feature points on the hand, calculating segments between the feature points, selecting information based on the proximity to one or more of the calculated segments from depth information, and calculating a mesh representing at least a portion of the hand mesh based on the selected depth information.
[0019] In some embodiments, the depth information includes a plurality of pixels, each of the plurality of pixels representing a distance to an object in the scene. The step of calculating the mesh includes grouping adjacent pixels that represent differences at distances below a threshold.
[0020] Some embodiments relate to a method of operating an AR system to render virtual objects in a scene including physical objects. The AR system includes at least one sensor and at least one processor. The method includes capturing information about the scene using at least one sensor, the information about the scene including depth information indicating distances to physical objects in the scene, processing the captured information using at least one processor to detect a hand in the scene and calculate points on the hand, selecting a subset of the depth information based on the proximity to the calculated points on the hand, and calculating a representation of the hand based on the selected depth information, the representation of the hand indicating the surface of the hand.
[0021] In some embodiments, the method further includes storing the calculated representation of the hand and continuously processing the captured information to update the stored representation of the hand.
[0022] In some embodiments, the step of calculating a hand representation includes calculating one or more parameters of the hand movement based on the captured information, projecting the hand position at a future time based on a latency associated with the step of rendering a virtual object using the calculated hand representation based on the one or more parameters of the movement, and morphing the calculated hand representation to represent the hand at the projected position.
[0023] In some embodiments, the method further includes rendering a selected portion of the virtual object based on the hand representation, the selected portion representing a portion of the virtual object not occluded by the hand.
[0024] In some embodiments, the depth information includes a depth map including a plurality of pixels each representing a distance. Based on the selected depth information, the step of calculating a hand representation includes identifying a group of pixels representing a surface segment.
[0025] In some embodiments, the step of calculating a hand representation includes defining a mesh representing the hand based on the identified group of pixels.
[0026] In some embodiments, the step of defining a mesh includes identifying triangular regions corresponding to the identified surface segments.
[0027] The foregoing summary is provided by way of example and not intended to be limiting. The present invention provides, for example, the following. (Item 1) A method of operating a computing system and reconstructing a hand to dynamically occlude a virtual object, the method comprising: Receiving, from an application that renders a virtual object in a scene, a query regarding data related to a hand in the scene; Capturing information of the scene from a device worn by a user, the device comprising one or more sensors, the information of the scene including depth information indicating a distance between the device worn by the user and a physical object in the scene; Detecting whether the physical object in the scene includes a hand; When the hand is detected, calculating a model of the hand based at least in part on the information of the scene; Using the model of the hand to mask depth information indicating a distance between the device worn by the user and a physical object in the scene; Calculating a hand mesh based on the depth information masked with respect to the model of the hand, the calculating including updating the hand mesh in real time as a relative location between the device and a change of the hand; Supplying the hand mesh to the application so that the application renders a portion of the virtual object not occluded by the hand mesh. A method comprising the above. (Item 2) The method according to Item 1, wherein the model of the hand includes a plurality of feature points of the hand indicating points on segments of the hand. (Item 3) The method according to Item 2, wherein at least a part of the plurality of feature points of the hand corresponds to joints of the hand and fingertips of the hand. (Item 4) The method further comprises: Determining a contour of the hand based on the plurality of feature points; Using the model of the hand to mask depth information indicating a distance between the device worn by the user and a physical object in the scene. Masking the depth information includes: Filtering and removing depth information outside the contour of the model of the hand; Generating a depth image of the hand based at least in part on the filtered depth information, the depth image including a plurality of pixels, each pixel indicating a distance to a point of the hand. The method according to item 2, including (Item 5) Filtering and removing depth information outside the contour of the hand model The method according to item 4, including removing depth information associated with physical objects in the scene (Item 6) Masking depth information indicating the distance between a device worn by the user and a physical object in the scene using the hand model includes associating a portion of the depth image with a hand segment Updating the hand mesh in real time includes selectively updating a portion of the hand mesh representing a suitable subset of the hand segment The method according to item 2 (Item 7) The method according to item 6, further including filling holes in the depth image before calculating the hand mesh (Item 8) Filling holes in the depth image The method according to item 7, including generating stereo depth information from a stereo camera of the device, where the stereo depth information corresponds to the area of the hole in the depth image (Item 9) Filling holes in the depth image The method according to item 7, including accessing surface information from a 3D model of the hand, where the surface information corresponds to the area of the hole in the depth image (Item 10) Calculating the hand mesh based on the depth information masked for the hand model Predicting a waiting time n from a query received at time t regarding data related to the hand in the scene from an application that renders the virtual object in the scene Predicting the hand posture at time t + the waiting time n Distorting the hand mesh using the posture predicted at time t + the waiting time n The method according to item 1, including (Item 11) The depth information indicating the distance between a device worn by the user and a physical object in the scene includes a sequence of depth images at a frame rate of at least 30 frames per second. The method according to item 1 (Item 12) An electronic system portable by a user, A device worn by the user, the device comprising a display configured to render virtual objects, and one or more sensors configured to capture information about the head pose of the user wearing the device and a scene including one or more physical objects, the information about the scene including depth information indicating a distance between the device and the one or more physical objects, a device, and A hand meshing component, the hand meshing component being configured to execute computer-executable instructions to detect a hand within the scene as the head pose changes and / or the hand moves, calculate a hand mesh of the detected hand, and update the hand mesh in real time, a hand meshing component, and An application, the application being configured to execute computer-executable instructions to render the virtual object within the scene, the application receiving, from the hand meshing component, the hand mesh and a portion of the virtual object occluded by the hand, an application An electronic system comprising. (Item 13) The hand meshing component Identifying feature points on the hand; Calculating segments between the feature points; Selecting information based on a proximity to one or more of the calculated segments from the depth information; and Calculating a mesh representing at least a portion of the hand mesh based on the selected depth information The portable electronic system according to item 12, configured to calculate a hand mesh by performing the above. (Item 14) The depth information includes a plurality of pixels, each of the plurality of pixels representing a distance to an object within the scene, Calculating the mesh includes grouping adjacent pixels representing a difference at a distance less than a threshold, The portable electronic system according to item 13. (Item 15) A method of operating an AR system to render a virtual object within a scene including a physical object, the AR system comprising at least one sensor and at least one processor, the method comprising Capturing scene information using the at least one sensor, wherein the scene information includes depth information indicating the distance to physical objects within the scene; Using the at least one processor; Processing the captured information to detect a hand within the scene and calculate points on the hand; Selecting a subset of the depth information based on the proximity to the calculated points on the hand; Calculating a representation of the hand based on the selected depth information, wherein the representation of the hand indicates the surface of the hand; A method comprising the above. (Item 16) Storing the calculated representation of the hand; Continuously processing the captured information and updating the stored representation of the hand; The method according to item 15, further comprising the above. (Item 17) Calculating the representation of the hand comprises: Calculating one or more parameters of the hand movement based on the captured information; Projecting the position of the hand at a future time determined based on the latency associated with rendering a virtual object using the calculated representation of the hand based on the one or more parameters of the movement; Morphing the calculated representation of the hand to represent the hand at the projected position; The method according to item 15, comprising the above. (Item 18) The method according to item 15, further comprising rendering a selected portion of the virtual object based on the representation of the hand, wherein the selected portion represents the portion of the virtual object not occluded by the hand. (Item 19) The depth information includes a depth map comprising a plurality of pixels each representing a distance; Calculating the representation of the hand based on the selected depth information includes identifying a group of pixels representing a surface segment; The method according to item 15. (Item 20) Calculating the representation of the hand includes defining a mesh representing the hand based on the identified group of pixels, the method according to item 19. (Item 21) Defining the mesh includes identifying triangular regions corresponding to the identified surface segments, the method according to item 20.
Brief Description of the Drawings
[0028] The accompanying drawings are not intended to be drawn to scale. In the drawings, each same or substantially same component illustrated in various figures is represented by like numerals. For purposes of clarity, not all components are labeled in all of the drawings.
[0029]
Figure 1
[0030]
Figure 2
[0031]
Figure 3
[0032]
Figure 4
[0033]
Figure 5A
[0034]
Figure 5B
[0035]
Figure 6
[0036]
Figure 7
[0037]
Figure 8
[0038]
Figure 9
[0039]
Figure 10
[0040]
Figure 11
[0041] What is described in this specification are methods and apparatuses for fast hand meshing for dynamic occlusion in an XR (Extended Reality) system. The XR system may create and use 3D reconstructions. To provide a realistic XR experience to the user, the XR system must understand the user's physical surroundings in order to correctly correlate the location of virtual objects with real objects. World reconstruction may be built from images and depth information about those physical surroundings collected using sensors that are part of the XR system. The world reconstruction may then be used by any of a plurality of components of such a system. For example, the world reconstruction may be used by components that perform visual occlusion processing, calculate physics-based interactions, or perform environment inference.
[0042] Occlusion processing identifies portions of a virtual object that should not be rendered and / or displayed to the user because objects in the physical world exist that block the user's view of where the virtual object should be perceived by the user. Physics-based interactions are calculated to determine where or how a virtual object appears to the user. For example, a virtual object may be rendered to appear to be resting on a physical object, moving through space, or colliding with the surface of a physical object. World reconstruction provides a model from which information about objects in the physical world can be obtained for such calculations.
[0043] When providing such a system, there are significant challenges. Substantial processing may be required to perform world reconstruction and calculate occlusion information. Additionally, the XR system must correctly understand how to position virtual objects in relation to the user's head, body, etc. As the user's position in relation to the physical environment changes, the relevant parts of the physical world may also change, which may require further processing. Further, 3D reconstruction data is often required to be updated as objects move within the physical world (e.g., a cup moves on a table). Updates to the data representing the environment the user is experiencing must be performed quickly without using too many computing resources of the computer generating the XR environment, as it is not possible to perform other functions while performing world reconstruction. Further, the processing of the reconstruction data by components that "consume" that data can exacerbate the demand for computer resources.
[0044] Dynamic occlusion processing identifies portions of virtual objects that should not be rendered and / or displayed to the user because there are physical objects that block the user's view of where the virtual object should be perceived by the user and the relative positions between the physical object and the virtual object change over time. Occlusion processing for considering the user's hand can be particularly important for providing a desirable XR experience. However, the inventors recognized and appreciated the true value that occlusion processing specifically improved for the hand can provide a more realistic XR experience for the user. The XR system can generate a mesh for an object used in occlusion processing, for example, based on a graphic image obtained at a frame rate of about 5 frames per second (fps). However, that rate may not meet the speed of location changes between the hand and virtual objects behind the hand, e.g., over 15 fps, over 30 fps, or over 45 fps, due to hand movement and / or head movement.
[0045] Users of XR devices may interact with the device by performing gestures with their hands. Since the hands are used directly in user interactions, latency is important. The user's hands can move at high speed during interaction with the device, for example, faster than the user moves to scan the physical environment for world reconstruction. Further, the user's hands are closer to the XR device worn by the user. Thus, the relative position between the user's hands and virtual objects behind the hands is also sensitive to head movement. If the hand representation used for occlusion processing is not updated fast enough to keep up with these sources of relative motion, the occlusion processing will not be based on the location of the hands and the occlusion processing will be inaccurate. If the virtual objects behind the hands are not rendered correctly to appear occluded by the hands during hand movement and / or head movement, the XR scene will appear unrealistic to the user. The virtual objects may appear on top of the hands as if the hands were transparent. Otherwise, the virtual objects may not appear to be in their intended location. The hands may appear to have the color pattern of the virtual objects, or other artifacts may appear. As a result, hand movement will break the user's immersion in the XR experience.
[0046] The inventors recognized and appreciated the true value that when an object occludes a virtual object, particularly high computational demands can be required when the object is the user's hand. However, the computational burden can be reduced by techniques that generate hand occlusion data at a high rate using low computational resources. The hand occlusion data may be generated by calculating a mesh of the hand from live depth sensor data obtained at a higher frequency than a graphic image. In some embodiments, the live depth sensor data may be obtained at a frame rate of at least 30 fps. To enable high-speed processing of the data, a small amount of data is processed to create a hand model for use in occlusion processing by masking the live depth data using a model in which the hand is represented by a plurality of segments simply identified from feature points. Further, to increase the accuracy of the occlusion processing, the hand occlusion data may be generated by predicting a change in the hand pose between the time of capture of the depth data and the time at which the hand mesh will be used for occlusion processing. The hand mesh may be distorted to represent the hand in the predicted pose.
[0047] Techniques as described herein may be used with or separately from many types of devices, including wearable or portable devices with limited computational resources, for many types of scenarios that provide cross-reality scenes. In some embodiments, the techniques may be implemented by a service that forms part of an XR system.
[0048] Figures 1-2 illustrate such a scenario. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3-6 illustrate an exemplary AR system that may operate in accordance with the techniques described herein and includes one or more processors, memory, sensors, and a user interface.
[0049] Referring to FIG. 1, an outdoor AR scene 4 is depicted, and to a user of AR technology, a physical-world park-like setting 6 is visible, featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these items, a user of AR technology also "sees" a robot image 10 standing on the concrete platform 8 of the physical world and a flying comic-like avatar character 2 that appears to be a anthropomorphic representation of a bumblebee, although these elements (e.g., avatar character 2 and robot image 10) do not exist within the physical world. Due to the extreme complexity of human visual perception and nervous system, it is difficult to produce AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements among other virtual or physical-world image elements.
[0050] Such an AR scene can be achieved using a system that includes a world reconstruction component that can construct and update a representation of the physical-world surface surrounding the user. This representation may be used for rendering occlusion, placing virtual objects in physics-based interaction states, and for virtual character path planning and navigation, or for other operations for which information about the physical world is used. FIG. 2 depicts another example of an indoor AR scene 200 that shows an exemplary world reconstruction use case that includes visual occlusion 202, physics-based interaction 204, and environment inference 206, according to some embodiments.
[0051] The exemplary scene 200 is a living room having a wall, a bookshelf on one side of the wall, a floor lamp at the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, the user of the AR technology may also perceive virtual objects such as an image on the wall behind the sofa, a bird flying in through the door, a deer peeking out from the bookshelf, and a figurine in the form of a windmill placed on the coffee table. With respect to the image on the wall, the AR technology requests information about not only the surface of the wall but also objects and surfaces in the room such as the lamp shape, which occludes the image and renders the virtual object correctly. With respect to the flying bird, the AR technology renders the bird with realistic physics and requests information about all objects and surfaces around the room to avoid the object and surface or their bounce-back when the bird collides. With respect to the deer, the AR technology requests information about surfaces such as the floor or the coffee table and calculates the place where the deer should be placed. With respect to the windmill, the system may identify that it is an object separate from the table and may infer that it is movable, while the corner of the shelf or the wall may be inferred to be stationary. Such specificities may be used in the inference regarding the parts of the scene used or updated in each of the various operations.
[0052] The scene may be presented to the user via a system that includes a plurality of components, including a user interface that may stimulate one or more user senses, including vision, sound, and / or touch. In addition, the system may include one or more sensors that may measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. Further, the system may include one or more computing devices with associated computer hardware such as memory. These components may be integrated within a single device or may be distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated within a wearable device.
[0053] Figure 3 depicts an AR system 302 configured to provide an experience of AR content that interacts with the physical world 306 according to some embodiments. The AR system 302 may include a display 308. In the illustrated embodiment, the display 308 may be worn by a user as part of a headset such that the user can wear the display across their eyes like a pair of goggles or glasses. At least a portion of the display may be transparent so that the user can observe the see-through reality 310. The see-through reality 310 may correspond to a portion of the physical world 306 within the current viewing perspective of the AR system 302 that the user can observe when wearing a headset that incorporates both the display and sensors of the AR system and obtains information about the physical world, corresponding to the user's viewing perspective.
[0054] The AR content may also be presented on the display 308 overlaid on the see-through reality 310. To provide an accurate interaction between the AR content and the see-through reality 310 on the display 308, the AR system 302 may include sensors 322 configured to capture information about the physical world 306.
[0055] The sensors 322 may include one or more depth sensors that output depth maps 312. Each depth map 312 may have a plurality of pixels, each representing the distance to a surface within the physical world 306 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensors and may be used to create depth maps. Such depth maps may be updated at the same rate as the depth sensors can form new images, which may be hundreds or thousands of times per second. However, the data may be noisy, incomplete, and may have holes, shown as black pixels on the depth maps.
[0056] The system may include other sensors such as an image sensor. The image sensor may obtain information that can be processed to represent the physical world in other ways. For example, an image may be processed in the world reconstruction component 316 to create a mesh that represents connected parts of objects within the physical world. For example, metadata about such objects, including color and surface texture, may also be obtained using sensors and stored as part of the world reconstruction.
[0057] The system may also obtain information about the user's head pose with respect to the physical world. In some embodiments, the sensor 310 may include an inertial measurement unit (IMU) that can be used to calculate and / or determine the head pose 314. The head pose 314 for the depth map may indicate, for example, the current viewpoint of the sensor that captures the depth map, with six degrees of freedom (6DoF). However, the head pose 314 may be used for other purposes, such as associating image information with a particular part of the physical world or associating the position of a display worn on the user's head with the physical world. In some embodiments, the head pose information may be derived in ways other than from the IMU, such as by analyzing objects within the image.
[0058] The world reconstruction component 316 may receive the depth map 312 and the head pose 314 and any other data from the sensors and integrate that data into the reconstruction 318, which may appear as at least a single combined reconstruction. The reconstruction 318 may be more complete and less noisy than the sensor data. The world reconstruction component 316 may update the reconstruction 318 using the spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0059] The reconstruction 318 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same part of the physical world or different parts of the physical world. In the illustrated embodiment, on the left side of the reconstruction 318, a part of the physical world is presented as a global surface, and on the right side of the reconstruction 318, a part of the physical world is presented as a mesh.
[0060] The reconstruction 318 may be used for AR functions, such as producing a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation may change as the user moves or as objects within the physical world change. The side of the reconstruction 318 may be used, for example, by a component 320 that produces a changing global surface representation within world coordinates, which may also be used by other components.
[0061] AR content may be generated, for example, by an AR application 304 or the like based on this information. The AR application 304 may be a game program that implements one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. These functions may be implemented by querying the reconstruction 318 produced by the world reconstruction component 316 for data in different formats. In some embodiments, the component 320 may be configured to output an update when the representation within the area of interest of the physical world changes. The area of interest may be set to approximate a part of the physical world in the vicinity of the user of the system, such as a part within the user's field of view, or projected (predicted / decided) to occur within the user's field of view.
[0062] The AR application 304 may generate and update AR content using this information. The virtual portion of the AR content may be presented on the display 308 in combination with see-through reality 310 to create a realistic user experience.
[0063] In some embodiments, the AR experience may be provided to the user through a wearable display system. FIG. 4 illustrates an example of a wearable display system 80 (hereinafter referred to as "system 80"). The system 80 includes a head-mounted display device 62 (hereinafter referred to as "display device 62") and various mechanical and electronic modules and systems for supporting the functions of the display device 62. The display device 62 may be coupled to a frame 64, which is wearable by a display system user or viewer 60 (hereinafter referred to as "user 60") and is configured to position the display device 62 in front of the eyes of the user 60. According to various embodiments, the display device 62 may be a sequential display. The display device 62 may be monocular or binocular. In some embodiments, the display device 62 may be an example of the display 308 in FIG. 3.
[0064] In some embodiments, a speaker 66 is coupled to the frame 64 and positioned proximate to the ear canal of the user 60. In some embodiments, another speaker (not shown) is positioned adjacent to the other ear canal of the user 60 to provide stereo / adjustable sound control. The display device 62 is operably coupled to a local data processing module 70 by means of a wired conductor or wireless connectivity 68, etc., which may be mounted in various configurations such as fixed to the frame 64, fixed to a helmet or hat worn by the user 60, built into headphones, or alternatively removably attached to the user 60 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0065] The local data processing module 70 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which may be used to assist in data processing, caching, and storage. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope (e.g., operatively coupled to frame 64 or otherwise attachable to user 60), and / or b) possibly data obtained and / or processed using the remote processing module 72 and / or the remote data repository 74 for passage to the display device 62 after processing or reading. The local data processing module 70 may be operatively coupled to the remote processing module 72 and the remote data repository 74, respectively, via communication links 76, 78, such as a wired or wireless communication link, such that these remote modules 72, 74 are operatively coupled to each other and available as resources to the local processing and data module 70. In some embodiments, the world reconstruction component 316 in FIG. 3 may be implemented at least partially within the local data processing module 70. For example, the local data processing module 70 may be configured to execute computer-executable instructions to generate a physical world representation based at least in part on at least a portion of the data.
[0066] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 70 may include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the calculation budget of the module 70 but will enable a smaller device. In some embodiments, the world reconstruction component 316 may use a calculation budget less than a single ARM core so that the remaining calculation budget of the single ARM core can be accessed for other uses, such as extracting a mesh, etc., and generate a physical world representation in real time over an unspecified space.
[0067] In some embodiments, the remote data repository 74 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed within the local data processing module 70, enabling fully autonomous use from the remote module. World reconstruction may be stored, for example, in whole or in part within this repository 74.
[0068] In some embodiments, the local data processing module 70 is operably coupled to the battery 82. In some embodiments, the battery 82 is a removable power source, such as a commercially available battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during non-operation time of the system 80 and a removable battery, so that the user 60 can operate the system 80 for a longer time period without having to connect to a power source, charge the lithium-ion battery, or shut down the system 80 and replace the battery.
[0069] Figure 5A illustrates a user 30 wearing an AR display system that renders AR content as the user 30 moves through a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at position 34, and the AR display system records surrounding information of a passable world (e.g., a digital representation of a physical object in the physical world that can be memorized and updated as the physical object changes) with respect to the mapped features or the positional relationship to the directional audio input at position 34. The position 34 is aggregated into data input 36 and processed by a passable world module 38, which can be implemented, at least, by processing on the remote processing module 72 of FIG. 4. In some embodiments, the passable world module 38 may include a world reconstruction component 316.
[0070] The passable world module 38 determines where and how the AR content 40 can be placed in the physical world such that the AR content 40 is determined from the data input 36. The AR content is "installed" in the physical world by presenting both the physical world and the representation of the AR content via a user interface, and the AR content is rendered as if it were interacting with the objects in the physical world, and the objects in the physical world are presented as if the AR content were obscuring the user's view of those objects when appropriate. In some embodiments, the AR content may be installed by appropriately selecting a portion of a fixed element 42 (e.g., a table) from the reconstruction (e.g., reconstruction 318) and determining the shape and position of the AR content 40. As an example, the fixed element may be a table, and the virtual content may be positioned to appear as if it were on that table. In some embodiments, the AR content may be installed within the structures in the field of view 44, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be installed with respect to a mapped mesh model 46 of the physical world.
[0071] As described, the fixed element 42 serves as a substitute for any fixed element within the physical world, which may be stored within the passable world module 38 such that the user 30 can perceive the content on the fixed element 42 without the system having to map it to the fixed element 42 each time it is visible to the user 30. The fixed element 42 may thus be a mapped mesh model from a previous modeling session or determined from a separate user, provided it is stored on the passable world module 38 for future reference by multiple users. Thus, the passable world module 38 allows the user 30's device to recognize the environment 32 from a previously mapped environment and displayed AR content without first mapping the environment 32, saving calculation processes and cycles and avoiding latency for any rendered AR content.
[0072] The mapped mesh model 46 of the physical world may be created by the AR display system, interacts with the AR content 40, and the appropriate surfaces and metrics for displaying it can be mapped and stored within the passable world module 38 for future retrieval by the user 30 or other users without the need for remapping or modeling. In some embodiments, the data input 36 is an input such as a geographical location, user identification, and current activity, which indicates to the passable world module 38 the fixed element 42 of one or more available fixed elements, the AR content 40 last placed on the fixed element 42, and whether to display that same content (such AR content being "persistent" content regardless of whether the user is viewing a particular passable world model).
[0073] In embodiments where an object is considered to be fixed, the passable world module 38 may be updated at any time to account for the variability of the physical world. The model of the fixed object may be updated very infrequently. Other objects within the physical world may be considered to be moving or otherwise not fixed. To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects at a much higher frequency than the frequency used to update the fixed objects. To enable accurate tracking of all objects within the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0074] FIG. 5B is a schematic diagram of the viewing optical system assembly 48 and associated components. In some embodiments, two eye-tracking cameras 50 that are directed towards the user's eyes 49 detect metrics of the user's eyes 49, such as eye shape, eyelid occlusion, pupil direction, and glare on the user's eyes 49. In some embodiments, one of the sensors is a depth sensor 51, such as a time-of-flight sensor, that emits signals into the world and detects the reflection of those signals from neighboring objects to determine the distance to a given object. The depth sensor may, for example, quickly determine whether an object is entering the user's field of view as a result of either the movement of those objects or a change in the user's posture. However, information about the position of an object within the user's field of view may alternatively or additionally be collected using other sensors. The depth information may, for example, be obtained from a stereoscopic image sensor or a plenoptic sensor.
[0075] In some embodiments, the world camera 52 records a view that exceeds the peripheral vision, maps the environment 32, and detects inputs that can affect the AR content. In some embodiments, the world camera 52 and / or the camera 53 may be grayscale and / or color image sensors, which may output grayscale and / or color image frames at fixed time intervals. The camera 53 may further capture a physical world image within the user's field of view at a specific time. The pixels of the frame-based image sensor may be sampled iteratively even if their values are invariant. The world camera 52, the camera 53, and the depth sensor 51 each have individual fields of view of 54, 55, and 56, respectively, and collect and record data from a physical world scene such as the physical world environment 32 depicted in FIG. 5A.
[0076] The inertial measurement unit 57 may determine the movement and orientation of the visual optics assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 51 is operably coupled to the eye tracking camera 50 as a confirmation of the focus adjustment measured with respect to the actual distance that the user's eye 49 is looking at.
[0077] It should be understood that the visual optics assembly 48 may include some of the components illustrated in FIG. 5B, and may include components instead of or in addition to the illustrated components. In some embodiments, for example, the visual optics assembly 48 may include two world cameras 52 instead of four. Alternatively, or in addition, the cameras 52 and 53 do not need to capture a visible light image of their full fields of view. The visual optics assembly 48 may include other types of components. In some embodiments, the visual optics assembly 48 may include one or more dynamic vision sensors (DVSs), and the pixels thereof may respond asynchronously to relative changes in light intensity that exceed a threshold.
[0078] In some embodiments, the visual optical system assembly 48 may not include the time-of-flight based depth sensor 51. In some embodiments, for example, the visual optical system assembly 48 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of incident light, from which depth information can be determined. For example, a plenoptic camera may include an image sensor overlaid with a transmissive diffraction mask (TDM). Alternatively, or in addition, a plenoptic camera may include an image sensor containing angle sensing pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a source of depth information instead of, or in addition to, the depth sensor 51.
[0079] Also, it should be understood that the component configuration in FIG. 5B is shown as an example. The visual optical system assembly 48 may include components with any suitable configuration, which may be set to provide the user with a practical maximum field of view for a particular set of components. For example, if the visual optical system assembly 48 has one world camera 52, the world camera may be installed within the central region of the visual optical system assembly instead of on the side.
[0080] Information from sensors within the visual optics system assembly 48 may be coupled to one or more of the processors within the system. The processor may generate data that may be rendered to cause the user to perceive virtual content that interacts with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in a single scene by modulating the opacity of a display device that the user views through the physical world. The opacity may be controlled to create the appearance of the virtual object and to block out physical objects within the physical world that are occluded by the virtual object from view by the user. In some embodiments, the image data may be modified to be perceived by the user as if the virtual content is realistically interacting with the physical world when viewed through the user interface (e.g., clipping the content to account for occlusion), and may include only the virtual content. Regardless of how the content is presented to the user, the model of the physical world is required to accurately calculate the properties of virtual objects that may be affected by physical objects, including the shape, position, motion, and visibility of the virtual objects. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 318.
[0081] The model may be created from data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created from data collected by multiple users, which may be aggregated within a computing device remote from all users (and may be "in the cloud").
[0082] The model may be created, at least in part, by a world reconstruction system, such as the world reconstruction component 316 of FIG. 3, which is depicted in more detail in FIG. 6. The world reconstruction component 316 may include a perception module 160 that can generate, update, and store a representation for a portion of the physical world. In some embodiments, the perception module 160 may represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world and may include surface information indicating whether a surface exists within the volume represented by the voxel. The voxels may be assigned a value indicating whether the corresponding volume has been determined to be empty, has been determined to contain a surface of a physical object, or has not yet been measured using the sensor and thus its value is unknown. It should be understood that the values may be stored in computer memory in any suitable manner, including not storing information regarding voxels that have been determined to be empty or unknown, indicating that voxels determined to be empty or unknown need not be explicitly stored.
[0083] In addition to generating information for the persistent world representation, the perception module 160 may identify and output indications of changes in the area surrounding the user of the AR system. Such indications of change may trigger other functions, such as triggering an update to the volumetric data stored as part of the persistent world, or triggering component 304 to generate and update AR content.
[0084] In some embodiments, the perception module 160 may identify changes based on a signed distance function (SDF) model. The perception module 160 may be configured to receive sensor data such as, for example, a depth map 160a and a head pose 160b, and then fuse the sensor data into an SDF model 160c. The depth map 160a may directly provide SDF information, and an image may be processed to become SDF information. The SDF information represents the distance from the sensor used to capture that information. Since those sensors may be part of the wearable unit, the SDF information may represent the physical world from the line of sight of the wearable unit and thus the user's line of sight. The head pose 160b may enable the SDF information to be associated with voxels within the physical world.
[0085] In some embodiments, the perception module 160 may generate, update, and store a representation for a portion of the physical world within the perception range. The perception range may be determined at least in part based on the reconstruction range of the sensor, which may be determined at least in part based on the limits of the sensor's observation range. As a specific example, an active depth sensor operating using active IR pulses may reliably operate over a certain range of distances and create the sensor's observation range, which may be several centimeters or tens of centimeters to several meters.
[0086] The world reconstruction component 316 may include additional modules that may interact with the perception module 160. In some embodiments, the persistent world module 162 may receive a representation of the physical world based on data obtained by the perception module 160. The persistent world module 162 may also include various formats of the representation of the physical world. For example, volumetric metadata 162b such as voxels may be stored together with a mesh 162c and a plane 162d. In some embodiments, other information such as a depth map may also be stored.
[0087] In some embodiments, the perception module 160 may include modules that generate a representation for the physical world in various formats, such as, for example, mesh 160d, plane, and semantics 160e. These modules may generate the representation based on data within the perception range of one or more sensors at the time the representation is generated, data captured at previous times, and information within the persistent world 162. In some embodiments, these components may act on depth information captured using a depth sensor. However, the AR system may include a visual sensor and generate such a representation by analyzing monocular or binocular visual information.
[0088] In some embodiments, these modules may act on regions of the physical world. Those modules may be triggered to update a sub-region of the physical world when the perception module 160 detects a change in the physical world within that sub-region. Such changes may be detected, for example, by detecting a new surface within the SDF model 160c or by other criteria such as changing the values of a sufficient number of voxels representing the sub-region.
[0089] The world reconstruction component 316 may include a component 164 that can receive a representation of the physical world from the perception module 160. Information about the physical world may be pulled by these components, for example, according to usage requests from an application. In some embodiments, the information may be pushed to the usage components via an indication of a change in a pre-identified region or a change in the physical world representation within the perception range. The component 164 may include, for example, a game program and other components that perform processing for visual occlusion, physics-based interactions, and environmental inference.
[0090] In response to a query from component 164, the perception module 160 may transmit a representation of the physical world in one or more formats. For example, when component 164 indicates that the use is for visual occlusion or physics-based interaction, the perception module 160 may transmit a representation of the surface. When component 164 indicates that the use is for environmental inference, the perception module 160 may transmit a mesh, plane, and semantics of the physical world.
[0091] In some embodiments, the perception module 160 may include a component that formats information and provides it to component 164. An example of such a component may be a raycasting component 160f. The using component (e.g., component 164) may query for information about the physical world from a particular perspective, for example. The raycasting component 160f may select from one or more representations of the physical world data within the field of view from that perspective.
[0092] Information about the physical world may also be used for occlusion processing. That information may be used by a visual occlusion component 164a, which may be part of the world reconstruction component 316. The visual occlusion component 164a may supply, for example, an application with information indicating portions of a visual object that are occluded by a physical object. Alternatively, or in addition, the visual occlusion component 164a may provide an application with information about a physical object, which may use that information for occlusion processing. As described above, accurate information about the hand position is important for occlusion processing. In one embodiment, as described herein, the visual occlusion component 164a may maintain a hand model in response to a request by an application and provide that model to the application when requested. FIG. 7 illustrates an example of such a process, which may be implemented across one or more of the components illustrated in FIG. 6 or, in some embodiments, by different or additional components.
[0093] FIG. 7 depicts an AR system 700 configured to generate a hand mesh in real time for dynamic occlusion processing according to some embodiments. The AR system 700 may be implemented on an AR device. The AR system 700 may include a data collection portion 702 configured to use sensors on the AR device to capture the pose of a user (e.g., head pose, hand pose, and the like) wearing the AR device (e.g., a display device 62) and information about the scene. The information about the scene may include depth information indicating the distance between the AR device and physical objects in the scene.
[0094] The data collection part 702 includes a hand-tracking component. The hand-tracking component may process sensor data such as depth and image information and detect one or both hands in the scene. Other sensor data may also be processed to detect one or both hands in the scene. Once detected, one or both hands can be represented in a sparse way, for example, by a set of feature points. The feature points can represent, for example, joints, fingertips, or other boundaries of hand segments. The information collected or generated by the data collection part 702 may be passed to the hand meshing part 704 to be used to generate a richer model of one or both hands, for example, a mesh, based on the sparse representation.
[0095] The hand meshing part 704 is configured to calculate the hand mesh of the detected one or both hands and update the hand mesh in real time as the posture changes and / or the hand moves.
[0096] The AR system 700 may include an application 706 configured to receive the hand mesh from the hand meshing part 704 and render one or more virtual objects in the scene. In some embodiments, the application 706 may receive occlusion data from the hand meshing part 704. In some embodiments, the occlusion data may indicate the portions of the virtual objects occluded by one or both hands. In some embodiments, for example, in the illustrated embodiment, the occlusion data may be a model of one or both hands, from which the application 706 may calculate the occlusion data. As a specific example, the occlusion data may be the hand mesh of one or both hands received from the hand meshing part 704.
[0097] Figure 8 is a flowchart illustrating a method 800 for generating a hand mesh in real time for dynamic occlusion, according to some embodiments. In some embodiments, method 800 may be implemented by one or more processors within AR system 700. Method 800 may begin when the hand meshing component 704 of AR system 700 receives a query regarding data related to one or both hands within the scene from application 706 of AR system 700 (act 802). Method 800 may include a step (act 804) of detecting one or both hands within the scene based on information about the scene captured by data collection portion 702 of AR system 700.
[0098] When one or both hands are detected, method 800 may include a step (act 806) of calculating one or more models of the one or both hands based on the information about the scene. The one or more models of the one or both hands may be sparse and may indicate the positions of feature points on the hand rather than the surface. Those feature points may represent joints or the end portions of hand segments. The feature points may be recognized from sensor data about the one or both hands, including, for example, stereo images of the one or both hands. Depth information, and in some instances, monocular images of the one or both hands, may be used instead or in addition to identify the feature points.
[0099] U.S. Provisional Patent Application No. 62 / 850,542, entitled “Hand Pose Estimation,” describes exemplary methods and apparatuses for obtaining information about the position and pose of a hand and modeling the hand based on the obtained information. A copy of the filing version of U.S. Application No. 62 / 850,542 is attached hereto and incorporated herein by reference in its entirety for all purposes. Techniques as described in that application may be used to construct a sparse model of the hand.
[0100] In some embodiments, one or more models of one or both hands may be calculated based on information about the scene captured by the sensors of the AR device. Examples of scene information include the exemplary image of FIG. 9A captured by one sensor corresponding to one eye, the two exemplary images of FIG. 9B captured by two sensors corresponding to the left and right eyes, and the exemplary depth image of FIG. 9C that can be obtained at least partially from the image of FIG. 9A or the image of FIG. 9B.
[0101] In some embodiments, one or more models of one or both hands may include a plurality of feature points of the hand, which may indicate points on the hand segment. Some of the feature points may correspond to the joints of the hand and the fingertips of the hand. FIGS. 9E and 9F depict schematic diagrams illustrating an exemplary 8 - feature - point model of the hand and an exemplary 22 - feature - point model of the hand, respectively.
[0102] In some embodiments, the feature points of the hand model may be used to determine the contour of the hand. FIG. 9D depicts an exemplary contour of the hand, which may be determined based on the feature points of the hand model. For example, adjacent feature points may be connected by lines, schematically illustrated in FIGS. 9E and 9F, and the contour of the hand may be shown as the distance from the lines. The distance from the lines may be determined from the image of the hand, information about the human anatomical structure, and / or other information. It should be understood that once the feature points of the hand are identified, the hand model may be updated at a later time using information previously obtained about the hand. For example, the length of the line connecting the feature points may not change.
[0103] The sparse hand model may be used to select a limited amount of data from which a richer model of the hand, including surface information, may be constructed, for example. In some embodiments, the selection may be made by using the hand contour to mask additional data such as depth data. Thus, method 800 may include masking depth information indicative of the distance between the AR device and physical objects in the scene, using one or more models of one or both hands (act 808).
[0104] FIG. 10 depicts a flowchart illustrating details of the step of masking depth information (act 808) using one or more models of one or both hands, according to some embodiments. Act 808 may include filtering out depth information outside the contour of one or more models of one or both hands (act 1002). This act results in removing depth information associated with physical objects other than the hand in the scene. Act 808 may include generating a depth image of one or both hands based on the filtered depth information (act 1004).
[0105] The depth image may include pixels that may indicate the distance to points of one or both hands, respectively. In some embodiments, the depth information may be captured such that the depth information is not captured for all surfaces of one or both hands. For example, the depth information may be captured using an active IR sensor. If the user is wearing a ring with a dark stone, for example, the IR may not reflect from the dark ring such that there will be holes in the collected insufficient information. Act 808 may include filling holes in the depth image (act 1006). In some embodiments, the holes may be filled by identifying the holes in the depth image and generating stereo depth information corresponding to the identified regions from the stereo camera of the AR device. In some embodiments, the holes may be filled by identifying the holes in the depth image and accessing surface information corresponding to the identified regions from one or more 3D models of one or both hands.
[0106] Optionally, the mesh of one or both hands may be calculated as a plurality of sub-meshes, and each sub-mesh represents a segment of one or both hands. The segment may correspond to a segment bounded by feature points. Since many of those segments are bounded by joints, the segments correspond to parts of the hand that can move at least independently of other segments of the hand so that the hand mesh can be updated quickly by updating the sub-mesh associated with the segment that has moved since the last hand mesh calculation. In such an embodiment, act 808 may include identifying feature points in the depth image (act 1008) that correspond to feature points of one or more models of one or both hands. Segments of the hand separated by feature points may be calculated. Act 808 may include associating a portion of the depth image with a segment of the hand separated by the identified feature points in the depth image (act 1010).
[0107] Method 800 may include calculating a mesh of one or both hands (act 810) based on depth information masked for one or more models of one or both hands. The mesh of one or both hands may be a representation of one or both hands showing the surface of one or both hands. FIG. 9G depicts a calculated dense hand mesh according to some embodiments. However, the present application is not limited to the step of calculating a dense hand mesh. In some embodiments, a sparse hand mesh may be sufficient for dynamic occlusion.
[0108] In some embodiments, the one or both hand meshes may be calculated from depth information. For example, the mesh may often be a set of regions, represented as triangles, that represent a part of the surface. Such regions may be identified by grouping adjacent pixels in a depth image where the difference in distance with respect to them is less than a threshold, indicating that the pixels are likely to be on the same surface. One or more triangles bounding such a region of pixels may be identified and added to the mesh. However, other techniques for forming the mesh from depth information may be used.
[0109] The step of calculating the one or both hand meshes may include the step of updating the one or both hand meshes in real time as the relative location between the AR device and the one or both hands changes. There may be a latency between the time when the scene information used to calculate the one or both hand meshes is captured and the time when the one or more calculated hand meshes are used, for example, by an application, for example, to render content. In some embodiments, the movement of the one or both hand segments may be tracked such that the future positions of those segments of the one or both hands can be projected / predicted. The one or both hand meshes may be distorted to conform to the projected location of the segments at the time when virtual objects processed using the one or both hand meshes will be rendered.
[0110] FIG. 11 depicts a flowchart illustrating details of the step (act 810) of calculating in real time one or both hand meshes based on depth information masked with respect to hand segmentation, according to some embodiments. Act 810 may include the step (act 1102) of determining a time t when the hand meshing portion 704 receives a query from the application 706 regarding data related to one or both hands.
[0111] Action 810 may include predicting the latency n from a query received from application 706 at time t (action 1104). Action 810 may include predicting the posture (e.g., hand posture) at time t + latency n of the query (action 1106). Action 810 may include distorting the mesh of one or both hands using the posture predicted at time t + latency n of the query (action 1108). In some embodiments, the step of predicting the hand posture may include predicting the movement of feature points of one or both hands in the depth image at time t + latency n of the query. Such prediction may be based on the step of tracking the position of the feature points over time. Such tracking enables the determination of motion parameters such as velocity or acceleration. The projection of the position may be based on extrapolation from the previous position to the future position assuming that the determined motion parameters remain the same. Alternatively, or in addition, the projection may be determined using a Kalman filter or similar projection techniques.
[0112] In some embodiments, the step of distorting the mesh of one or both hands using the predicted posture may include distorting the portion of the mesh of one or both hands corresponding to a subset of hand segments that are predicted to change at time t + latency n between feature points.
[0113] When distorting the previously calculated mesh to represent one or both hands at time t + latency n, multiple factors may be considered. The value of n may reflect, for example, the time required for processing to distort one or more meshes and for the application to use one or more meshes when rendering an object. The value may be estimated or measured from the operation, structure, or testing of the software. As another example, the value may be dynamically determined based on the step of adjusting a previously established latency based on measuring the latency during use or based on the processing load at the time when a request for one or more meshes is made.
[0114] When distorting one or more meshes, the amount of distortion may be based on the time when the data used to form the mesh of one or both hands was captured and the waiting time until one or more meshes will be used. In some embodiments, the mesh of one or both hands may be created other than in response to a request from an application. For example, once an application indicates that it is configured for occlusion processing, such as by making a call through an API, the AR system may periodically calculate one or more updated hand meshes. Alternatively, the hand tracking process may be continuously activated using a certain amount of the system's computing resources. In either case, the mesh of one or both hands may be updated relatively frequently, such as at least 30 times per second. Note that there may be a delay between when the data is captured to create the mesh and when a request for the mesh is received, and this delay may also be considered when distorting the hand mesh.
[0115] Method 800 may include supplying the application 706 with a mesh of one or both hands (act 812) such that the application renders portions of the virtual object that are not occluded by the mesh of one or both hands.
[0116] Although some aspects of some embodiments have been described so far, it should be understood that various modifications, corrections, and improvements will readily occur to those skilled in the art.
[0117] As an example, embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applied within an MR environment, or more generally, within other XR environments and VR environments.
[0118] As another example, the embodiments are described in relation to devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), discrete applications, and / or any suitable combination of devices, networks, and discrete applications.
[0119] Such modifications, corrections, and improvements are intended to be part of the present disclosure and are intended to be within the spirit and scope of the present disclosure. Further, while the advantages of the present disclosure are presented, it should be understood that not all embodiments of the present disclosure include all of the advantages described. Some embodiments may not implement any of the features described as advantageous herein and in some instances. Therefore, the foregoing description and drawings are merely examples.
[0120] The foregoing embodiments of the present disclosure can be implemented in any of a number of ways. For example, an embodiment may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided within a single computer or distributed among multiple computers. Such processors can include one or more processors within an integrated circuit component, including commercially available integrated circuit components known in the art, such as a CPU chip, a GPU chip, a microprocessor, a microcontroller, or a coprocessor. In some embodiments, the processor may be implemented within a custom circuit such as an ASIC or within a semi-custom circuit resulting from configuring a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores such that a subset of one or those cores can constitute the processor. However, the processor may be implemented using a circuit in any suitable format.
[0121] Furthermore, it should be understood that the computer can be embodied in any of several forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, the computer may be embodied in a device that is not generally considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or stationary electronic device.
[0122] In addition, the computer may have one or more input and output devices. These devices can be used, inter alia, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound generating device for audible presentation of output. Examples of input devices that can be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, the computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output devices are shown physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated within the same unit as the processor or other elements of the computing device. For example, a keyboard can be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated through a wireless connection.
[0123] Such computers may be interconnected by one or more networks in any suitable form, including forms such as a local area network or wide area network, such as a corporate network or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0124] In addition, the various methods and processes outlined in this specification may be encoded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may also be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine.
[0125] In this aspect, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy (registered trademark) disks, compact disk (CD), optical disk, digital video disk (DVD), magnetic tape, flash memory, circuit configurations within a field programmable gate array or other semiconductor device, or other tangible computer storage media) encoded with one or more programs that perform a method of implementing the various embodiments of the present disclosure discussed above when executed on one or more computers or other processors. As will be apparent from the foregoing examples, a computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such a computer-readable storage medium or media can be transportable such that the one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" includes only computer-readable media that can be regarded as a manufactured (i.e., a manufactured article) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.
[0126] As described above, the terms "program" or "software" are used herein in a general sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure. Additionally, according to one aspect of the present embodiment, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion among several different computers or processors to implement various aspects of the present disclosure.
[0127] Computer-executable instructions may take many forms, such as program modules, which are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0128] Also, data structures may be stored on a computer-readable medium in any suitable form. For simplicity of illustration, a data structure may be shown to have fields that are related through locations within the data structure. Such relationships may also be achieved by allocating storage for fields that accompany locations within the computer-readable medium that convey relationships between the fields. However, any suitable mechanism, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements, may be used to establish relationships between the information within the fields of a data structure.
[0129] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and thus, its use is not limited to the details and arrangements of the components described in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined with aspects described in other embodiments in any manner.
[0130] Also, the present disclosure may be embodied as a method, for which examples are provided. The acts performed as part of the method may be ordered in any suitable way. Thus, in the illustrative embodiments, although shown as consecutive acts, embodiments may be constructed in which acts are performed in a different order than illustrated, including performing some acts simultaneously.
[0131] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not by itself imply any priority, precedence, or order of one claim element over another, or the temporal order in which acts of a method are performed, but the ordinal terms are used only as labels to distinguish one claim element having a certain name from another element having the same name (for the purpose of using the ordinal terms).
[0132] Also, the phrases and terminology used herein are for the purpose of explanation and should not be regarded as limiting. The use of "comprising," "including," "having," "containing," "involving," and variations thereof herein means including the items listed thereafter and their equivalents and additional items.
Claims
1. A method of reconstructing a hand by operating a computing system to dynamically occlude a virtual object, the method comprising: Receiving, from an application that renders a virtual object within a scene, a query regarding data related to a hand within the scene; Capturing information about the scene from a device worn by a user, the device comprising one or more sensors, the information about the scene including depth information indicating a distance between the device worn by the user and a physical object within the scene; Detecting whether the physical object within the scene includes a hand; When the hand is detected, calculating a model of the hand based at least in part on the information about the scene; Using the model of the hand to mask the depth information indicating the distance between the device worn by the user and the physical object within the scene; Calculating a hand mesh based on the depth information masked with respect to the model of the hand, the calculating including updating the hand mesh in real time as a relative location between the device and a change in the hand; Supplying the hand mesh to the application such that the application renders a portion of the virtual object not occluded by the hand mesh. A method comprising the above.
2. The method according to claim 1, wherein the model of the hand includes a plurality of feature points of the hand indicating points on segments of the hand.
3. A method of rendering a virtual object within a scene including a physical object by operating an AR system, the AR system comprising at least one sensor and at least one processor, the method comprising: Capturing information about the scene using the at least one sensor, the information about the scene including depth information indicating a distance to a physical object within the scene; And Using the at least one processor to Process the captured information to detect a hand within the scene and calculate points on the hand; Selecting a subset of the depth information based on a proximity to the calculated points on the hand. Calculating the representation of the hand based on the selected depth information, wherein the representation of the hand indicates the surface of the hand, and A method comprising: **Claim 4** The method according to claim 2 or claim 3, wherein at least some of the plurality of feature points of the hand correspond to joints of the hand and fingertips of the hand. **Claim 5** The method further includes determining a contour of the hand based on the plurality of feature points, Masking the depth information is Filtering and removing the depth information outside the contour of the model of the hand, and Generating a depth image of the hand based at least in part on the filtered depth information, wherein the depth image includes a plurality of pixels, and each pixel indicates a distance to a point on the hand, and The method according to claim 2, comprising: **Claim 6** The method according to claim 5, wherein filtering and removing the depth information outside the contour of the model of the hand includes removing depth information associated with physical objects in the scene. **Claim 7** Masking the depth information indicating the distance between the device worn by the user and the physical object in the scene using the model of the hand includes associating a portion of the depth image with a segment of the hand, The method according to claim 5, wherein updating the mesh of the hand in real time includes selectively updating a portion of the mesh of the hand that represents a suitable subset of the segments of the hand. **Claim 8** The method according to claim 5 or claim 7, further including filling holes in the depth image before calculating the mesh of the hand. **Claim 9** The method according to claim 8, wherein filling holes in the depth image includes generating stereo depth information from a stereo camera of the device, and the stereo depth information corresponds to an area of the holes in the depth image. **Claim 10** The method according to claim 8, wherein filling holes in the depth image includes accessing surface information from a 3D model of the hand, and the surface information corresponds to an area of the holes in the depth image. **Claim 11** Calculating the mesh of the hand based on the depth information masked with respect to the model of the hand is Predicting a latency n from the query received at query time t regarding the data related to the hand within the scene from the application that renders the virtual object within the scene; Predicting a hand pose at a time of the query time t + the latency n; Distorting the hand mesh using the pose predicted at the time of the query time t + the latency n; The method according to claim 1, comprising:
12. The method according to claim 1 or claim 3, wherein the depth information includes a sequence of depth images at a frame rate of at least 30 frames per second.
13. A portable electronic system carried by a user, the portable electronic system comprising: A device worn by the user, the device comprising a display configured to render a virtual object, the device comprising one or more sensors configured to capture information about a head pose of the user wearing the device and a scene including one or more physical objects, the information about the scene including depth information indicating a distance between the device and the one or more physical objects, a device; A hand meshing component configured to detect a hand as the head pose changes and / or the hand within the scene moves, calculate a hand mesh of the detected hand, and update the hand mesh in real time by executing computer-executable instructions; An application configured to render the virtual object within the scene by executing computer-executable instructions, the application receiving the hand mesh and a portion of the virtual object occluded by the hand from the hand meshing component; A portable electronic system comprising:
14. The hand meshing component: Identifying a plurality of feature points on the hand; Calculating a plurality of segments between the plurality of feature points; Selecting information from the depth information based on a proximity to one or more of the calculated plurality of segments; Calculating a mesh representing at least a part of the hand mesh based on the selected depth information The portable electronic system according to claim 13, which is configured to calculate a hand mesh by performing the above.
15. The depth information includes a plurality of pixels, and each of the plurality of pixels represents the distance to an object in the scene. The portable electronic system according to claim 14, wherein calculating the mesh includes grouping adjacent pixels representing differences at distances less than a threshold value.
16. The method includes storing the calculated representation of the hand, updating the stored representation of the hand by continuously processing the captured information The method according to claim 3, further comprising.
17. Calculating the representation of the hand includes calculating one or more parameters of the hand movement based on the captured information, projecting the position of the hand at a future time determined based on the waiting time associated with rendering a virtual object using the calculated representation of the hand based on one or more parameters of the movement, projecting the position of the hand, representing the hand at the projected position by morphing the calculated representation of the hand The method according to claim 3, comprising.
18. The method further includes rendering a selected portion of the virtual object based on the representation of the hand, the selected portion representing a portion of the virtual object not occluded by the hand. The method according to claim 3.
19. The depth information includes a depth map including a plurality of pixels, and each of the plurality of pixels represents a distance. Calculating the representation of the hand based on the selected depth information includes identifying a group of pixels representing a surface segment. The method according to claim 3.
20. Calculating the representation of the hand includes defining a mesh representing the hand based on the identified group of pixels. The method according to claim 19.
21. Defining the mesh includes identifying a triangular region corresponding to the identified surface segment. The method according to claim 20.
Citation Information
Patent Citations
Image processing method and image processing device
JP2006343953A
A system for rendering a shared digital interface from each user's perspective.
JP2014514653A
Information processing apparatus, information processing system, information processing method, and program
JP2016115148A
Methods and systems for creating virtual and augmented reality.
JP2017529635A
Information processing apparatus, method for controlling information processing apparatus, and program
JP2018022292A