Occlusion and Collision Detection for Augmented Reality Applications

By generating and updating depth images, combining RGB images, generating RGBD images, and performing multi-level voxel collision detection, the problem of inaccurate placement of virtual objects in augmented reality technology is solved, and high-quality AR scene rendering is achieved.

CN114450717BActive Publication Date: 2025-06-20GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080068530.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-09-29
Publication Date
2025-06-20
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

In existing augmented reality technology, there are difficulties in accurate placement and real-time rendering of virtual objects in AR scenes, resulting in low rendering quality.

Method used

The computer system generates a depth image and divides it into multiple depth layers, adjusts the pixel position to update the depth image, generates an RGBD image, determines the collision between the virtual object and the multi-level voxel, and renders the virtual object based on this.

Benefits of technology

Accurate and real-time occlusion and collision detection in augmented reality environment, improving the rendering quality of AR scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114450717B_ABST
    Figure CN114450717B_ABST
Patent Text Reader

Abstract

Describes a technique for occlusion and collision detection in an AR scenario. For example, a depth sensor is used to generate a depth image. Distortion in the depth image is reduced or eliminated by at least dividing the depth image into multiple depth layers and moving depth pixels between these layers. An RGBD image is generated from the updated depth image and an RGB image generated substantially simultaneously with the depth image. Occlusion of virtual objects is detected based on the RGBD image. Additionally, a 3D model of the real-world environment is generated from the updated depth image and includes multi-level voxels. Collision with virtual objects is detected based on the multi-level voxels. Rendering of virtual objects in an AR scenario is based on occlusion and collision detection.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Augmented Reality (AR) superimposes virtual content on top of a user's view of the real world. With the development of AR software development kits (SDKs), the mobile industry has brought smartphone AR into the mainstream. AR SDKs typically provide six degrees-of-freedom (6DoF) tracking capabilities. A user can use a smartphone's camera to scan the environment, and the smartphone performs visual inertial odometry (VIO) in real time. Once the camera pose is continuously tracked, virtual objects can be placed into the AR scene to create the illusion of real and virtual objects merging together. The VIO system only creates a sparse representation of the real world.

[0002] When placing virtual objects into an AR scene, it is important that the placement is accurate and performed in real time. Otherwise, the rendering quality of the virtual objects will be low. Summary of the Invention

[0003] The present invention generally relates to methods and systems related to augmented reality applications. More specifically, embodiments of the present invention provide methods and systems for performing occlusion and collision detection in an AR environment. This can be applied to augmented reality and various applications in a computer's display system.

[0004] The present disclosure describes techniques for occlusion and collision detection in AR scenarios. For example, a computer system is used for occlusion and collision detection. The computer system is configured to perform various operations. The operations include generating a depth image in an augmented reality (AR) scenario and through a depth sensor of the computer system. The operations also include dividing the depth image into multiple depth layers, each depth layer corresponding to a depth range and including pixels having depth values within the depth range. The operations also include selecting a first depth layer having a first layer number and a second depth layer having a second layer number from the depth layers. The operations also include adjusting the first depth layer based on the first layer number, a first pixel in the first depth layer, the second layer number, and a second pixel in the second depth layer. The adjustment includes moving pixels from the second depth layer to the first depth layer. The operations also include updating the depth image based on the adjustment. The operations also include outputting the updated depth image to at least one AR application associated with the AR scenario.

[0005] For example, the total number of depth layers is based on the maximum depth of the depth sensor. The difference between the depth ranges of two consecutive depth layers is between 0.4 meters and 0.6 meters. The first depth layer and the second depth layer are selected based on the difference between the first layer number and the second layer number that is equal to or greater than 2. The first depth layer and the second depth layer are also selected based on the total number of first pixels and the total number of second pixels both being equal to or greater than a predefined threshold. The first layer number is greater than the second layer number. Adjusting the first depth layer includes performing morphological dilation from the first depth layer to the second depth layer. The size of the kernel for the morphological dilation is based on the difference between the first layer number and the second layer number. The morphological dilation is repeated iteratively for multiple iterations, and the number of iterations is based on the difference between the first layer number and the second layer number.

[0006] For example, the operation further includes: generating an RGB image in an AR scenario and via a red-green-blue (RGB) optical sensor of a computer system; generating an RGB-depth (RGBD) image based on the updated depth image and the RGB image; generating a set of three-dimensional (3D) points in the coordinate system of the AR scenario based on the updated depth image; generating a 3D model including a plurality of multi-level voxels. One multi-level voxel among the plurality of multi-level voxels is associated with one 3D point from the set. The operation further includes: determining a collision between a virtual object and the multi-level voxel; and rendering the virtual object in the AR scenario based on the depth of the virtual object and the RGBD image and based on the collision.

[0007] For example, a computer system includes: a depth sensor configured to generate a depth image in an augmented reality (AR) scenario; a red-green-blue (RGB) optical sensor configured to generate an RGB image in the AR scenario; one or more processors; and one or more memories storing computer-readable instructions that, when executed by the one or more processors, configure the computer system to perform an operation. The operation includes: updating the depth image by at least dividing the depth image into a plurality of depth layers and moving pixels from a first depth layer of the depth layers to a second depth layer; generating an RGB-depth (RGBD) image based on the updated depth image and the RGB image; generating a set of three-dimensional (3D) points in the coordinate system of the AR scenario based on the updated depth image; generating a 3D model including a plurality of multi-level voxels, wherein one multi-level voxel among the plurality of multi-level voxels is associated with one 3D point from the set; determining a collision between a virtual object and the multi-level voxel; and rendering the virtual object in the AR scenario based on the depth of the virtual object and the RGBD image and based on the collision.

[0008] For example, each depth layer corresponds to a depth range and includes pixels having depth values within the depth range. Updating the depth image further includes: selecting a first depth layer and a second depth layer from the depth layers based on a first layer number of the first depth layer and a second layer number of the second depth layer; and adjusting the second depth layer based on the first layer number, a first pixel in the first depth layer, the second layer number, and a second pixel in the second depth layer. The adjustment includes moving pixels from the first depth layer to the second depth layer.

[0009] For example, generating an RGBD image includes: registering a depth image with an RGB image based on an image resolution of the depth image, an image resolution of the RGB image, and a conversion between a depth sensor and an RGB optical sensor; performing depth densification on the depth image, the depth densification including a plurality of morphological dilations of the depth image; after the depth densification, filtering the depth image based on median filtering; and upsampling the filtered depth image to the image resolution of the RGB image based on the registration. Pixels in the RGBD image correspond to pixels in the RGB image and pixels in the upsampled depth image.

[0010] For example, rendering a virtual object includes: generating an alpha map from the depth image; and upsampling the depth image and the alpha map to the image resolution of the RGB image. In this example, rendering the virtual object includes: determining that a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image and a second pixel of the virtual object; determining a first depth of the first pixel from the RGBD image; determining that the first depth is less than or equal to a second depth of the second pixel; generating a smoothing factor for the first pixel based on the alpha map; and setting the RGB value of the pixel in the AR image based on a first RGB value of the first pixel, a second RGB value of the second pixel, and the smoothing factor. The smoothing factor is set to α = 1 - m i / 255, and the RGB value is set to and "α" is the smoothing factor, "i" is the pixel, "m i " is the value determined for the pixel from the alpha map, is the RGB value, "c i " is the first RGB value, and is the second RGB value.

[0011] For example, rendering a virtual object includes: determining that a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image and a second pixel of the virtual object; determining a first depth of the first pixel according to the RGBD image; determining that the first depth is greater than a second depth of the second pixel; and setting the RGB value of the pixel in the AR image to be equal to the RGB value of the second pixel.

[0012] For example, one or more non-transitory computer storage media store instructions that, when executed on a computer system, cause the computer system to perform operations. The operations include: generating a depth image in an augmented reality (AR) scenario and via a depth sensor of the computer system; generating an RGB image in the AR scenario and via red, green, and blue (RGB) optical sensors of the computer system; updating the depth image by at least dividing the depth image into a plurality of depth layers and moving pixels from a first depth layer of the depth layers to a second depth layer; generating an RGB depth (RGBD) image based on the updated depth image and the RGB image; generating a set of three-dimensional (3D) points in a coordinate system of the AR scenario based on the updated depth image; generating a 3D model including a plurality of multi-level voxels. One multi-level voxel among the plurality of multi-level voxels is associated with one 3D point from the set. The operations further include: determining a collision between a virtual object and the multi-level voxels; and rendering the virtual object in the AR scenario based on the depth of the virtual object and the RGBD image and based on the collision.

[0013] For example, the set of 3D points includes a point cloud. The multi-level voxels include first voxels at a first level having a first grid size and second voxels at a second level having a second grid size, the second grid size being smaller than the first grid size. In this example, generating the 3D model includes: dividing the coordinates of the 3D points by the first grid size to generate indices of the 3D points; hashing the indices to determine hash values; determining that a hash map does not include the hash values; and updating the hash map to include the hash values.

[0014] For example, rendering the virtual object includes: preventing the rendering collision by at least controlling the movement of the virtual object. The multi-level voxels include first voxels at a first level having a first grid size and second voxels at a second level having a second grid size, the second grid size being smaller than the first grid size. Determining the collision includes: generating one or more bounding boxes around the virtual object; determining a first intersection between the one or more bounding boxes and the first voxels; based on the first intersection, determining that the first voxels have a first hash value in the hash map; based on the first hash value being included in the hash map, determining a second intersection between the one or more bounding boxes and one of the second voxels; based on the second intersection, determining that the second voxels have a second hash value in the hash map; and detecting the collision based on the second hash value being included in the hash map.

[0015] For example, a multi-level voxel includes a first voxel having a first grid size at a first level and a second voxel having a second grid size at a second level, where the second grid size is smaller than the first grid size. In this example, determining a collision includes: storing an ordered queue including bits associated with one of the second voxels. Each bit is associated with a different depth image and indicates whether the second voxel corresponds to a 3D point visible in a different depth image. Determining a collision further includes: removing an end bit from the end of the ordered queue; inserting a start bit at the start of the ordered queue, where the start bit is associated with a depth image; determining that the total number of bits indicating the visibility of the second voxel in the ordered queue is greater than a predefined threshold; and detecting a collision based on the second voxel.

[0016] Many benefits superior to traditional techniques can be obtained through the present invention. For example, embodiments of the present invention provide methods and systems for providing accurate and real-time occlusion and collision detection with relatively low processing and storage usage. Occlusion and collision detection improve the quality of the AR scene rendered in an AR scenario. These and other embodiments of the present invention, as well as many of their advantages and features, are described in more detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Various embodiments in accordance with the present disclosure will be described with reference to the accompanying drawings, in which:

[0018] Figure 1 An example of a computer system for an AR application according to at least one aspect of the present disclosure is shown, the computer system including a depth sensor and a red, green, blue (RGB) optical sensor;

[0019] Figure 2 An example of an AR scene based on occlusion and collision rendering according to at least one aspect of the present disclosure is shown;

[0020] Figure 3 An example of an AR module for occlusion and collision rendering according to at least one aspect of the present disclosure is shown;

[0021] Figure 4 An example of depth map processing for updating a depth image according to at least one aspect of the present disclosure is shown;

[0022] Figure 5 An example of an update to a depth image according to at least one aspect of the present disclosure is shown;

[0023] Figure 6 An example of registration of a depth image and an RGB image according to at least one aspect of the present disclosure is shown;

[0024] Figure 7 An example of rendering an AR scene using an RGBD image based on occlusion detection according to at least one aspect of the present disclosure is shown;

[0025] Figure 8 An example of a three-dimensional (3D) model of a real-world environment and a bounding box of a virtual object in accordance with at least one aspect of the present disclosure is shown;

[0026] Figure 9 An example of a hash representation of a multi-level voxel in accordance with at least one aspect of the present disclosure is shown;

[0027] Figure 10 An example of a process for AR scene rendering based on occlusion and collision detection in accordance with at least one aspect of the present disclosure is shown;

[0028] Figure 11 An example of a process for updating a depth image in accordance with at least one aspect of the present disclosure is shown;

[0029] Figure 12 An example of a process for occlusion detection in accordance with at least one aspect of the present disclosure is shown;

[0030] Figure 13 An example of a process for collision detection in accordance with at least one aspect of the present disclosure is shown; and

[0031] Figure 14 An example computer system in accordance with an embodiment of the present disclosure is shown. Specific Embodiments

[0032] Various embodiments will be described in the following description. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these specific details. In addition, well-known features may be omitted or simplified in order not to obscure the described embodiments.

[0033] Embodiments of the present disclosure particularly relate to accurate and real-time detection of occlusion and collision between virtual objects and between virtual objects and real-world objects to facilitate rendering of AR scenes without visual markers or features of real-world objects. Occlusion and collision detection can rely on dense depth data in real time. However, it is challenging to generate such information from only a single red, green, blue (RGB) camera.

[0034] In embodiments of the present disclosure, depth sensors such as time-of-flight (ToF) cameras are used to acquire depth data and generate depth images. For example, a ToF camera measures the round-trip time of emitted light and resolves the depth value (distance) of points in a real-world scene. Such a camera can provide dense depth data at thirty to sixty frames per second (fps).

[0035] Applying depth data to visual occlusion and collision detection poses many key technical challenges. First, AR applications typically require real-time performance using limited computing resources. Second, ToF cameras have a unique sensing architecture and contain systematic and non-systematic biases. Specifically, the depth maps captured by ToF cameras have low depth accuracy and low spatial resolution, and also have errors caused by radiation, geometry, and illumination changes. Additionally, the depth maps need to be upsampled and registered to the resolution of the RGB camera to enable AR applications.

[0036] Embodiments of the present disclosure relate to a processing pipeline for computing visual occlusion and collision detection using RGB and ToF cameras on a computer system (e.g., a smart phone, a tablet, an AR headset, etc.). The ToF depth maps are processed to remove outliers and overcome sensor biases and errors. Then, a densification algorithm is applied to upsample the low-resolution depth maps to the resolution of the RGB images. An alpha map is also generated for blending between virtual and real objects along occlusion boundaries. A light-weighted voxelized representation of the real-world scene is also generated from the depth maps for fast collision detection. Thus, embodiments of the present disclosure describe a system that uses depth sensors (e.g., ToF cameras) on a computer system for multiple AR applications with computational efficiency and very good visual performance. The computer system is configured for depth map processing that removes outliers, densifies the depth maps, and enables blending at occlusion boundaries between real and virtual objects. The computer system is also configured to generate a light-weighted 3D representation of the scene and perform collision detection based on multi-level voxels.

[0037] Figure 1An example of a computer system 110 for an AR application according to at least one aspect of the present disclosure is shown. The computer system 110 includes a depth sensor 112 and an RGB optical sensor 114. The AR application can be implemented by the AR module 116 of the computer system 110. Generally, the RGB optical sensor 114 generates an RGB image of the real-world environment including, for example, real-world objects 130. The depth sensor 112 generates depth data regarding the real-world environment, where the data includes, for example, a depth map that shows the depth of the real-world objects 130 (e.g., the distance between the depth sensor 112 and the real-world objects 130). After the initialization of the AR scenario (where the initialization can include calibration and tracking), the AR module 116 renders an AR scene 120 of the real-world environment in the AR scenario, where the AR scene 120 can be presented at a graphical user interface (GUI) on a display of the computer system 110. The AR scene 120 shows a real-world object representation 122 of the real-world objects 130. In addition, the AR scene 120 shows virtual objects 124 that do not exist in the real-world environment. The AR module 116 can generate a red, green, blue, and depth (RGBD) image from the RGB image and the depth map to detect that at least a portion of the virtual objects 124 is occluded by the real-world object representation 122, and vice versa. The AR module 116 can additionally or alternatively generate a 3D model of the real-world environment based on the depth map, where the 3D model includes multi-level voxels. These voxels are used to detect collisions between at least a portion of the virtual objects 124 and the real-world object representation 122. The AR scene 120 can be rendered to appropriately show occlusions and avoid rendering collisions.

[0038] For example, the computer system 110 represents a suitable user device that includes, in addition to the depth sensor 112 and the RGB optical sensor 114, one or more graphics processing units (GPUs), one or more general-purpose processors (GPPs), and one or more memories that store computer-readable instructions executable by at least one of the processors to perform various functions of embodiments of the present disclosure. For example, the computer system 110 can be any one of a smart phone, a tablet computer, an AR headset, or a wearable AR device.

[0039] The depth sensor 112 has a known maximum depth range (e.g., maximum operating distance), and this maximum value can be stored locally and / or be accessible to the AR module 116. The depth sensor 112 can be a ToF camera. In this case, the depth map generated by the depth sensor 112 includes a depth image. The RGB optical sensor 114 can be a color camera. The depth image and the RGB image can have different resolutions. Generally, the resolution of the depth image is less than that of the RGB image. For example, the depth image has a resolution of 240×180, while the RGB image has a resolution of 1920×1280.

[0040] In addition, the depth sensor 112 and the RGB optical sensor 114 installed in the computer system 110 can be separated by a transformation (e.g., distance offset, field of view angle difference, etc.). This transformation can be known and its value can be stored locally and / or be accessible to the AR module 116. When using cameras, the ToF camera and the color camera can have similar fields of view. However, due to this transformation, the fields of view will be partially overlapped rather than completely overlapped.

[0041] The AR module 116 can be implemented as dedicated hardware and / or a combination of hardware and software (e.g., a general-purpose processor and computer-readable instructions stored in a memory and executable by the general-purpose processor). In addition to initializing the AR scenario and performing VIO, the AR module 116 can detect occlusions and collisions to appropriately render the AR scene 120.

[0042] In Figure 1 the illustrative example, a smart phone is used to show an AR scenario of the real-world environment. In particular, the AR scenario includes rendering an AR scene that includes a representation of a real-world table, on top of which a vase (or some other real-world object) is placed. A virtual ball (or some other virtual object) will be shown in the AR scene. In particular, the virtual ball will also be shown on top of the table. By tracking the occlusion between the virtual ball and the virtual vase (representing the real-world vase), when the pose of the virtual ball relative to the smart phone is behind the virtual vase, the virtual vase can occlude in a part of the AR scene. In other parts of the AR scene, when the change in the pose of the virtual vase relative to the smart phone is behind the virtual ball, the virtual ball can occlude the virtual vase. And in the remaining part of the AR scene, there is no occlusion. Additionally, the user of the smart phone can interact with the virtual ball to move the virtual ball on the top surface of the virtual table (representing the real-world table). By tracking the possible collisions between the virtual ball and the virtual objects, any interaction that may cause a collision will not be rendered. In other words, collision tracking can be used to control where the virtual ball can move in the AR scene.

[0043] Figure 2An example of an AR scene based on occlusion and collision rendering according to at least one aspect of the present disclosure is shown. As shown, different AR scenes are possible, but only one scene can appropriately show occlusion and avoid collisions. Here, the virtual object is described as being occluded by the real-world object representation. In addition, collisions between the virtual object and the real-world object representation are described. However, embodiments of the present disclosure are not limited thereto and are applicable to occlusion of real-world object representations by virtual objects and / or collisions between real-world object representations and virtual objects. Embodiments are similarly applied to occlusion between virtual objects and / or collisions between virtual objects.

[0044] As Figure 2 shown in the upper left side of , a first AR scene can be rendered, where the AR scene does not consider occlusion 210. In particular, the virtual object (shown as a sphere) should be occluded by the real-world object representation (shown as a cylinder). However, in the first AR scene, the virtual object is incorrectly rendered with a smaller depth than the real-world object representation and thus incorrectly appears to occlude the real-world object representation.

[0045] As Figure 2 shown in the upper right side of , a second AR scene can be rendered, where the AR scene does not consider collision 220. In particular, the virtual object (also shown as a sphere) should not collide with the real-world object representation (also shown as a cylinder). However, in the second AR scene, the virtual object incorrectly has the same depth as the real-world object representation and is in the same virtual space as the real-world object representation and thus incorrectly appears to collide with the real-world object representation.

[0046] As Figure 2 shown in the bottom center of , a correct AR scene 230 can be rendered, where the correct AR scene 210 considers occlusion and collision. AR scene 230 is Figure 1 an example of the AR scene 120 of . In the correct AR scene 230, the virtual object (also shown as a sphere) is shown as being occluded by the real-world object representation (also shown as a cylinder). Additionally, the virtual object is shown as not colliding with the real-world object representation. Embodiments of the present disclosure relate to occlusion and collision detection to support the rendering of correct AR scenes, such as correct AR scene 230.

[0047] Figure 3 An example of an AR module 300 for occlusion and collision rendering according to at least one aspect of the present disclosure is shown. AR module 300 is Figure 1An example of the AR module 116 and includes a plurality of computing components, such as a preprocessing component 310, a depth upsampling component 320, a visual occlusion component 330, a fast voxelization component 340, a collision detection component 350, and a rendering component 360. Each of these computing components 310 to 360 can be implemented as dedicated hardware and / or a combination of hardware and software.

[0048] The preprocessing component 310 processes a depth map (e.g., a depth image generated based on measurements using a ToF camera) to remove outliers. This processing is also shown in Figure 4 and Figure 5 The processed depth map is further processed in two processes 305 and 307. Processes 305 and 307 can be executed in parallel to reduce processing waiting time. In process 305, depth upsampling is performed by the depth upsampling component 320 to generate a high-resolution depth map. This high-resolution map is used by the visual occlusion component 330 to detect occlusions, and the output of the occlusion detection can be provided to the rendering component 360 for occlusion rendering. Figure 6 and Figure 7 The processing of process 305 is also shown.

[0049] In process 307, the fast voxelization component 340 converts the real-world scene into a 3D representation for collision detection, where this conversion depends on the processed depth map. The 3D representation includes multi-level voxels used by the collision detection component 350 to detect collisions, and the output of the collision detection can be provided to the rendering component 360 for collision rendering (e.g., to avoid collisions). Figure 8 and Figure 9 The processing of process 307 is also shown.

[0050] Figure 4 Shows an example of depth map processing for updating the depth image 400 according to at least one aspect of the present disclosure. Here, the depth image 400 is generated by a ToF camera, which is an example of a depth sensor. This update can be performed by, for example, Figure 3 the preprocessing component 310.

[0051] Specifically, both visual occlusion and collision detection require each pixel of the RGB image to have a reasonable depth value. However, the depth data from a ToF camera is often very noisy due to systematic errors and non-systematic errors. Specifically, systematic errors include infrared (IR) demodulation errors, amplitude ambiguity, and temperature errors. Generally, a longer exposure time increases the signal-to-noise ratio (SNR); however, this will reduce the frame rate.

[0052] In a typical AR application, the user often moves the ToF camera slowly. Therefore, outliers due to IR saturation and 3D structure distortion are dominant. Such outliers exist discontinuously along the depth between the foreground and the background. Specifically, pixels on background objects along the occlusion boundary tend to have abnormally small depth values. The larger the depth gap between the background and the foreground, the larger the affected area. Therefore, morphological-based image processing can be used to remove such outliers.

[0053] To process the foreground and the background differently, image segmentation is often used. However, precise segmentation is an expensive process. For computational efficiency, thresholding is used to divide the depth image 400 into multiple layers.

[0054] For example, the depth image 400 is divided into a number of depth layers, each having a layer number. The total number of depth layers depends on various factors. One factor is the maximum depth range of the ToF camera. Another factor is the thresholding. This factor can be used to control the depth range of each depth layer such that the depth layer represents a container including pixels having depth values within that depth range.

[0055] For example, the maximum depth range is three meters and the threshold is set to 0.5 meters (or a value between 0.4 meters and 0.6 meters). In this illustration, six layers will be created and the difference between two consecutive layers is 0.5 meters (the value of the thresholding). The first layer will include pixels with depths between 0 and 0.5 meters, the next layer will include pixels with depths between 0.5 meters and 1 meter, and so on, until the last layer including pixels with depths between 2.5 and 3.0 meters.

[0056] Additionally, if a layer has a pixel count less than a predefined threshold t pixel , then the layer can be ignored. Doing so can accelerate the processing of the depth image 400. Referring to the above description, if the fifth and sixth layers each include fewer than twenty pixels (or some other predefined threshold t pixel ), then these two layers are removed.

[0057] As Figure 4 shown, the resulting division of the depth image 400 includes four layers (shown as L1, L2, L3, and L4), where the first layer L1 has the minimum depth and the layer L4 has the maximum depth. In the above description, the layer L1 has a depth range between 0 and 0.5 meters, the layer L2 has a depth range between 0.5 meters and 1 meter, the layer L3 has a depth range between 1 meter and 1.5 meters, and the layer L4 has a depth range between 1.5 and 2.0 meters. The last two layers are removed due to having insufficient numbers of pixels.

[0058] As Figure 4As further shown, outliers exist at the boundaries between the layers and occur more frequently as the layers are spaced farther apart. For example, the frequency of outliers and the possible regions of outliers are much larger at the boundary between layer L1 and L4 than at the boundary between layer L2 and L4, and may not exist between layer L3 and L4. The boundaries including outliers are shown with diagonal shading.

[0059] When considering layer L1 and L4, the pixels in the shaded boundary have incorrect depth values (due to sensor errors as explained above). The depth values of these pixels are within the depth range of the first layer L1 (e.g., between 0 and 0.5 meters). However, in reality, the depth values of these pixels should be within the depth range of the fourth layer L4 (e.g., between 1.5 and 2.0 meters). Similarly, the pixels in the shaded boundary between the second layer L2 and the fourth layer L4 are incorrectly sensed as belonging to the second layer L2, while in fact they should belong to the fourth layer L4.

[0060] (e.g., by the preprocessing component 310) Update the depth image 400 to reduce or eliminate outliers. The update includes moving the pixels in the shaded boundary from the first layer L1 or the second layer L2 to the fourth layer L4.

[0061] Generally, each layer only contains pixels within a specific depth range. The thickness of each layer is λ = d max / l, where d max is the maximum working distance of the ToF camera and l is the number of layers. Each depth layer has a layer number, and the layer numbers are sorted in ascending order (e.g., L1 is the nearest layer and L l is the farthest layer). As Figure 4 shown, the diagonal shaded area represents the distortion area along the depth edge. If both sides of the depth edge are on consecutive layers (e.g., L1 and L2), the depth edge has less depth distortion and a smaller noise area. If the two sides are on non - consecutive layers, the depth distortion may be severe depending on the gap (e.g., the difference between the depth values of the foreground and the background). As Figure 4 shown, the distortion area between L1 and L4 is larger than the distortion area between L2 and L4. When dividing the depth map into layers, if the number of pixels on a layer is less than the threshold t pixel , then that layer is ignored. Based on the gaps between different depth layers, morphological dilation is performed on the depth layers from far to near. The purpose of dilation is to propagate the depth values from the far layer to the distortion areas on the nearer depth layers, which should belong to the far depth layer if not for the depth distortion problem.

[0062] Figure 5An example of updating a depth image according to at least one aspect of the present disclosure is shown. For example, the update involves performing morphological dilation on depth layers from far to near (e.g., from a relatively large depth range to a relatively small depth range). However, embodiments of the present disclosure are not limited thereto. For example, the update may additionally or alternatively include performing morphological erosion on depth layers from near to far.

[0063] In one embodiment, the update involves a set of update rules. The first update rule specifies that morphological dilation will be performed on depth layers from far to near. The second update rule specifies that depth distortions between consecutive depth layers can be ignored. In other words, when two depth layers are selected for morphological dilation, only non-consecutive layers can be selected (e.g., the difference between the layer numbers of the selected depth layers is equal to or greater than 2). The third update rule specifies that a depth layer having a number of pixels less than a threshold t pixel can be ignored. The fourth update rule specifies that the size of the kernel for morphological dilation can depend on the difference between the layer numbers of the selected layers. The fifth update rule specifies that for iterative application of the dilation operation on two selected layers, the size of the kernel decreases with the number of iterations. The sixth update rule specifies that morphological dilation can be iteratively applied to different pairs of layers in the selected depth layers.

[0064] As Figure 5 shown, the true depth image 510 (e.g., an undistorted depth image) is divided into a plurality of depth layers including layers L i and L j . In contrast, the received depth image 510 (e.g., the same depth image generated by a ToF camera and including edge distortion) can be similarly divided into a plurality of layers including layers L i and L j . However, due to edge distortion, layer L i in the received depth image 520 is different from (e.g., smaller than) layer L i in the true depth image 510. Similarly, layer L j in the received depth image 520 is different from (e.g., larger than) layer L j in the true depth image 510. The purpose of the update is to adjust layers L i and L j in the received depth image 520 to be close to layers L i and L j in the true depth image 510, respectively. This adjustment includes moving pixels of the distorted edge from one layer to another.

[0065] The update can start with selecting layers L i and L j。The difference between layer numbers (e.g., i - j) should be equal to or greater than 2. In this example, layer L i is deeper than layer L j Next, a morphological dilation operation is applied to layer L i where the size of the kernel is based on the difference "i - j", and this operation produces an intermediate processed image 540, where layer L i dilates while layer L j contracts. A masking operation 545 is applied to the processed image 540. In particular, before the dilation operation 530, a non - zero mask corresponding to layer L j is applied to the processed image 540. The result of the masking operation 545 is another intermediate processed image 550. A comparison operation 560 is applied to the processed image 550, whereby this processed image 550 can be compared with the received depth image 520 to determine changes to layer L i and L j The result of the comparison operation 560 is another processed image 570, and the changes to layer L i and L j are shown as shaded regions in the processed image 570. These different operations are repeated for different pairs of alternative layers. Figure 5 The update operation 580 in

[0066] corresponds to this iterative process. Once the iterative process is complete, an updated depth image 590 is generated. This updated depth image 590 approximates the true depth image 510.

[0067] Data: The depth image D is divided into l layers L1, L2, …, L l , and corresponding non - zero masks M1, M2, …, M l .

[0068]

[0069] In as combined with Figure 4 and Figure 5After the outliers shown are removed, the depth map is further processed for visual occlusion handling. Due to hardware limitations, ToF cameras typically have low spatial resolution. Therefore, embodiments of the present invention perform a registration operation to register depth samples to the frames of a high-resolution RGB image, and then densify the sparse depth map. The depth edges between occluding objects and the background need to be smooth in time while maintaining good discontinuity. Finally, this processing should be done in real time with limited computational resources. To overcome these challenges, the processing pipeline (e.g., corresponding to Figure 3 the upper process in

[0070] Figure 6 includes registration, morphology-based depth densification, smoothing filtering, upsampling, and alpha blending.

[0071] As shown, a depth sensor 610 and an RGB optical sensor 620 are mounted in a computer system. A transformation 630 exists between the depth sensor 610 and the RGB optical sensor 620. Although the two sensors 610 and 620 may have similar fields of view (FOVs), due to the transformation 630, their FOVs overlap only partially rather than completely. Figure 6 The overlap is shown as an FOV overlap 615 between the two innermost dashed lines.

[0072] The depth sensor 610 and the RGB optical sensor have different image resolutions. In other words, the depth sensor 610 generates a depth image 612, and the RGB optical sensor 620 generates an RGB image 622, where the depth image 612 has a lower image resolution than the RGB image 622. Figure 6 The lower resolution is shown by depicting the depth pixels 614 of the depth image 612 as sparse.

[0073] Due to the partial FOV overlap 615, the depth image 612 and the RGB image 622 also overlap only partially rather than completely. Figure 6 The overlap is shown as an image overlap 650 between the depth image 612 and the RGB image 622.

[0074] For example, registration only considers the depth pixels 614 that fall within the image overlap 650. Those outside the image overlap 650 are ignored (e.g., Figure 6Depth pixels to the left of the mid-image overlap 650. Similarly, RGB pixels falling within the image overlap 650 are considered and the remaining RGB pixels are ignored. The position (e.g., its pixel index) of the considered depth pixel 614 in the depth image 612 is associated with the position (e.g., its pixel index) of the considered RGB pixel in the RGB image 622 based on their correspondence. For example, when depth pixel "m" and RGB pixel "n" overlap in the image overlap 650, then these two pixels are associated. More specifically, based on the coordinates of pixel "m", the depth value of "m", and the intrinsic parameters of the depth camera, the depth pixel "m" is projected to 3D coordinates M. Then, based on the extrinsic transformation between the ToF camera and the RGB camera, the 3D point M is transformed to 3D point N. Then, based on the intrinsic parameters of the RGB camera, the 3D point N is projected onto the RGB pixel "n". The intrinsic parameters of the ToF camera and the RGB camera and the extrinsic transformation between the ToF camera and the RGB camera can be established during the device calibration step.

[0075] For example, a ToF camera is used and the ToF camera has a low resolution of 240×180. An RGB camera is also used to generate an RGB image at a higher resolution (1920×1280). To complete the entire pipeline in real time, the depth image is registered with the downsampled RGB image at 480×320.

[0076] Once the registration is complete (e.g., associations between depth pixels and RGB pixels are generated), a depth densification operation can be applied. For example, a computationally fast unguided depth upsampling method is used. The depth densification operation includes three morphological operations. First, dilation is performed with a diamond kernel to fill most of the empty pixels. Then, a complete kernel morphological closing operation is applied to fill most of the holes. Finally, to fill larger holes (which are usually very rare), a large full kernel dilation is performed. The kernel size can be carefully tuned based on different ToF cameras.

[0077] Thereafter, a filtering operation is applied. In particular, during densification, the morphological operations may generate incorrect depth values. Therefore, smoothing can be used to remove noise while preserving local edge information. A median filter can be used for this purpose. Simple depth thresholding is also used to generate a foreground mask for occluding objects. Then, Gaussian blur is applied to the mask to create an alpha map and also to smooth the edges in the depth image.

[0078] The filtered depth image (480x320) is upsampled to full RGB resolution (1920x1080) for visual occlusion rendering. This can be done in a GPU with nearest interpolation. At the same time, the alpha map is also scaled to full resolution.

[0079] During the rendering step, using the full-resolution depth image and the alpha map, alpha blending is used to synthesize the final image. An example of occlusion rendering is also shown in conjunction with the next figure.

[0080] Figure 7 An example of rendering an AR scene based on occlusion detection using an RGBD image 710 according to at least one aspect of the present disclosure is shown. For example, the RGBD image 710 represents a depth image that has been registered, depth-densified, filtered, and upsampled to the resolution of the RGB image. Each pixel in the RGBD corresponds to an RGB pixel of the RGB image and corresponds to a depth pixel of the depth image based on registration. Thus, each RGB pixel has the RGB value of the corresponding RGB pixel and the depth value of the corresponding depth pixel. The depth value in the RGBD image 710 is compared with the depth value of the virtual object 720, and if occlusion is detected based on the depth comparison, the alpha map 730 can be used for edge smoothing.

[0081] In particular, occlusion rendering involves a blending operation 750. For example, the RGBD image 710, the virtual object 720, and the alpha map 730 are input into the blending operation 750. This operation compares the depth values of the RGBD image 710 and the virtual object 720 for overlapping pixels.

[0082] When a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image 710 and a second pixel of the virtual object 720 (the first pixel and the second pixel are the same in the rendering buffer, the depths of the two pixels are compared to determine whether the second pixel should be occluded in the rendering. The depth of the first pixel can be determined from the RGBD image 710, the depth of the second pixel can be retrieved from the buffer and can be defined by the AR application. Then, the blending operation 750 compares this depth with the depth of the second pixel. If it is equal to or less than the depth of the second pixel, the first pixel occludes the second pixel. In this case, the blending operation 750 generates a smoothing factor for the first pixel based on the alpha map. The RGB value for the pixel in the AR image is set based on the first RGB value of the first pixel, the second RGB value of the second pixel, and the smoothing factor. For example, the smoothing factor is set to α = 1 - m i / 255, where the RGB value is set to and "α" is the smoothing factor, "i" is the pixel, "m i " is the value determined for the pixel from the alpha map, is the RGB value, "c i " is the first RGB value, and is the second RGB value. However, if it is greater than the depth of the second pixel, the first pixel does not occlude the second pixel. In this case, the blending operation 750 sets the RGB value for the pixel in the AR image to be equal to the RGB value of the second pixel (e.g., α = 1).

[0083] For example, the above rendering can be defined in an algorithm that can be implemented on a GPU. The algorithm can be expressed as:

[0084] Data: For pixel i: d i is the depth value from the ToF camera; m i is the alpha map value; c i is the color from the RGB camera; is the depth of the virtual object from the depth buffer; is the shadow color of the virtual object. is the final color of the current pixel i.

[0085]

[0086] Figure 8 Shows an example of a bounding box 860 of a 3D model 810 of a real-world environment and a virtual object 850 according to at least one aspect of the present disclosure. In particular, after outlier removal, the depth image is used to generate the 3D model 810 in the coordinate system 820 of the AR scenario. The 3D model includes multi-level voxels, where each such voxel includes a first voxel 830 in a first layer and a plurality of voxels 840 in a second layer, where the first layer has a lower resolution than the second layer. The bounding box 860 can define the virtual object with a certain margin. Fast collision detection is performed by determining whether the bounding box 860 overlaps with any first-level voxels 830, and if so, the second-level voxels 840 of the overlapping first-level voxels 830 are also considered to detect the collision position at a higher resolution. Rendering collisions can be prevented, thus, for example, prohibiting the virtual object 850 from moving to the collision position. Although a single bounding box 860 is shown in Figure 8 multiple bounding boxes are equally possible. For example, multiple bounding boxes can enclose the virtual object 850, can but need not be centered at the same point (e.g., the center of the virtual object 850) and / or can but need not have different sizes.

[0087] Figure 9 Shows an example of a hash representation 910 of multi-level voxels according to at least one aspect of the present disclosure. The multi-level voxels include respectively associated with Figure 8A first-level voxel 830 and a plurality of second-level voxels 840 that are similar to the first-level voxel. Here, hashing can accelerate collision detection by comparing the hash value of the voxel with the hash value of the position of one or more bounding boxes around the virtual object to determine whether there is an overlap.

[0088] For example, the hash representation 910 is defined as a hash map. The hash function 920 is applied to the first-level voxel 925. The resulting hash value is stored in the hash map and used as the spatial index 925 of the first-level voxel. Similarly, the hash function 930 is applied to the second-level voxel 935. The resulting hash value is also stored in the hash map and used as the spatial index of the second-level voxel 935.

[0089] There are some representations of 3D data, such as point clouds. The depth data captured by a depth sensor can be used for a fast 3D representation, and the data structure will support fast collision detection. As combined with Figure 8 and Figure 9 described, voxels can be used to represent the rough shape of a 3D scene, which reduces memory storage while also providing a faster construction speed and supporting fast collision detection. Although the data structure lacks geometric details, generally, for collision detection, detailed geometry may not be required unless an accurate response is needed.

[0090] For example, a cube is used as the unit voxel of the proposed data structure. The resolution of the data structure can be adjusted by changing the size of the unit voxel "c". As Figure 8 shown, a two-level voxel data structure is generated. The first level of this structure stores large voxels, which are indexed by a spatial hash function as Figure 9 shown, and the spatial hash function can be defined as: where, is an XOR operation, p1, p2, and p3 are large prime numbers (for example, p1 = 73856093; p2 = 19349663; p3 = 83492791), and n is the size of the hash table.

[0091] In the second level, each large voxel is subdivided into smaller voxels with a resolution of m * n * l. Then each small voxel is also indexed by a regular hash function.

[0092] When the AR scenario starts, the user scans the environment by moving their computer system around. Once the Simultaneous Localization and Mapping (SLAM) operation is successfully initialized, the 6DoF pose of the ToF camera is continuously tracked. 3D data structure reconstruction can be performed on each ToF frame at 30fps. Using the ToF camera pose, the ToF depth frame is first converted to a point cloud in the coordinate system of the AR scenario. Meanwhile, a plane detection step is performed to detect the horizontal support plane using a plane detection algorithm based on Random Sample Consensus (RANSAC). The support plane is where the 3D model will be placed. 3D point samples belonging to the plane can be removed to accelerate voxelization and reduce data storage. By using the motion sensing hardware on the computer system, the direction of gravity can be obtained. This enables the horizontal plane to be effectively found.

[0093] Then, for each remaining point, its coordinates (x, y, z) are divided by the first-level grid cell size c and rounded down to integer indices (i, j, k). Then, the integer indices are hashed using the spatial hashing function described above to check if the first-level voxel exists (e.g., is indexed in the hash map). If it does not exist, a new voxel is generated and the hash map is updated to include the hash value. If the voxel exists, the (x, y, z) coordinates are converted and rounded to integer indices of the second level: (i', j', k'). This index is also hashed to check if the second-level voxel exists (e.g., is indexed in the hash map). If it does not exist, a second-level voxel is generated and its hash value is stored in the hash map.

[0094] To improve robustness and temporal consistency, a queue with S bits is stored in each second-level voxel. Each bit stores a binary value to indicate whether the voxel has been "seen" by the current ToF frame or a past ToF frame (e.g., a sorted queue including multiple bits, where each bit is associated with a different depth image and indicates whether the second-level voxel corresponds to a 3D point visible in a different depth image). A "1" value can indicate the seen state. When processing a ToF frame, the oldest bit is popped from the queue and a new bit is inserted (e.g., the ending bit from the end of the sorted queue is removed and the starting bit is inserted at the beginning of the sorted queue). If the number of "1" bits is greater than the threshold t s , then the voxel is used for collision detection in the current frame.

[0095] This two - level data structure can be used for fast collision detection because each voxel represents an axis - aligned bounding box (AABB). In AR applications, virtual objects can also be represented by AABBs. During collision detection, first, all first - level voxels that may intersect the AABB of the virtual object are found. Then, a spatial hash function is used to look up these voxels in the hash map. Each lookup can be done in constant time. If a voxel exists in the map, the second - level valid voxels are checked for collision detection. All m*n*l voxels are iterated to check if such a voxel exists in the hash map. If any voxels exist, an intersection test is performed between the second - level voxels and the AABB of the virtual object. To improve robustness, a collision is detected only when the number of colliding voxels is greater than a threshold t c When a collision between the static scene and the moving virtual object is detected, the movement of the object is stopped to simulate the visual effect of collision avoidance.

[0096] Figures 10 to 13 An example process for occlusion and collision detection and rendering according to an embodiment of the present disclosure is shown. The process is described in connection with a computer system 110 as an example of a Figure 1 computer system. Some or all of the operations of the process can be implemented via specific hardware on the computer system and / or can be implemented as computer - readable instructions stored on a non - transitory computer - readable medium of the computer system. The stored computer - readable instructions represent programmable modules that include code executable by a processor of the computer system. Execution of these instructions configures the computer system to perform the corresponding operations. Each programmable module in combination with the processor represents a means for performing the corresponding operation. Although the operations are shown in a particular order, it should be understood that the particular order is not required and one or more operations can be omitted, skipped, and / or reordered.

[0097] Figure 10 An example of a process for AR scene rendering based on occlusion and collision detection according to at least one aspect of the present disclosure is shown. For example, the process starts at operation 1002, where the computer system generates a depth image. For example, the computer system includes a ToF camera, and the ToF camera is operated to generate a depth image in an AR scenario.

[0098] For example, the process includes operation 1004, where the computer system generates an RGB image. For example, the computer system includes an RGB camera, and the RGB camera is operated to generate an RGB image in an AR scenario. The depth image and the RGB image can be generated simultaneously or substantially simultaneously (e.g., within an acceptable time difference relative to each other, e.g., within a few milliseconds).

[0099] For example, the process includes operation 1006, where the computer system updates the depth image. For example, preprocessing of the depth image is performed to remove outliers by dividing the depth image into multiple depth layers and moving at least one pixel from a first depth layer of these depth layers to a second depth layer. Typically, the update is iterative between different layers and follows a set of update rules as described in conjunction with Figure 4 and Figure 5 .

[0100] For example, the process includes operation 1008, where the computer system generates an RGBD image. For example, the RGBD image is generated based on the updated depth image and the RGB image. In particular, as described in conjunction with Figure 6 , registration, depth densification, filtering, and upsampling are performed on the depth image. Each pixel in the RGBD image corresponds to a pixel in the RGB image and a pixel in the updated depth image. The RGBD pixel has the RGB value of the RGB pixel and the depth value of the depth pixel.

[0101] For example, the process includes operation 1010, where the computer system determines occlusion between a virtual object and the RGBD image. For example, the depth of each RGBD pixel (or a set of RGBD pixels overlapping with the virtual object) is compared with the depth of the virtual object. If the depth of the RGBD pixel is less than or equal to the depth of the virtual object, the RGBD pixel occludes the virtual object. Subsequently, a smoothing factor is set based on the alpha map.

[0102] For example, the process includes operation 1012, where the computer system generates a 3D model. For example, the computer system generates a set of 3D points, such as a point cloud (as described in conjunction with Figure 8 ) in the coordinate system of the AR scenario. Each 3D point corresponds to a depth pixel. For each 3D point, a multi-level voxel can be defined and each voxel can be indexed using a hash value (as described in conjunction with Figure 9 ).

[0103] For example, the process includes operation 1014, where the computer system determines a collision between a virtual object and another object in the AR scene (e.g., one shown in the RGBD image). For example, one or more bounding boxes are defined around the virtual object. A collision between the bounding box and the first-level voxel triggers the detection of second-level voxels that collide with the bounding box.

[0104] For example, the process includes operation 1016, where the computer system renders an AR image based on occlusion determination and collision determination. For example, the computer system renders a virtual object in the AR scene of the AR scenario based on the depth of the virtual object, the RGBD image, and the collision. In particular, when the virtual object is deeper than some RGBD pixels, a smoothing factor is applied to a given alpha map. In addition, when a collision is detected, the movement of the virtual object can be stopped to simulate the visual effect of avoiding the collision.

[0105] Figure 11 An example of a process for updating a depth image according to at least one aspect of the present disclosure is shown. The process can be implemented as Figure 10 a sub-operation of operation 1006.

[0106] For example, Figure 11 the process begins at operation 1102, where the computer system generates a depth image. For example, the process includes operation 1104, where the computer system divides the depth image into a plurality of depth layers. Each depth layer corresponds to a depth range and includes pixels having depth values within the depth range. If the total number of pixels included in a depth layer is less than a predetermined threshold, that depth layer can be ignored from the remaining operations of the process.

[0107] For example, the process includes operation 1106, where the computer system selects a first depth layer and a second depth layer. The first depth layer has a first layer number. The second depth layer has a second layer number. The selection can be based on, for example, the layer numbers of the depth layers. In particular, the selection rule can specify that two consecutive depth layers cannot be selected. In this case, the difference between the first layer number and the second layer number is equal to or greater than 2.

[0108] For example, the process includes operation 1108, where the computer system adjusts the first layer. Different adjustment operations are possible. For example, morphological dilation is possible. In this case, the first layer number is greater than the second layer number (e.g., the first depth layer is deeper than the second depth layer), and the morphological dilation operation is applied to the depth layers from far to near. In another example, morphological erosion is possible. In this case, the first layer number is less than the second layer number (e.g., the second depth layer is deeper than the first depth layer), and the morphological erosion operation is applied to the depth layers from near to far. The size of the kernel can depend on the difference between the layer numbers. The adjustment can be repeated for different pairs of selectable layers.

[0109] For example, the process includes operation 1110, where the computer system updates the depth image. For example, once the morphological dilation operation (and / or morphological erosion operation) is completed, the adjusted layers form the updated depth image.

[0110] For example, the process includes operation 1112, where the computer system outputs a depth image to at least one AR application. For example, the updated depth image is sent to a first application pipeline for detecting occlusions. The updated depth image is also sent to a second application pipeline for detecting collisions.

[0111] Figure 12 An example of a process for occlusion detection according to at least one aspect of the present disclosure is shown. The process can be implemented as Figure 10 sub-operations of operations 1008 to 1010.

[0112] For example, Figure 12 the process begins at operation 1202, where the computer system registers a depth image with an RGB image. For example, the registration depends on a known transformation between a TOF camera and an RGB camera. In particular, the overlap between the two images is determined, and only the pixels falling within the overlap are considered. For these pixels, the computer system determines the overlapping pairs of depth pixels and RGB pixels. The index of the depth pixel in the depth image is associated with the index of the RGB pixel in the RGB image with which the depth pixel and the RGB pixel overlap.

[0113] For example, the process includes operation 1204, where the computer system performs depth densification on the depth image. For example, one or more morphological dilation operations are applied to the depth image.

[0114] For example, the process includes operation 1206, where the computer system filters the depth image after depth densification and generates an alpha map. For example, a median filter is applied. A foreground mask is also applied to the depth image, and Gaussian blur is applied to the mask to generate the alpha map.

[0115] For example, the process includes operation 1208, where the computer system upsamples the depth image after filtering. For example, the depth image is upsampled to the resolution of the RGB image. Similarly, the alpha map is upsampled to the resolution of the RGB image.

[0116] For example, the process includes operation 1210, where the computer system detects occlusions. For example, the depth of each pixel from the upsampled depth image (or the set of depth pixels overlapping with the virtual object) is compared with the depth of the virtual object. If the depth pixel is deeper than the virtual object, an occlusion is detected.

[0117] For example, the process includes operation 1212, where the computer system renders pixels in the AR image based on occlusion detection. For example, the occlusion detection identifies different depth pixels that occlude the virtual object. Based on this registration, the corresponding RGB pixels are determined. A smoothing factor is set based on the values of these RGB pixels from the alpha map. Rendering is performed according to the value of the smoothing factor, the RGB pixels of the RGB image, and the RGB pixels of the virtual object.

[0118] Figure 13 An example of a process for collision detection according to at least one aspect of the present disclosure is shown. The process can be implemented as Figure 10 sub-operations of operations 1012 to 1014.

[0119] For example, Figure 13 the process begins with operation 1302, where the computer system generates a 3D model including a plurality of multi-level voxels. For example, the depth image updated according to Figure 11 is converted into a point cloud in the coordinate system of the AR scenario. Each coordinate (x, y, z) of the point cloud is used to define the first-level voxels and the second-level voxels (if not present) at a higher resolution.

[0120] For example, the process includes operation 1304, where the computer system updates the hash map. For example, for each voxel (at the first level or the second level), the coordinates (x, y, z) are divided by the resolution of the voxel level and rounded to generate an index, and a hashing operation is applied to the index. The resulting hash value is looked up in the hash map, and if not present, the hash map is updated to include the hash value.

[0121] For example, the process includes operation 1306, where the computer system updates the sorting queue. For example, the sorting queue is stored in each second-level voxel and contains bits with binary values. A '1' bit indicates that the second-level voxel corresponds to the visible part at the moment of capturing the past ToF frame. A '0' bit indicates otherwise. When processing the ToF image, the oldest bit is removed, and the latest bit corresponding to the current ToF image is inserted into the sorting queue. Only when the number of '1' bits is greater than a predetermined threshold, the second-level voxels are considered for collision detection.

[0122] For example, the process includes operation 1308, where the computer system detects a collision. For example, one or more bounding boxes are defined around the virtual object. The computer system locates all first-level voxels that potentially intersect the bounding box. The computer system then uses a hash function applied to the first-level voxels to look up these candidate voxels in a hash map. If there is a voxel in the hash map, the second-level voxels included in the first-level voxel are checked for collision detection based on the corresponding hash values of the second-level voxels in the hash map. If any voxels exist, an intersection test is performed between the second-level voxels and the bounding box of the virtual object. A collision is detected only if the number of colliding voxels is greater than a threshold.

[0123] For example, the process includes operation 1310, where the computer system renders pixels in the AR image based on the collision detection. For example, the collision detection identifies different depth pixels that potentially collide with the virtual object. Based on this registration, the corresponding RGB pixels are determined. Rendering is performed to avoid placing the object in a way that overlaps these RGB pixels.

[0124] Figure 14 An example of components of a computer system 1400 according to a particular embodiment is shown. Computer system 1400 is Figure 1 an example of computer system 110, and although these components are shown as belonging to the same computer system 1400, computer system 1400 can also be distributed.

[0125] Computer system 1400 includes at least a processor 1402, a memory 1404, a storage device 1406, an input / output (I / O) peripheral 1408, a communication peripheral 1410, and an interface bus 1412. Interface bus 1412 is configured to communicate, send, and transfer data, control, and commands among the various components of computer system 1400. Memory 1404 and storage device 1406 include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard disk drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage such as memory, and other tangible storage media. Any such computer-readable storage media can be configured to store instructions or program code embodying aspects of the present disclosure. Memory 1404 and storage device 1406 also include computer-readable signal media. A computer-readable signal media includes a propagated data signal that includes computer-readable program code therein. Such a propagated signal takes any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signal media includes any computer-readable media that is not a computer-readable storage media and that can communicate, propagate, or transmit a program for use in conjunction with computer system 1400.

[0126] In addition, the memory 1404 includes an operating system, programs, and applications. The processor 1402 is configured to execute the stored instructions and includes, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. The memory 1404 and / or the processor 1402 may be virtualized and may be hosted within another computer system such as a cloud network or a data center. The I / O peripheral devices 1408 include user interfaces such as a keyboard, a screen (e.g., a touch screen), a microphone, a speaker, other input / output devices, and computing components such as a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripheral devices. The I / O peripheral devices 1408 are connected to the processor 1402 through any of the ports coupled to the interface bus 1412. The communication peripheral device 1410 is configured to facilitate communication between the computer system 1400 and other computing devices through a communication network and includes, for example, a network interface controller, a modem, wireless and wired interface cards, an antenna, and other communication peripheral devices.

[0127] Although the subject matter has been described in detail with respect to specific embodiments of the subject matter, it will be understood that those skilled in the art can readily generate alterations, variations, and equivalents of such embodiments upon obtaining an understanding of the foregoing. Accordingly, it should be understood that the present disclosure is presented for purposes of illustration and not limitation and does not exclude including such modifications, variations, and / or additions to the subject matter that would be readily apparent to a person of ordinary skill in the art. Indeed, the methods and systems described herein may be implemented in a variety of other forms; furthermore, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the present disclosure. The appended claims and their equivalents are intended to cover such forms or modifications that would fall within the scope and spirit of the present disclosure.

[0128] Unless otherwise specifically stated, it should be understood that in this specification, discussions using terms such as "processing," "computing," "determining," and "identifying" refer to actions or processes of, for example, one or more computers or one or more similar electronic computing devices that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, a register, or other information storage device, a transmission device, or a display device of a computing platform.

[0129] One or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include microprocessor-based multipurpose computer systems that access stored software that programs or configures the computer system from a general-purpose computing device into a special-purpose computing device that implements one or more embodiments of the subject matter. Any suitable programming, scripting, or other type of language or combination of languages can be used to implement the teachings contained herein in software that will be used to program or configure the computing device.

[0130] Embodiments of the methods disclosed herein can be performed in the operation of such computing devices. The order of the blocks presented in the above examples can vary, e.g., the action blocks can be reordered, combined, and / or decomposed into sub-blocks. Specific blocks or processes can be performed in parallel.

[0131] Conditional language used herein, e.g., "wherein," "can," "could," "might," "may," "for example," etc., unless specifically stated otherwise or otherwise understood within the context in which it is used, is generally intended to mean that a particular example includes a particular feature, element, and / or step while other examples do not. Thus, such conditional language is generally not intended to imply that a feature, element, and / or step is required in any way for one or more examples or that one or more examples necessarily include logic for deciding whether such a feature, element, and / or step is included in any particular example or will be performed in any particular example with or without author input or prompting.

[0132] The terms "comprising," "including," "having," etc. are synonyms and are used inclusively in an open-ended manner and do not exclude additional elements, features, acts, operations, etc. Further, the term "or" is used in its inclusive sense (and not in its exclusive sense) such that when, for example, used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. The use of "adapted to" or "configured to" herein means open and inclusive language that does not exclude devices adapted to or configured to perform additional tasks or steps. Further, the use of "based on" is meant to be open and inclusive because a process, step, calculation, or other action "based on" one or more recited conditions or values can in fact be based on additional conditions or values beyond those recited. Similarly, the use of "at least partially based on" is meant to be open and inclusive because a process, step, calculation, or other action "at least partially based on" one or more recited conditions or values can in practice be based on additional conditions or values beyond those recited. The headings, lists, and numbers included herein are for ease of explanation only and are not restrictive.

[0133] The various features and processes described above can be used independently of each other or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Additionally, in some embodiments, specific method or process blocks may be omitted. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated therewith can be performed in a suitable other order. For example, the described blocks or states can be performed in an order different from that specifically disclosed or multiple blocks or states can be combined in a single block or state. Exemplary action blocks or states can be performed serially, in parallel, or in a specific other manner. Action blocks or states can be added to or removed from the disclosed examples. Similarly, the example systems and components described herein can be configured differently from those described. For example, elements can be added, removed, or rearranged compared to the disclosed examples.

Claims

1. A method implemented by a computer system, the method comprising: In an augmented reality (AR) scenario, a depth image is generated by a depth sensor of the computer system; The depth image is updated by at least dividing the depth image into a plurality of depth layers and moving pixels from a first depth layer of the plurality of depth layers to a second depth layer; The updated depth image is output to at least one AR application associated with the AR scenario, the at least one AR application including a first application pipeline for detecting occlusion and a second application pipeline for detecting collisions; The outputting the updated depth image to the second application pipeline for detecting collisions includes: In the AR scenario, an RGB image is generated by an RGB optical sensor of the computer system; Based on the updated depth image and the RGB image, an RGB-depth (RGBD) image is generated; Based on the updated depth image, a set of three-dimensional (3D) points is generated in the coordinate system of the AR scenario; A 3D model including a plurality of multi-level voxels is generated based on the depth image, wherein one multi-level voxel of the plurality of multi-level voxels is associated with one 3D point from the set; Determining a collision between a virtual object and the multi-level voxel in the AR scenario according to the 3D model; and In the AR scenario, rendering the virtual object based on the depth of the virtual object and the RGBD image and based on the collision.

2. The method according to claim 1, wherein, The updating the depth image by at least dividing the depth image into a plurality of depth layers and moving pixels from a first depth layer of the plurality of depth layers to a second depth layer includes: Dividing the depth image into a plurality of depth layers, each depth layer corresponding to a depth range and including pixels having depth values within the depth range; Selecting a first depth layer having a first layer number and a second depth layer having a second layer number from the plurality of depth layers; Adjusting the first depth layer based on the first layer number, a first pixel in the first depth layer, the second layer number, and a second pixel in the second depth layer, wherein the adjustment includes moving pixels from the first depth layer to the second depth layer; wherein the first layer number is greater than the second layer number, the depth of the first depth layer is deeper than the depth of the second depth layer, and adjusting the first depth layer includes performing morphological dilation from the first depth layer to the second depth layer; Updating the depth image based on the adjustment.

3. The method according to claim 2, wherein, The total number of the depth layers is based on the maximum depth of the depth sensor.

4. The method according to claim 2, wherein, The difference between the depth ranges of two consecutive depth layers is between 0.4 meters and 0.6 meters.

5. The method according to claim 2, wherein, The selecting a first depth layer having a first layer number and a second depth layer having a second layer number from the plurality of depth layers includes: Determining the first layer number corresponding to the first depth layer and the second layer number corresponding to the second depth layer based on the numbers of the plurality of depth layers, wherein the difference between the first layer number and the second layer number is equal to or greater than 2.

6. The method according to claim 5, wherein, The first depth layer and the second depth layer are also selected based on the total number of the first pixels and the total number of the second pixels both being equal to or greater than a predefined threshold.

7. The method according to claim 2, wherein, The size of the kernel for the morphological dilation is based on the difference between the first layer number and the second layer number.

8. The method according to claim 2, wherein, The morphological dilation is repeated iteratively for a number of iterations, and the number of iterations is based on the difference between the first layer number and the second layer number.

9. The method according to claim 1, the set of 3D points includes a point cloud, wherein, The multi-level voxel includes a first voxel with a first grid size at a first level and a second voxel with a second grid size at a second level, the second grid size being smaller than the first grid size, and wherein, generating the 3D model includes: Dividing the coordinates of the 3D points by the first grid size to generate indices of the 3D points; Hashing the indices to determine hash values; Determining that the hash map does not include the hash values; and Updating the hash map to include the hash values.

10. The method according to claim 1, wherein, Rendering the virtual object includes: preventing the collision from being rendered by at least controlling the movement of the virtual object, wherein, the multi-level voxel includes a first voxel with a first grid size at a first level and a second voxel with a second grid size at a second level, the second grid size being smaller than the first grid size, and determining a collision between the virtual object in the AR scenario and the multi-level voxel includes: Generating one or more bounding boxes around the virtual object; Determining a first intersection between the one or more bounding boxes and the first voxel; Based on the first intersection, determining that the first voxel has a first hash value in the hash map; Based on the first hash value being included in the hash map, determining a second intersection between the one or more bounding boxes and a second voxel among the second voxels; Based on the second intersection, determining that the second voxel has a second hash value in the hash map; and Detecting the collision based on the second hash value being included in the hash map.

11. The method according to claim 1, wherein, The multi-level voxel includes a first voxel with a first grid size at a first level and a second voxel with a second grid size at a second level, the second grid size being smaller than the first grid size, and determining a collision between the virtual object in the AR scenario and the multi-level voxel includes: Storing a sorted queue including bits associated with a second voxel among the second voxels, wherein each bit is associated with a different depth image and indicates whether the second voxel corresponds to a 3D point visible in the different depth image; Removing an end bit from the end of the sorted queue; Inserting a start bit at the beginning of the sorted queue, wherein the start bit is associated with the depth image; Determining that the total number of bits indicating the visibility of the second voxel in the sorted queue is greater than a predefined threshold; and Detecting the collision based on the second voxel.

12. The method according to claim 1, wherein, Outputting the updated depth image to a first application pipeline for detecting occlusion includes: In the AR scenario, generating an RGB image by the RGB optical sensor of the computer system; Generate a red-green-blue depth (RGBD) image based on the updated depth image and the RGB image; Determine that a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image and a second pixel of the virtual object; Detect occlusion of the virtual object in the AR scenario based on the depths of the first pixel and the second pixel; Render the virtual object in the AR scenario based on the occlusion; 13. The method according to claim 12, wherein, Generating the RGBD image includes: Registering the depth image with the RGB image based on the image resolution of the depth image, the image resolution of the RGB image, and a transformation between the depth sensor and the RGB optical sensor; Performing depth densification on the depth image, the depth densification including a plurality of morphological dilations of the depth image; After the depth densification, filtering the depth image based on median filtering; and Based on the registration, upsampling the filtered depth image to the image resolution of the RGB image, wherein pixels in the RGBD image correspond to pixels in the RGB image and pixels in the upsampled depth image; 14. The method according to claim 12, wherein, Generating the RGBD image includes: Generating an alpha map from the depth image; and Upsampling the depth image and the alpha map to the image resolution of the RGB image; 15. The method according to claim 11, wherein, Rendering the virtual object includes: Determine that a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image and a second pixel of the virtual object; Determine a first depth of the first pixel from the RGBD image; Determine that the first depth is less than or equal to a second depth of the second pixel; Generate a smoothing factor for the first pixel based on the alpha map; and Set the RGB value of the pixel in the AR image based on a first RGB value of the first pixel, a second RGB value of the second pixel, and the smoothing factor; 16. The method according to claim 15, wherein, The smoothing factor is set to α = 1 - m i / 255, and the RGB value is set to and "α" is the smoothing factor, "i" is the pixel, "m i " is the value determined for the pixel from the alpha map, is the RGB value, "ci" is the first RGB value, and is the second RGB value.

17. The method according to claim 12, wherein, Rendering the virtual object includes: Determine that a pixel to be rendered in the AR image corresponds to a first pixel of the RGBD image and a second pixel of the virtual object; Determine a first depth of the first pixel from the RGBD image; Determine that the first depth is greater than a second depth of the second pixel; and Set the RGB value of the pixel in the AR image to be equal to the RGB value of the second pixel; 18. A computer system, comprising: A depth sensor configured to generate a depth image in an augmented reality (AR) scenario; A red-green-blue (RGB) optical sensor configured to generate an RGB image in the AR scenario; One or more processors; And One or more memories storing computer-readable instructions that, when executed by the one or more processors, configure the computer system to perform the steps of the method recited in claims 1 to 17.

19. One or more non-transitory computer storage media storing instructions that, when executed on a computer system, cause the computer system to perform the steps of the method of claims 1 to 17.

Citation Information

Patent Citations

  • Optimizing head mounted displays for augmented reality

    US20170206712A1