Method and apparatus for combining augmented reality objects in a real-world image

By storing and blending depth data within the AR system, the method addresses the inaccuracies in depth data, improving the rendering of 3D objects in relation to real-world objects and enhancing the overall user experience.

JP7699611B2Active Publication Date: 2025-06-27GOOGLE LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2022571306
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-05-22
Publication Date
2025-06-27
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

Current augmented reality (AR) systems face inaccuracies and instabilities in depth data, leading to undesirable user experiences when rendering 3D objects in relation to real-world objects.

Method used

The method involves storing depth data associated with real-world objects and using this stored data to improve the accuracy of depth and position processing when rendering 3D objects within the AR system. This is achieved by blending stored depth images with new depth images and generating a real-world image that accurately represents the depth and position of objects.

Benefits of technology

This approach enhances the accuracy of depth processing, resulting in a more stable and desirable user experience by ensuring that 3D objects are rendered correctly in relation to real-world objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699611000003
    Figure 0007699611000003
  • Figure 0007699611000004
    Figure 0007699611000004
  • Figure 0007699611000005
    Figure 0007699611000005
Patent Text Reader

Abstract

A method includes receiving a first depth image associated with a first frame at a first time of an augmented reality (AR) application, the first depth image representing at least a first portion of a real world space, the method further includes storing the first depth image and receiving a second depth image associated with a second frame at a second time after the first time of the AR application, the second depth image representing at least a second portion of the real world space, the method further includes generating a real world image by blending at least the stored first depth image with the second depth image, receiving a rendered AR object, combining the AR object in the real world image, and displaying the real world image combined with the AR object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Field Embodiments relate to scene representation in an augmented reality system.

Background Art

[0002] Background Augmented reality (AR) may include fusing three-dimensional (3D) graphics with real-world geometry. When a 3D object moves around in real-world geometry, the 3D object can appear in front of or behind a real-world object when rendered on an AR display. For example, a humanoid object can appear in front of or behind furniture, a half-wall, a tree, etc. in real-world geometry when rendered on an AR display.

[0003] However, current AR systems can have inaccurate and / or unstable depth data such that, upon rendering on an AR display, some parts of a real-world object and / or 3D object may (e.g., at the correct depth and / or position) appear and / or may not appear. For example, when a humanoid emerges from behind a real-world object (e.g., a half-wall), a part of the humanoid (e.g., a leg) may not appear when that part of the humanoid should appear upon rendering on an AR display. Alternatively, when a humanoid moves to a position behind a real-world object (e.g., a half-wall), a part of the humanoid (e.g., a leg) may appear when that part of the humanoid should not appear upon rendering on an AR display. This can result in a less-than-desirable user experience in current AR systems.

Summary of the Invention

[0004] Summary In a general scenario, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or a method including one or more processors and a memory storing instructions may execute a process in a certain way, the method including receiving a first depth image associated with a first frame at a first time of an augmented reality (AR) application, the first depth image representing at least a first portion of the real-world space, the method further including storing the first depth image and receiving a second depth image associated with a second frame at a second time after the first time of the AR application, the second depth image representing at least a second portion of the real-world space, the method further including generating a real-world image by at least blending the stored first depth image with the second depth image, receiving a rendered AR object, combining the AR object within the real-world image, and displaying the real-world image combined with the AR object.

[0005] The embodiment may include one or more of the following features. For example, the first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application. The first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application, and the method may further include selecting a portion of the plurality of depth images stored in the buffer and generating a data structure based on the portion of the plurality of depth images, where the data structure represents the real-world space, the data structure includes depth information, position information, and orientation information, and the method may further include storing the generated data structure. The first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application, and the method may further include receiving a portion of the plurality of depth images stored in the buffer and generating a plurality of surface elements (surfels) based on the portion of the plurality of depth images, where the plurality of surfels represents the real-world space, and the method may further include storing the generated plurality of surfels.

[0006] For example, the method may further include receiving a data structure including depth information, position information, and orientation information, rendering the data structure as a third depth image, and blending the third depth image with the real-world image. The method may further include receiving a plurality of surfels representing the real-world space, rendering the plurality of surfels as a third depth image, and blending the third depth image with the real-world image. Combining AR objects within the real-world image may include, based on depth, replacing a portion of the pixels within the real-world image with a portion of the pixels within the AR object.

[0007] Blending the stored first depth image with the second depth image may include replacing a portion of the pixels in the second depth image with a portion of the stored first depth image. The second depth image may lack at least one pixel, and blending the stored first depth image with the second depth image may include replacing at least one pixel with a portion of the stored first depth image. The method further includes receiving a plurality of surfaces representing the real-world space and rendering the plurality of surfaces. The second depth image may lack at least one pixel, and the method further includes replacing at least one pixel with a portion of the rendered plurality of surfaces. The stored first depth image may include a position confidence indicating the likelihood that the first depth image represents the real-world space at a certain position.

[0008] In another general aspect, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or method including one or more processors and a memory storing instructions can execute a certain process in a certain way, the method including receiving depth data associated with a frame of an augmented reality (AR) application, the depth data representing at least a portion of the real-world space, the method further including storing the depth data in a buffer associated with the AR application as one of a plurality of depth images representing the frame of the AR application, selecting a portion of the plurality of depth images stored in the buffer, generating a data structure based on the portion of the plurality of depth images, the data structure representing the real-world space, the data structure including depth information, position information, and orientation information, and the method further including storing the generated data structure.

[0009] Embodiments may include one or more of the following features. For example, the data structure may include a plurality of surface elements (surfels). The data structure may be stored in association with a server. Selecting a portion of the plurality of depth images may include selecting a plurality of images from a plurality of buffers on a plurality of devices executing an AR application. The stored depth data may include a position confidence indicating the likelihood that the depth data represents the real-world space at a certain position.

[0010] In yet another general aspect, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that may be executed on a computer system), and / or method including one or more processors and a memory storing instructions may execute a certain process in a certain way, the method including receiving first depth data associated with a frame of an augmented reality (AR) application, the first depth data representing at least a portion of the real-world space, the method further including receiving a data structure representing at least a second portion of the real-world space associated with the AR application, the data structure including depth information, position information, and orientation information, the method further including generating a real-world image by blending at least the first depth data with the data structure, receiving an AR object, combining the real-world image with the AR object, and displaying the real-world image combined with the AR object.

[0011] Embodiments may include one or more of the following features. For example, combining AR objects within a real-world image may include replacing a portion of the pixels within the real-world image with a portion of the pixels within the AR object, based on depth. Blending stored first depth data with a data structure may include replacing a portion of the pixels within a second depth image with a portion of the stored first depth image. The first depth data may lack at least one pixel, and blending the first depth data with a data structure may include replacing at least one pixel with a portion of the data structure. The data structure may include a plurality of surface elements (surfels). The data structure may include a plurality of surfels, the first depth data may lack at least one pixel, and the method may further include replacing at least one pixel with a portion of the plurality of surfels. A data structure representing a real-world space may include a position confidence indicating the likelihood that the depth data represents the real-world space at a location. The data structure may be received from a server.

[0012] Brief Description of the Drawings Exemplary embodiments will be more fully understood from the following detailed description and the accompanying drawings, in which like elements are represented by like reference numerals, which are given by way of example only and thus are not limiting of the exemplary embodiments.

Brief Description of the Drawings

[0013]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0014] Note that these figures are intended to illustrate general characteristics of methods, structures, and / or materials utilized in certain exemplary embodiments and to supplement the following description. However, these drawings are not to scale and may not precisely reflect the exact structure or performance characteristics of any given embodiment, and should not be construed as defining or limiting the range of values or characteristics encompassed by the exemplary embodiments. For example, the relative thicknesses and positioning of molecules, layers, regions, and / or structural elements may be reduced or exaggerated for clarity. The use of like or the same reference numerals in the various drawings is intended to indicate the presence of like or the same elements or features.

[0015] Detailed Description At least one problem with current augmented reality (AR) systems is the potential for latency in the depth processing of three-dimensional (3D) objects as they move around in the real-world space (e.g., geometry). For example, inaccurate rendering (e.g., depth and / or position) of a portion of a 3D object can occur as the 3D object moves behind an object in the real-world geometry (e.g., a real-world object). An exemplary implementation solves this problem by storing depth data associated with real-world objects and using the stored depth data when fusing and rendering the 3D objects with the real-world geometry. At least one advantage of this technique is that object depth and / or position processing can be more accurate as the 3D object changes position in the real-world geometry. Higher depth processing accuracy can result in a more desirable user experience compared to current AR systems.

[0016] In an exemplary implementation, an input depth frame can be stored in a buffer. As an AR video (e.g., captured by a camera) progresses from frame to frame, previous frame data (stored in the buffer) is re-projected onto the current frame. In other words, the latest input values are merged with previous values. This data integration process can be based on a statistical analysis of errors in the input depth frame and / or the accumulated depth frame.

[0017] Figure 1A shows a real-world space 100 and shows a user 130 within the real-world space 100. Real-world objects and AR objects are shown together in this figure as they would be seen by user 130 via a mobile device. A scene (e.g., of a room) as seen by user 130 of the AR system is shown by dashed lines. The real-world space 100 can include at least one real-world object 135. An AR system associated with the mobile device can be configured to place an AR object 140 within the real-world space 100. Figure 1A shows that the AR object 140 is placed at a depth behind the real-world object 135. However, based on the depth and position of the real-world object 135 as compared to the depth and position of the AR object 140, only a portion 145 of the AR object is behind the real-world object 135.

[0018] Figure 1B again illustrates the AR object 140 within the real-world space 100. In Figure 1B, the AR object has been repositioned and is placed at a depth in front of the real-world object 135. Figure 1C again illustrates the AR object 140 within the real-world space 100. In Figure 1C, the AR object remains in a location at a depth in front of the real-world object 135. However, depth processing of the real-world object 135 (e.g., missing depth data and / or lack of accurate depth data) is delaying the rendering of the real-world space 100 (e.g., on the display of the AR system). This delay in depth processing or data acquisition causes a portion 150 of the real-world object to be shown as being in front of the AR object 140 when all of the AR object 140 should be rendered in front of the real-world object 135.

[0019] Figure 1C shows an undesirable result when rendering the real-world space 100. The result shown by Figure 1C can occur when the depth information associated with the real-world space 100 (and more specifically, regarding the real-world object 135) is incomplete. This result can occur if the depth information is incomplete after the AR object 140 has been moved and / or if, for example, the user 130 looks away from the real-world object 135 and then looks back at the real-world object 135. For example, Figure 1A may correspond to a first frame in the AR system, and Figures 1B and / or 1C may correspond to successive frames after the frame corresponding to Figure 1A.

[0020] According to an exemplary implementation, depth information associated with a frame corresponding to FIG. 1A can be stored in a memory (e.g., a buffer). The depth information can be stored as geometry generated based on the real-world space 100 and / or the rendering of the real-world space 100 (e.g., a plurality of pixels including depth) (as will be described in more detail below). Frames sequentially rendered after the frame corresponding to FIG. 1A can use the stored depth information. This can result in a rendering that uses the complete depth information (associated with the real-world object 135). Thus, frames after the AR object 140 is moved can include the complete depth information (associated with the real-world object 135) and can be rendered as shown in FIG. 1B. Exemplary implementations can prevent or minimize rendering as shown in FIG. 1C.

[0021] FIG. 1D is a block diagram of a signal flow for generating an augmented reality (AR) image according to an exemplary implementation. As shown in FIG. 1D, the signal flow includes a rendered depth image 105 block, a buffer 110 block, a rendered and stored image 115 block, a blend 120 block, and a display 125 block.

[0022] In an exemplary implementation, depth data associated with an image viewed by a user via a mobile device (e.g., as shown in FIG. 1A) may correspond to a rendered depth image 105 and may be stored in buffer 110. A frame subsequent to the current frame may be captured and displayed by the mobile device. Each of these frames may be stored in buffer 110. These stored frames may be used to generate more complete depth data representing the depth of objects within the real-world space 100. As new frames are rendered (e.g., as shown in FIG. 1B), the stored depth data may be used to supplement (or, as represented by blend block 120, blended with) the depth data associated with the current rendering of the real-world space 100. Thus, when an AR object (e.g., AR object 140) is repositioned (or a new AR object is positioned), the AR object may be rendered at the correct depth relative to the real-world object.

[0023] The rendered depth image 105 block can be a rendering of an image (or a frame of a video) captured by a camera of a device (e.g., a mobile phone, a tablet, a headset, etc.) that executes an AR application. The rendered depth image 105 (or associated depth data) can be stored in buffer 110. In an exemplary implementation, multiple rendered depth images can be stored or accumulated within buffer 110. The accumulated frames of the rendered depth image can represent that multiple rendered depth images 105 are stored in buffer 105. Alternatively, or in addition, the accumulated frame can represent that a blended depth image is stored within buffer 105 (as represented by the dashed line). As a result of the accumulation of the blended depth image, missing depth data or invalid depth data can be replaced over time by valid depth data. In other words, as the AR system captures images, the valid depth data from the captured images can be accumulated (and stored) over time. The multiple rendered depth images (or depth images) represent frames of the AR application.

[0024] The rendered depth image 105 block can include an image having depth information and / or color information. The depth information can include a depth map having depth values for each pixel in the image. The depth information can include depth layers each having a number (e.g., an index or z-index) indicating a layer order. The depth information can be a layered depth image (LDI) having multiple ordered depths for each pixel in the image. The color information can be the color (e.g., RGB, YUV, etc.) for each pixel in the image. The depth image can be an image where each pixel represents a distance from the camera position. In some cases, the input depth image can be a sparse image, and some (e.g., some or most) of the pixels can be blank or marked as invalid.

[0025] The rendered, stored depth image 115 block can be a rendering of an image retrieved from storage (e.g., memory) of a device (e.g., a mobile phone, tablet, headset, etc.) that executes an AR application and / or from a server having memory accessible using the device that executes the AR application. As shown, the rendered, stored depth image 115 block can be read from buffer 110. The rendered, stored depth image 115 block can include an image having depth information and / or color information. The depth information can include a depth map having depth values for each pixel in the image. The depth information can include depth layers each having a number (e.g., an index or z-index) indicating a layer order. The depth information can be a layered depth image (LDI) having multiple ordered depths for each pixel in the image. The color information can be a color (e.g., RGB, YUV, etc.) for each pixel in the image.

[0026] The blend 120 block is configured to blend the rendered depth image 105 with the rendered, stored depth image 115, and the display 125 block is configured to display the resulting blended image. In an exemplary implementation, the blended image (or associated depth data) can be stored in buffer 110. In an exemplary implementation, multiple blended images can be stored or accumulated within buffer 110.

[0027] Blending two or more depth images can include combining a portion of each image. For example, data missing (e.g., depth data, color data, pixels, etc.) from the rendered depth image 105 block can be filled using data from the rendered, stored depth image 115 block. For example, pixels having the same position and / or the same position and the same depth can be combined. The position can be based on a distance and direction from a reference point (or home position). The position can be based on a coordinate system (e.g., an x, y grid).

[0028] As described above, some of the pixels in the depth image may be blank or marked as invalid. Thus, in an exemplary implementation, pixels that are missing or marked as invalid in the rendered depth image 105 may be filled with pixels from the rendered, stored depth image 115 that have the same position and layer as the missing or invalid pixels. In one exemplary implementation, the pixels from the rendered depth image 105 and the rendered, stored depth image 115 are in the same position and in a layer with the same index value. Blending the two images may include selecting pixels from the rendered, stored depth image 115 and discarding pixels from the rendered depth image 105. Alternatively, blending the two images may include selecting pixels from the rendered depth image 105 and discarding pixels from the rendered, stored depth image 115. Alternatively, blending the two images may include averaging the colors and assigning the averaged color to the respective position and respective layer. Other techniques for blending two images are within the scope of the present disclosure.

[0029] FIG. 2 is a block diagram of a signal flow for storing geometry according to an exemplary implementation. As shown in FIG. 2, the signal flow includes a position 205 block, an image 210 block, a depth 215 block, a stored geometry 220 block, a geometry construction 225 block, and a position confidence 230 block. The image 210 block may be data associated with an image captured by a camera of the AR system. The depth 215 block may be depth data captured by a depth sensor associated with the AR system and / or calculated based on the image data.

[0030] Each image captured by an AR system (e.g., the camera of the AR system) may include color data (represented by the image 210 block) and depth data (represented by the depth 215 block). The camera that captures the image can be at a certain position within the real-world space (e.g., real-world space 100) and can be directed in a certain direction within the real-world space. That position and direction can be represented by the position 205 block. The position 205, the image 210, and the depth 215 can be used to generate geometry data associated with the real-world space (represented by the geometry construction 225 block). The generated geometry can be added to and stored in the previously generated geometry. Adding geometry can also include replacing and / or substituting data. The geometry can include data representing objects within the real-world space. How a particular data representing an object corresponds to an object at a certain position can be determined and stored (represented by the position confidence 230 block).

[0031] The position 205 block can be position information associated with the AR system. The position information can be associated with a reference point (e.g., the starting (reference or home) position in the real-world geometry) and / or a global (global position) reference point (e.g., from a global positioning sensor). The position 205 can be the distance and direction from the reference point (or home position). The position 205 can be based on a coordinate system (e.g., an x, y grid) and direction.

[0032] The geometry construction 225 block may be configured to generate a data structure (e.g., an n-tuple, a tree, etc.) representing real-world geometry associated with the AR system. The geometry construction 225 block may generate a data structure for a position. In some implementations, the position 205, the image 210, and the depth 215 may be read from the buffer 110. The geometry construction 225 block may use the image 210, the depth 215, and the position 205 to generate a data structure. That is, the data structure may include color information, texture information, depth information, position information, and orientation information. The depth information may include depth layers each having a number (e.g., an index or a z-index) indicating a layer order. The depth information may be a layered depth image (LDI) having multiple ordered depths for each pixel in the image. The depth information may include a depth map. The texture information, the depth information, the position information, and the orientation information may be elements of a geometric object having connectivity (e.g., a polygon mesh). The texture information, the depth information, the position information, and the orientation information may be elements of a geometric object without connectivity (e.g., a surface element or a surfel). Further, the depth data may be stored in a grid-based data structure such as a voxel or an octree that may store a signed distance function (SDF) in sampling in 3D space. The geometry construction 225 block may add the data structure to the stored geometry 220.

[0033] The geometry construction 225 block may also generate a confidence value based on a comparison between the newly generated data structure and a stored data structure (e.g., previously generated). The more similar the newly generated data structure is to the stored data structure, the higher the position confidence. In other words, for a certain position, if there is a close match between the two data structures, there is some higher probability (e.g., high confidence) that using the stored data structure to render an image will result in an accurate image for the current real-world scene. The position confidence may indicate the likelihood that data, a data structure, a part of a data structure, depth data, etc. represents the real-world space at a position.

[0034] A hierarchical depth image (LDI) can be an image-based representation of a three-dimensional (3D) scene. The LDI can include a two-dimensional (2D) array or group of hierarchical depth pixels. Each hierarchical depth pixel can include a set of LDI samples sorted along one line of sight when viewed from a single camera position or viewpoint. The camera can also be referred to as an LDI camera. Other ways to refer to LDI samples can include, but are not limited to, points, depth pixels, or hierarchical depth pixel samples. For each LDI sample, a camera called the source camera provides the data associated with the LDI sample. The representation of an LDI pixel can include color information, alpha channel information, depth information (the distance between the pixel and the camera), an identifier of the source camera of the LDI sample (e.g., a number, pointer, or reference to the camera), and other attributes that can support the rendering of the LDI in three-dimensional (3D) space. For example, the alpha channel information can be used to determine the opacity level of the pixel.

[0035] LDI samples within that partition plane can be projected onto points within the source camera window space. The source camera window space can include a plurality of pixels. The points can be projected onto at least one pixel included in the image plane of the source camera. Then, the points can be projected back from the source camera window space onto the partition plane. The projection can result in a surface element (surfel) at a certain position of the point within the partition plane. The size of the surfel can be determined by an image filter defined within the source camera window space, which can also be referred to as the source filter. The surfel can have an associated color based on the color of the pixel onto which the point was projected within the image plane of the source camera.

[0036] A target camera including a target camera window space can be selected. A surface in the partition plane can be projected onto a surface footprint within the target camera window space. The surface footprint can cover, overlap with, or include one or more pixels included in the image plane of the target camera window space. The one or more pixels can be filled with a color and / or depth associated with the surface. Each of the one or more pixels can include a plurality of pixel samples or points. Each of the plurality of pixel samples can be projected from the target camera window space onto the partition plane. Each of the plurality of pixel samples can be projected from the partition plane into the source camera window space and the current position can be identified for each pixel sample within the source camera window space. Based on the identified current position of each pixel sample in the source camera window space, a color weight can be applied to each pixel sample. The partition plane and the texture map can be combined to form a model of a scene for 3D real-time rendering in the AR space.

[0037] Figure 3 is a block diagram showing the rasterization of a surface element (surfel) included in a partition plane. An application executed in a computing system can generate an image of a scene from various positions within the field of view of a device (e.g., a mobile phone, a tablet, a headset, etc.) executing an AR application. Each scene image can include a plurality of pixel samples or points including values for associated color information, depth information, and surface normals. A point can be a position in 3D space without volume, size, and extent. A point can represent the position of a pixel as seen from the center of the source camera through the center of the pixel. The number of pixels within the view of an image is determined based on the resolution of the image.

[0038] For example, when projected onto the surface representation, a pixel can be considered a surfel. A surfel can be used to efficiently render complex geometric objects in 3D space in real time (at an interactive frame rate). A surfel can include one or more samples (points) contained in the raw LDI. A surfel can be a point primitive lacking any specific connectivity. Thus, since there is no need to calculate topology information such as adjacency information, a surfel can be used to model dynamic geometry. The attributes of a surfel can include, but are not limited to, depth, texture color, as well as a normalized vector and position.

[0039] Each scene image can be assembled into a data structure (e.g., an LDI) that can be used in a simplified version of the representation of the scene for 3D real-time rendering (drawing) in the AR space by an AR application. For example, multiple pixel samples or points can be grouped into multiple partitions. A partition can be a plane or polygon that includes a subset of the multiple pixel samples or points representing the scene image. A partition plane (e.g., partition plane 304) can be at a position in the 3D image space where a subset of the points is located in 3D space. In some implementations, a quadrilateralization algorithm can create a polygon approximation that can be used to create the partition plane. In some implementations, an iterative partitioning algorithm can create a polygon approximation that can be used to create the partition plane.

[0040] A texture map can be created (generated) for each partition plane. Each partition plane and its associated texture map can be combined to form a model (simplified representation) of the scene for 3D real-time rendering (drawing) in the AR space by an AR application. The algorithm executed by the AR application when rendering the model of the scene can be based on the algorithm used to create each of the partition planes.

[0041] Referring to FIG. 3A, the points included in the partition plane are rasterized to create (generate) a texture map for the partition plane. Point 302 can be one of many points included in partition plane 304. The color of each point included in the partition plane is incorporated into the texture map. The texture rasterization algorithm can create (generate) a texture map for a given partition plane that includes an RGBA texture.

[0042] For example, referring to FIGS. 3A-3E, the texture rasterizer can select a source camera that includes a source camera window space 306 that can be the image plane of the source camera. The texture rasterizer can select a target camera that includes a target camera window space 336 that can be the image plane of the target camera. The texture rasterizer can select a source camera and a target camera for a given partition plane, and the partition plane includes samples (points) from the raw LDI. In some implementations, the selection of the target camera is based on the camera that has the best view of the partition plane. In some implementations, the source camera is the camera that provided (generated) the samples included in the raw LDI. In some implementations, the target camera can be a virtual camera.

[0043] Each point included in the partition plane can be projected onto the source camera window space. Referring to FIG. 3A, the source camera window space 306 may include a plurality of pixels (e.g., pixels 360a - 360l). When projected, a point (e.g., point 302) is projected onto at least one pixel (e.g., pixel 360f) included in the image plane of the source camera (source camera window space 306), resulting in the projected point 308. A filter 310 of a specific size and shape (e.g., a circle having a specific radius as shown in the examples of FIGS. 3A - 3B) is included in the source camera window space 306. The filter 310 positions the projected point 308 at the center of the filter 310 (the filter 310 is arranged around the projected point 308). The filter 310 may define a shape that can completely or partially overlap the pixels included in the source camera window space 306. In the example shown in FIG. 3A, the filter 310 completely or partially overlaps the shaded pixels (e.g., pixels 360b - c, pixels 360e - f, and pixels 360h - l) including the pixel 360f that includes the projected point 308. The shape of the filter 310 may also define the size of the point 302 projected within the source camera window space 306.

[0044] Referring to FIGS. 3A - 3B, when the projected point 308 is projected back into the partition plane 304, it results in a surface 322 at the position of the point 302 in the partition plane 304. For example, a plurality of light rays (e.g., light rays 310a - 310d) can be drawn from the corners of the pixels included in the filter 310 (where the filter 310 overlaps) to the partition plane 304. The intersection of the light rays and the partition plane 304 can define a surface footprint 320 for the surface 322. Additionally, the surface can be depicted as a mathematical function of a circle. The surface can be approximated by polygons and rasterized. The surface can be simplified to a forward - facing rectangle and depicted as having a uniform depth. A circular filter (e.g., filter 310) can define the size of the point 302 in the image plane of the source camera (e.g., source camera window space 306). The filter 310 (e.g., a circle having a specific radius as shown in the example of FIGS. 3A - 3B) can be projected back onto the partition plane 304 along with the projected point 308. When the projected point 308 is projected back into the partition plane 304, for the example shown in FIGS. 3A - 3C, it results in a quadrilateral surface footprint (e.g., surface footprint 320) that defines the point 302 in the partition plane 304. Additionally, the surface footprint 320 provides (defines) the size of the point 302 in the partition plane 304. The surface footprint 320 is a 3D shape with respect to the surface 322.

[0045] Pixels 360f within the source camera window space 306, which is defined by the filter 310 and contains the projected point 308, are projected back into the partition plane 304, resulting in the surfels 322. The texture rasterizer performs projecting the point 302 from the partition plane 304 to a pixel (e.g., pixel 360f) within the source camera window space 306 as defined by the filter, and then projecting that pixel (e.g., pixel 360f) back into the partition plane 304. As a result, the point 302 is changed to a surfel 322 with an associated color based on the color of that pixel. The size associated with the filter 310 may determine the size of the surfel footprint. In some implementations, the size of the surfel footprint may be approximately the same for each surfel footprint. Additionally, or alternatively, the position of the source camera relative to the partition plane may also contribute to determining the size of the surfel footprint.

[0046] Pixels projected into the partition plane from a first source camera result in a surfel footprint that is larger than the surfel footprint resulting from pixels projected into the partition plane from a second source camera when the position of the first source camera is closer to the partition plane than the position of the second source camera. For each surfel, the best source camera can be selected. Thus, each surfel may be associated with a different source camera.

[0047] As described, the partition plane 304 may include a plurality of pixel samples or a subset of points representing the scene image. Projecting the plurality of points included in the partition plane 304 into the source camera window space 306 and then projecting them back into the partition plane 304, the partition plane 304 may include a plurality of surfels having various surfel footprints.

[0048] Each surface included in the partition plane has an associated color. The color associated with a surface can be the color of the projected pixel from the source camera. For example, the color associated with surface 322 can be the color of pixel 360f. Creating a texture map for the surfaces included in the partition plane provides the colors necessary for rendering (drawing) the partition plane in a 3D real-time scene in the AR space.

[0049] Generally, the partition plane can be input into a texture rasterizer. The texture rasterizer can generate and output a texture map and a matrix for the partition plane. The output texture map can include an RGBA texture. The matrix can transform the coordinates of a point from the world or camera space, using the view matrix, into the eye space (view space). The eye space enables each coordinate to be viewed from the perspective of the camera or observer. The partition plane can include the plane and vectors of the surfaces in the LDI eye space.

[0050] The texture rasterizer can define a target camera. In some implementations, the target camera can be the same as the source camera. In some cases, the target camera can be a different camera from the source camera. Referring to FIG. 3C, the surface included in the partition plane (e.g., surface 322 included in partition plane 304) can be projected into the target camera window space (e.g., target camera window space 336). The projection of surface 322 includes projecting point 302 to pixel 340e (projected point 318) and projecting the surface footprint 320 (projected surface footprint 330).

[0051] The texture rasterizer may define a texture map (image of the texture) as pixels within the target camera window space 336. Projecting the partition plane 304 into the target camera window space 336 results in a surface footprint 330 that includes pixel samples or points 312a - 312e for the surface 322. Since the points 312a - 312e included in the surface footprint 330 define the texture for the partition plane 304, the texture rasterizer may use the points 312a - 312e to determine the color values of the surface.

[0052] The target camera window space may be an image plane. The image plane may include pixels 340a - 340l. The image plane of the target camera has an associated resolution based on the number of pixels included. For example, as shown in FIGS. 3A - 3C, the resolution of the image plane of the target camera (target camera window space 336) is the same as the resolution of the image plane of the source camera (source camera window space 306). In some implementations, the resolution of the image plane of the target camera may be different from the resolution of the image plane of the source camera.

[0053] The target camera window space 336 may include a plurality of pixels 340a - 340l. The projected surface footprint 330 may include (cover) pixels 340a - 340i included in the target camera window space 336. Referring to FIG. 3C, the pixels 340a - 340i included in (covered by) the surface footprint 330 are shown shaded. The pixels 340a - 340i covered by (overlapped by) the projected surface footprint 330 may be filled (colored) with the color associated with the surface 322. The pixels 340a - 340i may define a texture (texture map) for the surface 322.

[0054] The projected surface footprint 330 can be filled with the color of the projected surface 332. In some implementations, one or more pixels that are partially included in the projected surface footprint 330 (where the projected surface footprint 330 overlaps) can be filled with the color associated with the projected surface 332. For example, pixels 340a - 340d and pixels 340f - 340i are partially covered by the projected surface footprint 330.

[0055] Nine pixels (e.g., pixels 340a - 340i) are shown as being included in the projected surface footprint 330 (where the projected surface footprint 330 overlaps). In some implementations, fewer than nine pixels may be included in the projected surface footprint 330 (the projected surface footprint 330 may overlap fewer than nine pixels). In some implementations, more than nine pixels may be included in the projected surface footprint 330 (the projected surface footprint 330 may overlap more than nine pixels). For example, the number of pixels that may be included in the projected surface footprint 330 (the projected surface footprint 330 may overlap) can be on the order of 1 - 9 pixels. However, the projected surface footprint 330 can be the same size as the entire image.

[0056] Referring to FIGS. 3C-3D, a pixel can include a plurality of pixel samples or points. For purposes of explanation, FIGS. 3C-3D show points 312a-312i that are pixel samples or points for each respective pixel 340a-340i covered by the projected surface footprint 330. Each pixel 340a-340l can include a plurality of pixel samples. Points (e.g., points 312a-312i) covered by the projected surface footprint 330 can be projected back into the partition plane 304, resulting in a projected surface footprint 340 that includes projected points (e.g., projected points 342a-342i, respectively). The projection includes projecting point 318, resulting in a point 344 that is projected into the projected surface 332.

[0057] The projected surface footprint 340 shown in FIG. 3D is the same size, the same shape as the surface footprint 320 shown in FIG. 3C, and is in the same position within the same 3D space. The color associated with the color of the projected surface 332 is the color associated with the surface 322.

[0058] Referring to FIGS. 3C-3E, each pixel sample or point included within the projected surface footprint included within the partition plane can be projected back into the source camera window space. The projection can identify the current position for a pixel sample or point within the source camera window space. For example, points 342a-342i can be projected back into the source camera window space 306 as projected points 362a-362i. For example, the projected point 362a identifies the position within the source camera window space 306 for the projected point 342a, which is the projection of the point 312a from the target camera window space 336. For example, the point 312a can be generated as a point included in the pixel 340a within the target camera window space 336.

[0059] Filter 310 may have an associated function that gives a weight to each pixel sample or point included in (overlapped by) filter 310 with respect to the color of the point. The color weight of the point may be based on the distance of the point from the point located at the center of filter 310 (e.g., projected point 354). When point 354 and surfel 332 are projected back into source camera window space 306, they may result in surfel 352 centered at point 354.

[0060] Referring to FIGS. 3A-3E, partition plane 304 may include a plurality of surfels and associated surfel footprints. Each surfel and its associated footprint may be projected back into source camera window space 306 and filter 310 may be applied. In some cases, one surfel footprint may overlap an adjacent surfel footprint. By making the function of filter 310 a bell curve, a smooth blend between adjacent surfels is ensured. The blend of the colors of adjacent surfels is improved.

[0061] The partition plane may include a plurality of surfels (e.g., two or more surfels). In some implementations, each surfel may be associated with a different source camera. In some implementations, each surfel may be associated with the same source camera. In some implementations, some surfels may be associated with a first source camera and other surfels may be associated with a second source camera. For example, the best source camera may be selected for a particular surfel.

[0062] The color for a point or pixel sample included in a surfel footprint (e.g., surfel footprint 340) projected back into the source camera window space (e.g., source camera window space 306) and included in a filter (e.g., filter 310) can be determined (calculated) based on the color value for the surfel included at the center of the filter (e.g., the color of surfel 332 which is the color of surfel 352). An exemplary calculation for determining the color of a pixel sample or point (color (p1)) is shown by Equation 1,

[0063] [Number]

[0064] where p1 is a pixel sample or point, the surfel color value is the color value of the surfel included in the same filter as pixel sample or point p1, and the weight value (p1) is the weight value for pixel sample or point p1,

[0065] [Number]

[0066] is the sum of all weights, and n = total number of weights. Surfel-based fusion for a depth map can take as input a sequence of depth images. A depth image can be an image where pixels represent the distance from the camera position. This input image can be a sparse image. For example, some or most of the input pixels can be blank and / or marked as invalid. Each depth pixel can have an associated confidence value (e.g., in the range of 0 to 1). Further, the depth image can have a corresponding luminance image representing the same scene and camera elements (e.g., number of pixels, image length and width, etc.). In an exemplary implementation, the accumulated depth frames (e.g., stored in buffer 110) can be used as input to 3D surfel generation.

[0067] In some implementations, the techniques described herein may be implemented in an application programming interface (API). Inclusion in the API may provide access to additional data for these techniques. For example, the estimated pose of each camera frame with respect to world coordinates may be determined based on the tracking provided by the API. The tracking may also be used in position determination, ambient light estimation, background and flat surface determination, etc. In other words, access to the API data may help to execute some of the techniques described herein.

[0068] In some implementations, a portion of the depth frames (e.g., every other, one out of three, one out of four, less than all, etc.). In other words, a lower frame rate may be used for 3D surfel generation (e.g., as compared to the frames captured as depth 215 and / or image 210 during an AR session). In some implementations, the frames for 3D surfel generation may be stored in a second (not shown) buffer and / or a portion of buffer 110.). In some implementations, the input frames may be stored in a second buffer and used for input to surfel generation at a lower frame rate. As a result, data from multiple frames may be used, but the actual processing of the multiple frames may be performed at a lower frame rate.

[0069] Figures 4A, 4B, 4C, and 4D graphically illustrate the accumulation of geometry according to an exemplary implementation. The geometry can be stored as a data structure (as discussed above). For example, the data structure can be an n-tuple. (such as a surface). In an exemplary implementation, a surface can be generated from the input depth by first estimating the normal orientation associated with each pixel in the input. Given the depth value and normal vector associated with each pixel, the surface can be generated by clustering these pixels and generating a disk (as described above) represented in world coordinates. The size of these disks can be based on a number of adjacent pixels that share the same depth and orientation (and optionally color). These surfaces can be stored over the frames of an AR session, and when each new depth frame is integrated, the surfaces can be updated and / or merged based on this new information. If the new depth information does not match the previous surface data (e.g., if something moves within the scene), the original surface can be penalized (e.g., a decrease in confidence), deleted, and / or replaced with new geometry.

[0070] As shown in FIG. 4A, grid 405 can represent a portion of the real-world geometry (or partition plane) that can be displayed on the AR display. There is a first object 410 on grid 405. Circle 415 is a graphical representation of a data structure (e.g., a surface footprint) that can contain information about a portion of grid 405 that includes the first object 410. In an exemplary implementation, the data structure (or data structures, such as a surface) can be stored as stored geometry 220. Additionally, the data structure can include a confidence, which can be stored as confidence 230. ) As shown in FIG. 4B, the number of circles 415 is increasing, and a second object 420 is added to the grid 405. The circle 425 is a graphical representation of a data structure (e.g., a surface footprint) that may contain information regarding a portion of the grid 405 that includes the second object 420. FIG. 4B may represent that the grid 405 represents a portion of the real world captured at a different time (e.g., a later time) than FIG. 4A.

[0071] As shown in FIG. 4C, the number of circles 415 is increasing, and the second object 420 is moving on the grid 405. The circle 425 is in two parts shown as circles 425-1 and 425-2 respectively associated with the second object 420 at different positions on the grid 405. Since the circle 425-1 remains at the position where 420 was not in FIG. 4B, the circle 425-1 may be penalized (e.g., a decrease in confidence). FIG. 4C may represent that the grid 405 represents a portion of the real world captured at a different time (e.g., a later time) than FIG. 4B.

[0072] As shown in FIG. 4D, the number of circles 415 is increasing, the circle 425-1 has been removed, and the number of circles 425-2 is increasing (the circle 425-2 may represent the object 420). In FIG. 4D, the object 410 and the object 420 may be fully represented (e.g., as a data structure or a surface). FIG. 4D may represent that the grid 405 represents a portion of the real world captured at a different time (e.g., a later time) than FIG. 4C.

[0073] FIG. 5 is a block diagram of a signal flow for generating an image for display according to an exemplary implementation. As shown in FIG. 5, the signal flow includes a rendered image 505 block, a stored geometry 220 block, a rendering 510 block, a blend 515 block, an AR object 520 block, a post-processing 525 block, and a display 530 block. In an exemplary implementation, stored geometry (e.g., representing the real-world space captured and stored over time) can be used to generate more complete depth data representing the depth of an object within the real-world space. When a new frame is rendered, the stored geometry can be used to supplement the depth data associated with the current rendering of the real-world space (or blended with it as represented by the blend 515 block). The AR object can be combined with the blended rendering to generate an image for display on the display of the AR system.

[0074] The rendered image 505 block can be a rendered depth image (e.g., based on depth 215) and / or a rendered color image (e.g., based on image 210). In other words, the rendered image 505 can include color information, depth information, orientation information, layer information, object information, etc. In some implementations, the depth information can be blank (e.g., incomplete pixels or missing pixels) and / or marked as invalid.

[0075] The rendering 510 block may be a rendering of an image retrieved from the stored geometry 220 block. The stored geometry 220 block may represent the storage (e.g., memory) of a device (such as a mobile phone, tablet, headset, etc.) that executes an AR application and / or a server having memory accessible using the device that executes the AR application. As shown, the rendering 510 block may be read from the stored geometry 220 block. The rendering 510 block may include an image having depth information and / or color information. The depth information may include a depth map having depth values for each pixel in the image. The depth information may include depth layers each having a number (e.g., index or z-index) indicating a layer order. The depth information may be a layered depth image (LDI) having multiple ordered depths for each pixel in the image. The color information may be the color (e.g., RGB, YUV, etc.) for each pixel in the image.

[0076] In an exemplary implementation, the stored geometry 220 block may include a data structure that includes a surface. Accordingly, the rendering 510 block may use a projection technique. The projection technique may include point-based rendering or splatting. Point-based rendering or splatting may include assigning variables to pixels within the pixel space. The variables may include color, texture, depth, direction, etc. The variables may be read from at least one surface (e.g., based on the surface and pixel positions).

[0077] The blend block 515 is configured to blend the rendered image 505 with the rendering 510 of the stored geometry 220. By blending the rendered image 505 with the rendering 510 of the stored geometry 220, a representation of the real world or a real-world image can be generated. Blending two or more images can include combining portions of each image. For example, data missing from the rendered image 505 block (e.g., depth data, color data, pixels, etc.) can be filled using data from the rendering 510 of the stored geometry 220. For example, pixels having the same position and / or the same position and the same depth can be combined. The position can be based on the distance and direction from a reference point (or home position). The position can be based on a coordinate system (e.g., an x, y grid).

[0078] As described above, some of the pixels in the depth image can be blank or marked as invalid. Thus, in an exemplary implementation, pixels missing or marked as invalid (e.g., missing depth information or invalid depth information) in the rendered image 505 (e.g., having depth) can be filled with pixels (e.g., depth information) from the rendering 510 of the stored geometry 220 having the same position and layer as the missing or invalid pixels. In one exemplary implementation, pixels from the rendered image 505 and the rendering 510 of the stored geometry 220 are in layers having the same position and having the same index value. Blending two images can include selecting pixels from the rendering 510 of the stored geometry 220 and discarding pixels from the rendered image 105.

[0079] Alternatively, blending two images may include selecting pixels from the rendered image 105 and discarding pixels from the rendering 510 of the stored geometry 220. Alternatively, blending two images may include averaging colors and assigning the averaged colors to the respective positions and respective layers. Other techniques for blending two images are within the scope of the present disclosure.

[0080] Further, prior to blending, the images may be projected onto each other. For example, an image may be captured while the mobile device is in motion. The previous frame (e.g., stored in a buffer and / or generated by the rendering (510) block) may be re-projected onto the current frame (e.g., the rendered image 505). This implementation may enable (or help enable) alignment of objects and / or observed features across frames.

[0081] The blend block 515 can also combine the AR object 520 with the real-world image. The AR object can be an image generated by an AR application for placement into the real-world space (by the AR application). As described above, blending and / or combining two or more images can include combining portions of each image. In an exemplary implementation, combining the AR object 520 with the real-world image can include depth-based occlusion. For example, if a portion of the AR object 520 is at a depth (e.g., layer) in front of a portion of a real-world object, the portion of the real-world object can be removed from the combined image. Further, if a portion of the AR object 520 is at a depth (e.g., layer) behind a portion of a real-world object, the portion of the AR object 520 can be removed from the combined image. The advantage of using the stored geometry 220 is that if the real-world image includes depth information marked as blank or invalid at the location of a portion of the AR object 520, the depth information of the stored geometry 220 can be used instead of the blank or invalid depth information. Thus, depth-based occlusion can be more accurate using the stored geometry 220.

[0082] The post-processing 525 block can improve the quality of the resulting image or frame. The display 530 block is configured to display the resulting post-processed blended and / or combined image. For example, the resulting image or frame can be filtered to smooth or sharpen transitions between colors. The resulting image or frame can be filtered to remove artifacts (e.g., errors including colors or depths that are likely not part of the image). The resulting image or frame can be filtered to remove AR and real-world discontinuities (such as AR elements that should be blocked by real-world elements).

[0083] In an exemplary implementation, the stored geometry 220 can be used without the rendered image 505. In other words, the real-world space can be the stored real-world space. Thus, the stored geometry 220 can be a complete (or highly refined) representation of the real-world space. In this implementation, the AR object 520 is combined with the rendered, stored real world.

[0084] In some implementations, the techniques described herein can be implemented in an application programming interface (API). The API can be an element of a developer toolkit. Including these techniques in an API accessible to developers can enable many use cases. For example, the real-world space can represent a living space (e.g., a living room, a dining room, etc.). The AR object 520 can be a furniture object (e.g., a sofa, a chair, a table, etc.). A user of the AR application can place the furniture object in the living space as needed. In a further implementation, the AR application can be configured to remove an object in the real-world space. For example, an existing piece of furniture can be removed and replaced with an AR image of another piece of furniture. As described above, the stored geometry 220 can include surfaces. In this embodiment, a portion of the surface can be removed (e.g., deleted from the stored geometry 220 or prevented from being rendered by the rendering 510 block).

[0085] FIGS. 6 and 7 are flowcharts of methods according to exemplary embodiments. The methods described with respect to FIGS. 6 and 7 may be performed due to the execution of software code stored in a memory (e.g., a non-transitory computer-readable storage medium) associated with the apparatus and executed by at least one processor associated with the apparatus.

[0086] However, alternative embodiments such as a system embodied as a dedicated processor are contemplated. The dedicated processor can be a graphics processing unit (GPU). The GPU can be a component of a graphics card. The graphics card can also include video memory, a random access memory digital-to-analog converter (RAMDAC), and driver software. The video memory can be a frame buffer that stores digital data representing an image, a video frame, an object in an image, or a scene in a frame. The RAMDAC can be configured to read the contents of the video memory, convert the contents into an analog RGB signal, and transmit the analog signal to a display or monitor. The driver software can be software code stored in the aforementioned memory. The software code can be configured to implement the methods described herein.

[0087] The methods described below are described as being executed by a processor and / or a dedicated processor, but the methods are not necessarily executed by the same processor. In other words, at least one processor and / or at least one dedicated processor may execute the methods described below with respect to FIGS. 6 and 7.

[0088] FIG. 6 is a diagram showing a method for storing geometry according to an exemplary implementation. As shown in step S605, color data is received. For example, an AR application can operate. The AR application can operate on a computing device (e.g., a mobile phone, a tablet, a headset, etc.) including a camera. The color data can be captured by the camera and communicated. Step S605 starts the first stage of a two-stage technique.

[0089] In step S610, depth data is received. For example, the camera may include a function of capturing depth data. In other words, the camera may capture color data (e.g., RGB) and depth (D) data. The camera may be an RGBD camera. Alternatively, or in addition, the AR application may be configured to generate depth data from color data. As described above, the depth data may include pixels that are blank (e.g., incomplete or missing pixels) and / or pixels marked as invalid.

[0090] In step S615, a position is received. For example, the AR application may be configured to determine the position of a computing device within the real-world space. The position may be based on the distance and direction from a reference point (or home position). The position may be based on a coordinate system (e.g., an x, y grid). The position may also include depth (e.g., the distance from an object). The AR application may be configured to generate a reference during the initialization of the AR application. The reference may be a position in the real-world space and / or a global (global position) reference point (e.g., from a global positioning sensor).

[0091] In step S620, image data is stored. For example, image data based on color data and depth data is stored. The image data may be pixels, point (or point cloud) data, polygon (e.g., triangle) data, mesh data, a surface, etc. As described above, the depth data may include pixels that are blank (e.g., incomplete or missing pixels) and / or pixels marked as invalid. Therefore, the image data may have missing or invalid depth information. As described above, the image data may be stored in a buffer (e.g., buffer 110). In this first stage, this image data may be reprojected (e.g., rendered), blended with current data (e.g., as captured by the camera), and combined with AR objects for display by the AR application.

[0092] In step S625, the stored image data is read out. For example, the image data stored in advance can be read out. The stored image data can be read out from a buffer (for example, buffer 110). In an exemplary implementation, the stored image data includes a plurality of frames captured by the camera of the device that executes the AR application. Further, a part of the plurality of frames can be read out (for example, every other frame, one out of three frames, one out of four frames, less than all, etc.). Step S625 starts the second stage of the two-stage technique.

[0093] In step S630, the current geometry is constructed. For example, a data structure (for example, an n-tuple, a tree, etc.) representing the real-world geometry (for example, the geometry of the real-world space) associated with the AR application or system can be constructed. The data structure can include a certain geometry for a certain position (for example, for a certain object at a certain position). The data structure can include color information, texture information, depth information, position information, and orientation information. The depth information can include depth layers each having a number (for example, an index or a z-index) indicating the layer order. The depth information can be a layered depth image (LDI) having a plurality of ordered depths for each pixel in the image. The depth information can include a depth map. The texture information, depth information, position information, and orientation information can be elements of a geometric object having connectivity (for example, a polygonal mesh). The texture information, depth information, position information, and orientation information can be elements of a geometric object without connectivity (for example, a surface element or a surfel).

[0094] In step S635, the stored geometric data is updated. For example, the geometry can be added to an existing data structure (e.g., the stored geometry 220). Updating the stored data can include adding to the stored geometric data, modifying the stored geometric data, replacing a part of the stored geometric data, and / or deleting a part of the stored geometric data. As described above, the geometric data can be a surface. Thus, updating the stored geometric data can include adding a surface, modifying a surface, replacing a surface, and / or deleting a surface. Over time (e.g., as the data structure becomes a complete representation of the real-world space), the stored geometry can be used to generate an image (e.g., a frame) within the AR application instead of the stored image data. In other words, the stored geometric data (or a part thereof) can be rendered, blended with the current data (e.g., captured by a camera), and combined with AR objects for display by the AR application. In some implementations, the stored geometric data (or a part thereof) can be rendered and combined with AR objects for display by the AR application (without the current data).

[0095] In step S640, the position reliability is updated. For example, the reliability can be a numerical value within a range (e.g., 0 to 1). A larger (or smaller) numerical value can indicate a higher reliability, and a smaller (or larger) numerical value can indicate a lower reliability. For example, a reliability of 1 can be a high reliability, and a reliability of 0 can be a low reliability. The reliability can indicate how well the data structure (or a part of the data structure) represents the real-world space.

[0096] For example, if an object in the real-world space repeatedly appears at a certain position and depth, the data structure representing that object in the real-world space may have a high associated confidence level (e.g., a value approaching 1). In an exemplary implementation, a surface representing an object may have a high associated confidence level (e.g., a value approaching 1). If an object in the real-world space appears at a first position and depth within a first frame, and at a second position and depth within a second frame, the data structure representing that object in the real-world space may have a low associated confidence level (e.g., a value approaching 0). In an exemplary implementation, a surface representing an object may have a low associated confidence level (e.g., a value approaching 0). As described above, the data structure representing an object that has moved within the real-world space may ultimately be deleted at that position and depth within the data structure. As described above, the real-world space may be represented by a plurality of data structures (e.g., surfaces).

[0097] These surfaces have several use cases, including object scanning, room reconstruction, physical collision, free space detection, path planning, etc. Further, one of our core use cases is to feedback the surface data to the final output depth map.

[0098] FIG. 7 is a diagram showing a method for generating an image according to an exemplary implementation. As shown in step S705, a rendered image is received. For example, an AR application may operate. The AR application may operate on a computing device (e.g., a mobile phone, a tablet, a headset, etc.) including a camera. An image (or a frame of a video) may be received from the camera. The received image (or a frame of a video) may be rendered.

[0099] In step S710, an augmented reality (AR) object is received. For example, the AR object can be an object generated by an AR application for placement into the real-world space (by the AR application). Thus, the AR object can be received from elements of the AR application configured to generate the AR object. The AR object can include color information, depth information, orientation information, position information, and the like.

[0100] In step S715, stored geometric data is received. For example, the geometric data can be stored on a computing device. For example, the geometric data can be stored on a server. For example, the geometric data can be stored on a cloud (or remote) memory device. Thus, the geometric data can be received from a computing device, a server, and / or a cloud memory. The geometric data can be received via wired or wireless communication.

[0101] In step S720, the stored geometric data is rendered. For example, the geometric data is rendered as an image. Rendering the geometric data can generate at least a portion of an image representing the real-world space. Rendering can use projection techniques. The projection techniques can include point-based rendering or splatting. Point-based rendering or splatting can include assigning variables to pixels within the pixel space. The variables can include color, texture, depth, orientation, and the like. The variables can be read from at least one surface (e.g., based on the position of the surface and the pixel).

[0102] In step S725, the rendered image and the rendered geometric data are blended and combined with the AR object. For example, the rendered image and the rendered geometric data can be blended. Then, the AR object can be combined with the resulting image. By blending the rendered image with the rendering of the stored geometry, a representation of the real world or a real-world image can be generated. Blending two or more images can include combining portions of each image (described in more detail above).

[0103] As described above, some of the pixels in the depth image can be blank or marked as invalid. Thus, in an exemplary implementation, pixels that are missing or marked as invalid in the rendered image (e.g., missing depth information or invalid depth information) can be filled with pixels (e.g., depth information) from the rendering of the stored geometric data having the same position and layer as the missing or invalid pixels.

[0104] In an exemplary implementation, combining the AR object with the real-world image can include occlusion based on depth. For example, if a portion of the AR object is at a depth (e.g., layer) in front of a portion of the real-world object, the portion of the real-world object can be removed from the combined image. Further, if a portion of the AR object is at a depth (e.g., layer) behind a portion of the real-world object, the portion of the AR object can be removed from the combined image. The advantage of using the stored geometric data is that if the real-world image includes depth information that is blank or marked as invalid at the location of a portion of the AR object, the depth information of the stored geometric data can be used instead of the blank or invalid depth information. Thus, occlusion based on depth can be more accurate using the stored geometric data.

[0105] In step S730, post - blend processing is executed. For example, the post - blend processing can improve the quality of the resulting image or frame. The resulting image or frame can be filtered to smooth color transitions or sharpen color transitions. The resulting image or frame can be filtered to remove artifacts (e.g., errors including colors or depths that are likely not to belong to the image). The resulting image or frame can be filtered to remove AR and real - world discontinuities (such as AR elements that should be blocked by real - world elements). In step S735, the image is displayed. For example, the resulting post - processed blended image can be displayed.

[0106] FIG. 8 is a block diagram of a system for storing multi - device geometry according to an exemplary implementation. As shown in FIG. 8, the system includes device 1 805, device 2 810, device n 815, server 820, and memory 825. Device 1 805 includes a position 205 block, an image 210 block, and a depth 215 block. Device 2 810 includes a position 205 block, an image 210 block, and a depth 215 block. Device n 825 includes a position 205 block, an image 210 block, and a depth 215 block. Server 820 includes a geometry construction 225 block. Memory 825 includes a stored geometry 220 block and a position reliability 230 block.

[0107] In the exemplary implementation of FIG. 8, a plurality of devices (e.g., device 1 805, device 2 810, …, device n) operate together in an AR environment. The server receives image, depth, and position data from each of the plurality of devices. The geometry construction 225 block may generate geometric data using the image, depth, and position data from each of the plurality of devices. The server 820 may store the geometric data and reliability in the stored geometry 220 block and position reliability 230 block of the memory 825. Although the memory 825 is shown separately from the server 820, the memory 825 may be included in the server 820.

[0108] The plurality of devices includes a buffer 110 and may communicate a portion of the frames stored in the buffer 110 to the server 820. The server 820 may communicate the stored geometry (e.g., as a surface) to the plurality of devices. Thus, each of the plurality of devices may utilize a more complete real-world space in the AR application as compared to the real-world space generated by an individual device. In other words, each of the plurality of devices may utilize portions of the real-world space that may not be seen by an individual device. In the embodiment of FIG. 8, each of the plurality of devices is configured to project (e.g., render) and blend images.

[0109] FIG. 9 illustrates a block diagram of a system for a multi-device augmented reality system according to an exemplary implementation. As shown in FIG. 9, the system includes device 1 805, device 2 810, device n 815, server 820, and memory 825. Device 1 805 includes a position 205 block, an image 210 block, and a display 920 block. Device 2 810 includes a position 205 block, an image 210 block, and a display 920 block. Device n 815 includes a position 205 block, an image 210 block, and a display 920 block. The memory 825 includes a stored geometry 220 block and a position reliability 230 block.

[0110] In the embodiment of FIG. 9, the plurality of devices are configured to receive a projected (e.g., rendered) and blended image from server 920. The rendered and blended image represents the real-world space (e.g., real-world space 100). Accordingly, server 820 includes a post-blend 915 block, an image blender 910 block, and an image renderer 905 block. However, each of the plurality of devices is configured to combine an AR object with the rendering of the real-world space.

[0111] FIG. 10 shows an example of a computing device 1000 and a mobile computing device 1050 that may be used with the techniques described herein. Computing device 1000 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. Computing device 1050 is intended to represent various forms of mobile devices, such as a personal digital assistant, cellular phone, smartphone, and other similar computing devices. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit embodiments of the invention described in this document and / or claimed.

[0112] The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 connected to the memory 1004 and a high-speed expansion port 1010, and a low-speed interface 1012 connected to a low-speed bus 1014 and the storage device 1006. Each of the components 1002, 1004, 1006, 1008, 1010, and 1012 is interconnected using various buses and may be mounted appropriately on a common motherboard or in other manners. The processor 1002 is capable of processing instructions executed within the computing device 1000, and these instructions include instructions stored in the memory 1004 or on the storage device 1006 for displaying graphic information for a GUI on an external input / output device such as a display 1016 coupled to the high-speed interface 1008. In other implementation examples, multiple processors and / or multiple buses may be used as appropriate along with multiple memories and multiple types of memories. Also, multiple computing devices 1000 may be connected, and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0113] The memory 1004 stores information within the computing device 1000. In one implementation example, the memory 1004 is one or more volatile memory units. In another implementation example, the memory 1004 is one or more non-volatile memory units. The memory 1004 may also be another form of computer-readable medium such as a magnetic disk or an optical disk.

[0114] The memory device 1006 can provide mass storage for the computing device 1000. In one implementation example, the memory device 1006 may be a computer-readable medium, such as a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a device in a storage area network or other configuration, or may include such a computer-readable medium. A computer program product can be tangibly embodied in an information carrier. The computer program product may also include instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer-readable medium or a machine-readable medium, such as the memory 1004, the memory device 1006, or the memory on the processor 1002.

[0115] The high-speed controller 1008 manages bandwidth-intensive operations for the computing device 1000, while the low-speed controller 1012 manages less bandwidth-intensive operations. Such an assignment of functions is merely illustrative. In one implementation example, the high-speed controller 1008 is coupled to the memory 1004, to the display 1016 (e.g., via a graphics processor or accelerator), and to a high-speed expansion port 1010 that can receive various expansion cards (not shown). In this implementation example, the low-speed controller 1012 is coupled to the memory device 1006 and a low-speed expansion port 1014. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth (registered trademark), Ethernet (registered trademark), wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, etc., or to a networking device such as a switch or a router, e.g., via a network adapter.

[0116] Computing device 1000 can be implemented in many different forms as shown in the figures. For example, it can be implemented as a standard server 1020, or multiple times as a group of such servers. It can also be implemented as part of a rack server system 1024. Additionally, it can be implemented in a personal computer such as a laptop computer 1022. Alternatively, components from computing device 1000 can be combined with other components in a mobile device (not shown) such as device 1050. Each of such devices may include one or more of computing devices 1000, 1050, and the entire system may be composed of multiple computing devices 1000, 1050 that communicate with each other.

[0117] Computing device 1050 includes, among other components, a processor 1052, a memory 1064, an input / output device such as a display 1054, a communication interface 1066, and a transceiver 1068. Device 1050 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of components 1050, 1052, 1064, 1054, 1066, and 1068 are interconnected using various buses, and some of these components may be mounted appropriately on a common motherboard or in other manners.

[0118] Processor 1052 is capable of executing instructions within computing device 1050, including instructions stored in memory 1064. The processor can be implemented as a chipset of chips including multiple separate analog and digital processors. The processor can provide coordination among other components of device 1050, such as, for example, a user interface, applications executed by device 1050, and control of wireless communication by device 1050.

[0119] Processor 1052 can communicate with a user via a control interface 1058 and a display interface 1056 coupled to a display 1054. The display 1054 can be, for example, a TFT LCD (Thin-Film-Transistor Liquid Crystal Display), or an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 1056 can include appropriate circuitry for driving the display 1054 to present graphical and other information to the user. The control interface 1058 can receive commands from the user and convert them for sending to the processor 1052. Additionally, an external interface 1062 can be provided in communication with the processor 1052 to enable proximity area communication between the device 1050 and other devices. The external interface 1062 can provide, for example, wired communication in one implementation, or wireless communication in other implementations, and multiple interfaces may also be used.

[0120] Memory 1064 stores information within computing device 1050. Memory 1064 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Extended memory 1074 is also provided and can be connected to device 1050 via expansion interface 1072. Expansion interface 1072 can include, for example, a SIMM (Single In Line Memory Module) card interface. Such extended memory 1074 may provide additional storage space for device 1050, or may also store applications or other information for device 1050. Specifically, extended memory 1074 may include instructions for executing or supplementing the above-described process, and may also include secure information. For this reason, for example, extended memory 1074 may be provided as a security module for device 1050 and may be programmed with instructions that permit secure use of device 1050. Additionally, secure applications may be provided via the SIMM card along with additional information, such as by placing identification information on the SIMM card in a hack-proof manner.

[0121] The memory may include, for example, flash memory and / or NVRAM memory as described below. In one implementation example, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 1064, extended memory 1074, or memory on processor 1052, and can be received, for example, through transceiver 1068 or external interface 1062.

[0122] Device 1050 may communicate wirelessly via a communication interface 1066 that may include a digital signal processing circuit as needed. The communication interface 1066 may provide communication under various modes or protocols, such as GSM (registered trademark) voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA (registered trademark), CDMA2000, or GPRS, among others. Such communication may occur, for example, via a radio frequency transceiver 1068. In addition, short-range communication may occur using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 1070 may provide additional navigation-related and location-related wireless data to device 1050, and such data may be appropriately used by applications running on device 1050.

[0123] Device 1050 may also perform voice communication using a voice codec 1060 that can receive verbal information from a user and convert it into usable digital information. The voice codec 1060 may also generate sounds audible to the user, such as via a speaker, for example, in the handset of device 1050. Such sounds may include sounds from a voice call, recorded sounds (such as voice messages, music files, etc.), or sounds generated by an application operating on device 1050.

[0124] The computing device 1050 may be realized in many different forms as shown in the figure. For example, it may be realized as a mobile phone 1080. It may also be realized as part of a smartphone 1082, a personal digital assistant, or other similar mobile devices.

[0125] In a general scenario, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or a method including one or more processors and a memory storing instructions may execute a certain process in a certain way, the method including receiving a first depth image associated with a first frame at a first time of an augmented reality (AR) application, the first depth image representing at least a first portion of the real-world space, the method further including storing the first depth image and receiving a second depth image associated with a second frame at a second time after the first time of the AR application, the second depth image representing at least a second portion of the real-world space, the method further including generating a real-world image by at least blending the stored first depth image with the second depth image, receiving a rendered AR object, combining the AR object within the real-world image, and displaying the real-world image combined with the AR object.

[0126] The embodiment may include one or more of the following features. For example, the first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application. The first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application. The method may further include selecting a portion of the plurality of depth images stored in the buffer and generating a data structure based on the portion of the plurality of depth images. The data structure represents the real-world space and includes depth information, position information, and orientation information. The method may further include storing the generated data structure. The first depth image may be one of a plurality of depth images representing frames of an AR application stored in a buffer associated with the AR application. The method may further include receiving a portion of the plurality of depth images stored in the buffer and generating a plurality of surface elements (surfels) based on the portion of the plurality of depth images. The plurality of surfels represents the real-world space. The method may further include storing the generated plurality of surfels.

[0127] For example, the method may further include receiving a data structure including depth information, position information, and orientation information, rendering the data structure as a third depth image, and blending the third depth image with the real-world image. The method may further include receiving a plurality of surfels representing the real-world space, rendering the plurality of surfels as a third depth image, and blending the third depth image with the real-world image. Combining an AR object within the real-world image may include replacing a portion of the pixels within the real-world image with a portion of the pixels within the AR object based on depth.

[0128] Blending the stored first depth image with the second depth image may include replacing a portion of the pixels in the second depth image with a portion of the stored first depth image. The second depth image may be missing at least one pixel, and blending the stored first depth image with the second depth image may include replacing at least one pixel with a portion of the stored first depth image. The method further includes receiving a plurality of surfaces representing the real-world space and rendering the plurality of surfaces. The second depth image may be missing at least one pixel, and the method further includes replacing at least one pixel with a portion of the rendered plurality of surfaces. The stored first depth image may include a position confidence indicating the likelihood that the first depth image represents the real-world space at a certain position.

[0129] In another general aspect, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or method including one or more processors and a memory storing instructions can execute a certain process in a certain way, the method including receiving depth data associated with a frame of an augmented reality (AR) application, the depth data representing at least a portion of the real-world space, the method further including storing the depth data in a buffer associated with the AR application as one of a plurality of depth images representing the frame of the AR application, selecting a portion of the plurality of depth images stored in the buffer, generating a data structure based on the portion of the plurality of depth images, the data structure representing the real-world space, the data structure including depth information, position information, and orientation information, and the method may further include storing the generated data structure.

[0130] Embodiments may include one or more of the following features. For example, the data structure may include a plurality of surface elements (surfels). The data structure may be stored in association with a server. Selecting a portion of a plurality of depth images may include selecting a plurality of images from a plurality of buffers on a plurality of devices executing an AR application. The stored depth data may include a position confidence indicating the likelihood that the depth data represents the real-world space at a certain position.

[0131] In yet another general aspect, an apparatus, device, system, non-transitory computer-readable medium (storing computer-executable program code that may be executed on a computer system), and / or method including one or more processors and a memory storing instructions may execute a certain process in a certain way, the method including receiving first depth data associated with a frame of an augmented reality (AR) application, the first depth data representing at least a portion of the real-world space, the method further including receiving a data structure representing at least a second portion of the real-world space associated with the AR application, the data structure including depth information, position information, and orientation information, the method further including generating a real-world image by blending at least the first depth data with the data structure, receiving an AR object, combining the real-world image with the AR object, and displaying the real-world image combined with the AR object.

[0132] Embodiments may include one or more of the following features. For example, combining an AR object within a real-world image may include replacing a portion of the pixels in the real-world image with a portion of the pixels within the AR object based on depth. Blending the stored first depth data with a data structure may include replacing a portion of the pixels in the second depth image with a portion of the stored first depth image. The first depth data may lack at least one pixel, and blending the first depth data with a data structure may include replacing at least one pixel with a portion of the data structure. The data structure may include a plurality of surface elements (surfels). The data structure may include a plurality of surfels, the first depth data may lack at least one pixel, and the method may further include replacing at least one pixel with a portion of the plurality of surfels. The data structure representing the real-world space may include a position confidence indicating the likelihood that the depth data represents the real-world space at a certain position. The data structure may be received from a server.

[0133] Exemplary embodiments may include various modifications and alternative forms, and those embodiments are shown by way of example in the drawings and described in detail above. However, there is no intention to limit the exemplary embodiments to the specific forms disclosed, and on the contrary, it should be understood that the exemplary embodiments cover all modifications, equivalents, and alternatives falling within the scope of the claims. The same numbers refer to the same elements throughout the description of the figures.

[0134] The various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive and transmit data and instructions from and to a storage system, at least one input device, and at least one output device. The various implementations of the systems and techniques described herein may be realized as a circuit, module, block, or system that can combine software aspects with hardware aspects, and / or, generally, may be so referred to herein. For example, a module may include functions / acts / computer program instructions that execute on a processor (e.g., a processor formed on a silicon substrate, a GaAs substrate, etc.) or some other programmable data processing device.

[0135] Some of the above-exemplified embodiments are described as a process or method shown as a flowchart. Although these flowcharts describe the operations as a sequential process, many of the operations may be performed in parallel, concurrently, or simultaneously. Additionally, the order of the operations may be rearranged. The process may end when those operations are completed, but may also have additional steps not included in the figure. These processes may correspond to methods, functions, procedures, subroutines, subprograms, etc.

[0136] Some of the above-described methods, which are illustrated by flowcharts, may be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium such as a storage medium or a computer-readable medium. The processor may perform the necessary tasks.

[0137] The specific structural details and functional details disclosed herein are merely representative for explaining exemplary embodiments. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to only the embodiments described herein.

[0138] Terms such as first, second, etc. may be used herein to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0139] It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, no intervening elements are present. Other phrases used to describe the relationship between elements should be interpreted in a similar manner (e.g., "between" and "directly between", "adjacent" and "directly adjacent", etc.).

[0140] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of exemplary embodiments. As used herein, the singular forms are also intended to include the plural forms unless the context clearly indicates otherwise. The terms "comprises" and / or "comprising," when used herein, specify the presence of the recited features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0141] Also, in some alternative realizations, the recited functions / acts may occur in a different order than shown in the figures. For example, two figures shown in succession may in fact be executed simultaneously, depending on the functionality / acts involved, or may sometimes be executed in the reverse order.

[0142] Unless defined otherwise, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the exemplary embodiments belong. Further, terms such as those defined in commonly used dictionaries are to be interpreted as having a meaning that coincides with their meaning in the context of the relevant art and are not to be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0143] Regarding software, or algorithms and symbolic representations of operations on data bits in a computer memory, the above exemplary embodiments and corresponding detailed description portions are presented. These descriptions and representations are for those skilled in the art to effectively convey the content of their research to other skilled persons. An algorithm, as the term is used herein and generally, is considered to be a consistent sequence of steps leading to a desired result. These steps require physical operations on physical quantities. Although not necessarily, usually these quantities take the form of optical, electrical, or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, or numbers, etc. has proven to be sometimes convenient mainly for reasons of common usage.

[0144] In the above exemplary embodiments, references to symbolic representations of acts and operations that may be implemented as program modules or functional processes (e.g., in the form of a flowchart) include routines, programs, objects, components, data structures, etc. that perform a particular task or implement a particular abstract data type and can be described and / or implemented using existing hardware with existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application specific integrated circuits, or field programmable gate array (FPGA) computers, etc.

[0145] However, it should be borne in mind that all of these and similar terms should be associated with appropriate physical quantities and are nothing more than convenient labels applied to these quantities. Unless specifically stated otherwise or as apparent from the description, terms such as processing, computing, calculating, or determining of a display refer to actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical electronic quantities within the registers and memories of the computer system and transform that data into other data similarly represented as physical quantities within the computer system memory, registers, or other such information storage, transmission, or display devices.

[0146] Also, aspects implemented by software of the exemplary embodiments are typically encoded on some form of non-transitory program storage medium or are implemented over some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or hard drive) or optical (e.g., a compact disc read-only memory, i.e., CD ROM), and may be read-only or random access. Similarly, the transmission medium may be a twisted pair wire, coaxial cable, optical fiber, or some other suitable transmission medium known in the art. The exemplary embodiments are not limited by these aspects of a given implementation.

[0147] Finally, while the appended claims recite particular combinations of the features described herein, the scope of the present disclosure is not limited to the specific combinations claimed, but rather extends to any combination of the features or embodiments disclosed herein, whether or not that specific combination is presently specifically recited in the appended claims.

Claims

1. A method comprising: receiving a first depth image associated with a first frame captured by a camera at a first time of an augmented reality (AR) application, the first depth image representing at least a first portion of a real-world space, the method further comprising: receiving a second depth image associated with a second frame captured by the camera at a second time after the first time of the AR application, the second depth image representing at least a second portion of the real-world space, the method further comprising: identifying missing depth information or invalid depth information in the second depth image, and generating a blended depth image by replacing the missing depth information or invalid depth information with depth information corresponding to the missing depth information or invalid depth information from the first depth image; displaying an AR object combined with a real-world image generated based on the blended depth image.

2. The method of claim 1, wherein the first depth image is one of a plurality of depth images representing frames of the AR application stored in a buffer associated with the AR application.

3. The first depth image is one of a plurality of depth images representing frames of the AR application stored in a buffer associated with the AR application, and the method further comprises: selecting a portion of the plurality of depth images stored in the buffer; and generating a data structure based on the portion of the plurality of depth images, the data structure representing the real-world space and including depth information, position information, and orientation information, the method further comprising: storing the generated data structure. The method of claim 1.

4. The first depth image is one of a plurality of depth images representing frames of the AR application stored in a buffer associated with the AR application, and the method further comprises: receiving a portion of the plurality of depth images stored in the buffer; and generating a plurality of surfels based on the portion of the plurality of depth images, the plurality of surfels representing the real-world space, the method further comprising: The method according to claim 1, comprising storing the plurality of generated surfels.

5. Furthermore, receiving a data structure including depth information, position information, and orientation information; rendering the data structure as a third depth image; and blending the third depth image with the real-world image, the method according to any one of claims 1 to 4.

6. Furthermore, receiving a plurality of surfels representing the real-world space; rendering the plurality of surfels as a third depth image; and blending the third depth image with the real-world image, the method according to any one of claims 1 to 4.

7. Combining the AR object in the real-world image includes replacing a part of the pixels in the real-world image with a part of the pixels in the AR object based on depth, the method according to any one of claims 1 to 6.

8. Generating the real-world image includes replacing a part of the pixels in the second depth image with a part of the first depth image, the method according to any one of claims 1 to 7.

9. The missing depth information or invalid depth information corresponds to at least one pixel, and generating the real-world image includes replacing the at least one pixel with a part of the first depth image, the method according to any one of claims 1 to 7.

10. Furthermore, receiving a plurality of surfels representing the real-world space; rendering the plurality of surfels; and the second depth image is missing at least one pixel, and generating the real-world image includes replacing the at least one pixel with a part of the rendered plurality of surfels, the method according to any one of claims 1 to 9.

11. The first depth image includes a position confidence indicating the likelihood that the first depth image represents the real-world space at a certain position, the method according to any one of claims 1 to 10.

12. A method comprising receiving depth data associated with a frame of an augmented reality (AR) application, the depth data representing at least a part of the real-world space, and the method further storing the depth data in a buffer associated with the AR application as one of a plurality of depth images representing frames of the AR application, wherein the frame is captured by a single camera, and the method further comprises identifying missing depth information or invalid depth information in the depth data; selecting a portion of the plurality of depth images corresponding to the missing depth information or invalid depth information stored in the buffer; generating a data structure by replacing the missing depth information or invalid depth information with the portion of the plurality of depth images, the data structure representing the real-world space, the data structure including depth information, position information, and orientation information, and the method further comprises storing the generated data structure. A method **Claim 13** The method according to claim 12, wherein the data structure includes a plurality of surface elements (surfels). **Claim 14** The method according to claim 12 or 13, wherein the data structure is stored in association with a server. **Claim 15** The method according to any one of claims 12 to 14, wherein selecting the portion of the plurality of depth images includes selecting the plurality of depth images from a plurality of buffers on a plurality of devices executing the AR application. **Claim 16** The method according to any one of claims 12 to 15, wherein the stored depth data includes a position reliability indicating a likelihood of representing the real-world space at a position where the depth data is located. **Claim 17** A method comprising: receiving first depth data associated with a first frame of an augmented reality (AR) application, the first depth data representing at least a portion of a real-world space, the first frame being captured by a camera, and the method further comprises receiving a data structure representing at least a second portion of the real-world space associated with the AR application, the data structure being associated with a second frame captured by the camera, the data structure including depth information, position information, and orientation information, and the method further comprises identifying missing depth information or invalid depth information in the first depth data, and Generating a blended depth image by replacing the missing depth information or invalid depth information with depth information corresponding to the missing depth information or invalid depth information from the data structure; A method comprising displaying an AR object combined with a real-world image generated based on the blended depth image.

18. The method according to claim 17, wherein combining the AR object in the real-world image includes replacing a part of the pixels in the real-world image with a part of the pixels in the AR object based on depth.

19. The method according to claim 17 or 18, wherein generating the real-world image includes replacing a part of the pixels in the first depth data with a part of the data structure.

20. The first depth data is missing at least one pixel, The method according to any one of claims 17 to 19, wherein generating the real-world image includes replacing the at least one pixel with a part of the data structure.

21. The method according to any one of claims 17 to 20, wherein the data structure includes a plurality of surface elements (surfels).

22. The data structure includes a plurality of surfels, The first depth data is missing at least one pixel, and the method further includes replacing the at least one pixel with a part of the plurality of surfels. The method according to any one of claims 17 to 21.

23. The method according to any one of claims 17 to 22, wherein the data structure representing the real-world space includes a position reliability indicating the likelihood that the first depth data represents the real-world space at a certain position.

24. The method according to any one of claims 17 to 23, wherein the data structure is received from a server.

25. A program for causing a computer to execute the method according to any one of claims 1 to 24.

26. An apparatus comprising one or more processors and a memory storing the program according to claim 25, which is executed by the one or more processors.

Citation Information

Patent Citations

  • Position attitude measurement device, position attitude measurement method, and program

    JP2012026895A

  • Image processor, image processing method and program

    JP2014106543A

  • Information processing device and information processing method

    JP2015125621A

  • Real-time 3d reconstruction using power-efficient depth sensors

    JP2016514384A

  • Information processor, information processing method, program, and system

    JP2019125345A