How to process depth maps

By identifying depth regions and transmitting depth value ranges in multi-view imaging, the method addresses spatial and temporal coherence issues, reducing computational complexity and bit rate, and enhances the efficiency of depth map processing for real-time applications.

JP2025526575APending Publication Date: 2025-08-15KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504172
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-26
Filing Date
2023-08-21
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing multi-view imaging methods face challenges with spatial and temporal coherence, computational demands, and limited view transmission due to irregular patch shapes and the lack of alpha channel support in video standards, particularly in telepresence applications with real-time, low latency, and low computational resources.

Method used

The method involves identifying depth regions in a depth map and transmitting depth value ranges, which allows for efficient segmentation and rendering by comparing depth values with the transmitted ranges, reducing computational complexity and bit rate requirements.

Benefits of technology

This approach provides accurate and efficient rendering of multi-view frames with reduced computational power and bit rate, enabling more views to be transmitted without the need for a segmentation mask, thus improving the processing of depth maps for subsequent transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526575000001_ABST
    Figure 2025526575000001_ABST
Patent Text Reader

Abstract

A method is provided for processing a depth map for subsequent transmission. Depth regions in the depth map are identified, each defining a portion of the depth map. Depth value ranges are determined for the depth regions and are further transmitted together with the depth map. The depth value ranges provide segmentation information for the depth regions within the depth map.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the processing of depth maps, and in particular to the processing of depth maps and the use of depth maps for subsequent transmission. [Background technology]

[0002] Multi-view imaging generally refers to the imaging of a scene along with obtaining the scene's geometry. Typically, the scene geometry is captured in a depth map or the like. The scene depth is used along with the image of the scene to synthesize novel images at new viewpoints. In other words, multi-view imaging can be used to view a scene from selected viewpoints instead of relying on the camera's viewpoint.

[0003] Multi-view imaging can involve segmenting an image with depth into separate patches. The image patches and depth patches are packed into an atlas. On the client side, the patches are first sorted by decreasing z-axis value relative to the virtual viewpoint (i.e., the viewpoint being synthesized). The view synthesis algorithm then accesses the patches in this order and, when the patches have similar depths, alternates between blending patches from different source views and compositing the blended view onto the previously synthesized output. Summary of the Invention [Problem to be solved by the invention]

[0004] However, placing a variable number of patches with irregular shapes in the atlas per frame destroys spatial and temporal coherence (thereby increasing the bit rate) and increases the number of pixels required to store segmentation-related pixel data. Atlases typically consist of a background sprite texture with a depth model and a variable number of rectangles, where transparency is used to encode the irregular patch shapes within the rectangles, or a separate segmentation map is stored in the atlas.

[0005] Furthermore, the actual placement of patches for texture, depth, and segmentation is puzzle-like, which can be computationally demanding due to bitrate and pixel area constraints. While this may not be much of an issue for professional broadcast applications with sufficient computational resources (e.g., using cloud computing), it is likely to be problematic for telepresence applications with real-time, low latency, and low computational requirements.

[0006] Finally, most existing video standards do not support an alpha channel by default. The alpha channel provides a transparency value for a pixel. As a result, segment shapes consume costly pixel space. If full-view textures and full-view images are placed directly on the image frame, the total number of views that can be transmitted is limited. For example, encoders typically support 4K at 60Hz or 8K at 30Hz.

[0007] Therefore, there is a need to improve the segmentation methods used in multi-view images.

[0008] US 10855965 B1 discloses a segmented 3D multi-view image generator that generates fewer multi-view images for partitions with less salient features.

[0009] US2019 / 158838(A1) discloses efficient Wedgelet-based coding for coding blocks of various sizes. [Means for solving the problem]

[0010] The invention is defined by the claims.

[0011] According to an example embodiment of the present invention, there is provided a method for processing a depth map for subsequent transmission, the method comprising: obtaining a depth map of a scene comprising a plurality of depth values; identifying one or more depth regions in the depth map, each of the one or more depth regions defining a portion of the depth map; The method includes determining one or more ranges of depth values for one or more depth regions, and transmitting the depth map and the one or more ranges of depth values.

[0012] Improved multiview frames can be rendered by using depth maps with segmented objects. However, the segmentation is typically done at the encoder before transmitting / broadcasting the multiview information (e.g., texture and depth), and therefore the multiview information must include the segmentation information in order for the decoder to utilize the segmentation during rendering.

[0013] It has been found that providing a depth range for each object essentially provides depth segmentation for the depth region with much lower bitrate requirements than, for example, providing a segmentation mask. This is because there is no need to transmit a segmentation mask (e.g., a segmentation map) to provide the segmentation information. The range of depth values allows segmentation to be achieved efficiently at the decoder by comparing the depth values in the depth map with the transmitted range of depth values. When a depth value falls within one of the ranges, it can be assumed that the depth value corresponds to the object from which the range of depth values was derived.

[0014] Furthermore, this provides multi-view information that can be coded at relatively low computational cost.

[0015] The identified depth region is a contiguous portion of the depth map and may contain a portion of the entire object.

[0016] A depth region may include a region of the depth map that has a relatively smooth depth gradient. A depth region preferably consists of an object or part thereof.

[0017] A range of depth values can be transmitted as two values indicating the maximum and minimum depth values. Of course, other low-bitrate methods of transmitting a range of depth values are known. In particular, transmitting a range of depth values for a depth region (rather than specifying which pixels have which depth within the range) provides a low-data method for enabling segmentation functions. Each range of depth values can correspond to a single depth region, or there can be the same range shared between different depth regions. An indication of which depth region corresponds to which depth value range can also be transmitted.

[0018] The method may further comprise defining a plurality of cells in the depth map, each cell corresponding to a portion of the depth map, and identifying one or more depth regions in the depth map includes identifying one or more depth regions within each of the cells.

[0019] Additionally, defining cells in a depth map that correspond to only a portion of the depth map allows different portions of the depth map to be coded separately. This can reduce the processing power required to render multiview frames using depth maps by rendering each frame based on a corresponding number of ranges of depth values within each cell. For example, cells containing only background depth information (i.e., a single range for the entire cell) can be accurately rendered without using back-to-front compositing, while cells containing two or more ranges of depth regions are preferably rendered using back-to-front compositing.

[0020] Preferably, the depth map is divided into more than 15 cells.

[0021] Defining a plurality of cells in the depth map may include dividing the depth map into a grid of cells.

[0022] A grid of cells includes a number of cells, each of which is of a known size. For example, the grid can include a number of equally shaped cells.

[0023] Knowing the size of each cell in the grid allows for more efficient coding and rendering of multiview frames using depth maps, since the position of each cell is known and consistent between frames.

[0024] The method may further include transmitting a number of different ranges of depth values in each cell.

[0025] Depth interval data should be compressed and decompressed with low computational complexity (and latency), but also should not take up too much space. In a binary data format, this can be achieved by indicating where a depth range ends in one cell and begins in another. To implement this, for each cell we have a 4-bit number that indicates how many intervals there are in the cell (up to 16 intervals per cell).

[0026] Only 16 or fewer distinct ranges of depth values can be transmitted.

[0027] This allows all range indices to be coded with 4 bits. The range of depth values can be coded using, for example, 16 bits each.

[0028] It would be realistic to assume that a limited number of depth ranges appear within a cell (e.g., up to 4, 10, 12, or 16). Using this information helps to reduce the bit rate with limited encoding / decoding complexity, thus adding relatively little latency.

[0029] Depth ranges can vary between cells. For example, there may be objects close together on one side of the depth map and objects farther away on the other side. Therefore, it may be advantageous to have a specific depth range for each cell. Having many cells per depth map and / or many depth intervals per cell may also impose a large overhead on depth ranges.

[0030] Identifying the one or more depth regions may be based on determining a depth smoothness between depth values of the depth map.

[0031] For example, first and / or second derivatives can be used to determine how quickly depth changes as a function of image coordinates. Depth changes are allowed within a segment only if the local changes are smooth. As the change in depth from one pixel to the next increases, the likelihood that this change will result in cover / occlusion or uncover / deocclusion during rendering increases. Essentially, depth smoothness is used to predict from a depth map in sensor space (i.e., as it was acquired) that the corresponding real-world 3D object will have a continuous surface.

[0032] The method may further include determining a mesh resolution based on the depth smoothness and transmitting the mesh resolution.

[0033] The present invention also provides a method for synthesizing an image at a target viewpoint, the method comprising: receiving at least one or more images and one or more depth maps of a scene; receiving one or more ranges of depth values for at least one of the depth maps, each range of depth values corresponding to one or more depth regions in one of the depth maps; and processing at least one of the depth maps together with at least one of the images using the range of depth values to synthesize an image at the target viewpoint.

[0034] By processing the depth map based on a range of depth values, the rendering step can provide more accurate results because the depth regions correspond to a range of depth values. Thus, any errors in depth values that may occur during rendering due to, for example, warping and / or blending steps, can be corrected by including a range of depth values.

[0035] Furthermore, receiving ranges of depth values for various depth regions allows the decoder side to process the depth data much more efficiently. For example, the ranges of values can be used to segment the received depth map. Having depth segmentation for the depth map can provide more accurate and efficient rendering of multi-view frames. Vertex and fragment shaders commonly used in rendering can also utilize the ranges of depth values, as described below.

[0036] For example, processing at least one of the depth maps using a range of depth values may comprise modifying positions of vertices in one or more meshes generated based on the depth maps to fit within corresponding ranges of depth values, and / or modifying transparency values in a transparency map generated based on the depth maps based on one or more ranges of depth values.

[0037] Processing at least one depth map may include rendering depth regions in the one or more depth maps in a back-to-front depth order.

[0038] Processing at least one of the depth maps may include selecting a mesh for each range of depth values, applying depth values of one or more depth regions corresponding to each range of depth values from the depth map to the corresponding mesh, modifying the positions of vertices in each mesh to fit within the corresponding range of depth values, and synthesizing an image at the target viewpoint using the image and the modified mesh or meshes.

[0039] The depth variation within a region (i.e., depth range) can be used to predict which mesh resolution to use. In practice, generating a mesh means allocating memory for a given mesh size. Applying depth values to a mesh means calculating 3D (x, y, z) mesh coordinates based on an input depth map (e.g., without projecting depth onto x, y, z coordinates).

[0040] For example, a mesh can be modified by a vertex shader during rendering. The depth value range provides limits on the positions of vertices within the mesh. This means that erroneous depth values that are outside the depth value range are corrected to be within that range, thus improving the accuracy of the mesh.

[0041] This is advantageous because the mesh can be used for warping operations in multi-view imaging. Thus, having accurate vertex coordinates in the mesh provides more accurate warping to the target viewpoint. Of course, the range of depth values can also be used for other geometry-based shaders / operations in the graphics pipeline to modify the mesh coordinates / geometry, etc.

[0042] Assuming a cell size of 128x128 pixels, various mesh sizes can be generated (e.g., 128x128 mesh, 64x64 mesh, 32x32 mesh, 16x16 mesh, and 8x8 mesh are stored in memory). Depending on the smoothness in the depth region, the depth region can use the lowest resolution mesh possible.

[0043] Smoothness of a depth map is a relative concept. Consider the depth map of a soccer ball. If the mesh is too coarse, texture rendering errors can become noticeable. Thus, while the soccer ball surface is locally fairly smooth, the curvature is still quite large, which may benefit from a higher resolution mesh. Another example is a book placed so that the cover is parallel to the depth image sensor. In this case, the entire cover can be perfectly represented by a single quadrilateral (or two triangles) in the mesh.

[0044] The mesh resolution to be used for a particular depth region can be received along with the range of depth values for that depth region.

[0045] The method may further include obtaining a mesh resolution based on the depth smoothness of each of the one or more depth maps, and selecting a mesh for each range of depth values is based on the mesh resolution.

[0046] A depth map with high depth smoothness (i.e., not much change between pixels) may not require as many vertices in the mesh, since gradual depth changes can be accurately depicted with a small number of vertices.

[0047] Preferably, the depth map is divided into various cells and depth segmentation is performed for each cell in the encoder. The segments resulting from the segmentation are encoded using multiple intervals (i.e., ranges of depth values) per cell, each interval separating a region of the cell.

[0048] Each region of multiple depth pixels has specific smoothness characteristics. For example, a given region may be of constant depth, and therefore a mesh consisting of a single quadrangle (two triangles) is sufficient. However, another region may contain ridges or curved surfaces. To represent that interval / region, a mesh with high spatial resolution is required. This allows portions of the depth map with high depth smoothness (depth regions within cells) to use a mesh with fewer vertices, thereby reducing computational complexity, while portions of the depth map with lower depth smoothness (higher variability of depth values) can use a mesh with more vertices to improve accuracy.

[0049] Processing at least one of the depth maps may include warping the one or more depth maps to the target viewpoint, generating a transparency map using the one or more warped depth maps, and modifying transparency values in the transparency map based on one or more ranges of depth values, and synthesizing the image at the target viewpoint is based on using the one or more warped depth maps and the transparency map.

[0050] A transparency map is typically used in a layered format-based multiview frame to set each pixel for each layer as transparent (or partially transparent). When a depth value in the depth map is outside the corresponding range of depth values, the transparency value can be modified so that the corresponding transparency value is set to 100% transparent. This allows erroneous depth values (e.g., caused by warping and / or blending steps) to be hidden in the multiview frame.

[0051] Additionally, the transparency map modification inherently segments the foreground edges of objects in each layer, since objects in different layers are likely to have depth values outside the range of depth values corresponding to that layer.

[0052] For example, this can be achieved in a fragment shader of the graphics pipeline. Of course, other pixel-based operations can be adapted based on the range of depth values.

[0053] The method may further include receiving an indication of the number and shape of cells in the grid of cells defined for each of the one or more depth maps, and rendering the depth region includes rendering the depth region separately for each cell.

[0054] The present invention also provides a computer program comprising computer program code which, when executed on a processor, causes the processor to perform all of the steps of any (or all) of the methods described above.

[0055] The present invention also provides a processor configured to execute the above-described computer program code.

[0056] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0057] For a better understanding of the present invention and to show more clearly how the same may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Figure 1] FIG. 10 illustrates an example partition with a grid plotted on a depth map. [Figure 2] 1 illustrates a single grid cell of a grid. [Figure 3] FIG. 3 shows the grid cells of FIG. 2 with depth regions A-C labeled. [Figure 4] FIG. 4 shows a one-dimensional profile of the cell shown in FIGS. 2 and 3 with a corresponding range of depth values. [Figure 5]Figure 4 shows a low-resolution mesh drawn over the cells in Figures 2 and 3. [Figure 6] FIG. 10 shows a mesh drawn for depth region A. [Figure 7] 10 shows a mesh drawn for depth region B. [Figure 8] 10 shows a mesh drawn for depth region C. [Figure 9] FIG. 1 illustrates how depth maps are processed in an encoder. [Figure 10] FIG. 10 is a diagram showing a method for synthesizing an image at a target viewpoint. DETAILED DESCRIPTION OF THE INVENTION

[0058] The present invention will now be described with reference to the drawings.

[0059] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems, and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar parts.

[0060] The present invention provides a method for processing a depth map for subsequent transmission: depth regions in the depth map are identified, each of the depth regions defining a portion of the depth map; depth value ranges are determined for the depth regions and are further transmitted with the depth map; the depth value ranges provide segmentation information for the depth regions in the depth map.

[0061] To avoid extra pixel space requirements for segmentation and the computational complexity associated with packing, it is proposed to transmit segmentation information as a set of thresholds for a depth map rather than as a spatial map. In particular, the segmentation information is transmitted as a range of depth values, with each range of depth values containing one or more objects in the scene. Preferably, the set of thresholds is specified for a predefined region geometry (e.g., a rectangle). Then, by applying the range of depth values to the depth map, an approximate segmentation map can be reconstructed after decoding by the client. The range of depth values can be transmitted as metadata. Note that the range of depth values may also be referred to as a threshold or a depth interval.

[0062] The first step is to create an upper bound and shape for the segment size by cutting the image region into a regular grid of pieces. For example, assuming an image size of 1920x1080 pixels, one can use grid cells of 128x128 pixels.

[0063] In practice, both the image and the corresponding depth map may be used with the grid cells. Thus, texture and depth may be defined for each grid cell and then used on the client side to synthesize a new image at the target viewpoint. However, in this case, processing the depth map is sufficient to define the segmentation information.

[0064] 1 shows an example subdivision having a grid 102 plotted on a depth map 100. The grid 102 defines a number of grid cells. Note that the depth map 100 may have a lower resolution than the corresponding image, for example, four times lower. Thus, the depth map may have a grid cell size of 32x32 pixels.

[0065] 1 shows a person 104 and a table 106, with the shading of the table 106, compared to the shading of the person 104, indicating that the table 106 is closer to the viewpoint of the depth map 100 than the person 104. The background 108 is also shown without shading, which can indicate that the depth value of the background 108 is uncertain or that the depth value is outside a predefined depth region.

[0066] Most grid cells will not contain any occluding edges, and depth segmentation and a front-to-back rendering approach will not be required. However, some grid cells will contain occluding edges. In this case, more distant objects (e.g., background 108) will need to be drawn first during rendering, and closer objects (e.g., person 104 and table 106) will need to be composited on top of the drawn result in a later rendering step. Therefore, in this case, having segmentation information during rendering is highly advantageous, as it allows for more accurate rendering of edges between objects at different depths.

[0067] Using depth segmentation information in combination with back-to-front view compositing with compositing also has computational complexity advantages. On the client device, the depth map is typically converted to a regular mesh in the vertex shader. For occluded edges, the mesh must have many vertices, and because the mesh is regular, the vertex shader is computationally expensive. Perhaps millions of vertices need to be rendered per view per frame. In contrast, depth segmentation reduces depth changes (within a grid cell) and therefore the number of vertices required for depth steps with occlusion. The foreground / background transition (i.e., the depth step) can be effectively transferred to the transparency map of the grid cell. For many of the cells shown in Figure 1, four mesh vertices, one at each corner of the grid cell, are sufficient. This is especially true for cells containing only one object.

[0068] FIG. 2 shows a single grid cell 200 of grid 102. In this case, three objects are shown at different depths. The closest object is a rod 202, then a circle 204, and finally a background 206. The depth of circle 204 varies from left to right. Using the smoothness of depth, connected regions can be labeled within cell 200. In this case, rod 202 separates background 206 into connected regions 1 and 5, and circle 204 separates into connected regions 2 and 4. The portion of rod 202 within cell 200 is labeled as connected region 3.

[0069] It is important to realize that for client-side occlusion-aware rendering, regions do not necessarily have to be connected regions. In other words, pixels with the same or similar depth are blended across the view, but composited back-to-front per depth layer. This allows us to convert connected regions 1-5 into depth regions A, B, C based on the range of depth values.

[0070] Figure 3 shows grid cell 200 from Figure 2 labeled with depth regions A-C. As can be seen, the original connected region pairs [1, 5] and [2, 4] can be represented by single depth regions A and B, respectively. Connected region 3 can be represented by depth region C. As a result, each pixel in cell 200 receives one of three depth class labels (A, B, or C), with each class corresponding to a single depth range (minimum and maximum depth values).

[0071] 4 illustrates a one-dimensional profile of the cell 200 shown in FIGS. 2 and 3 with a corresponding range of depth values. The range of depth values is sometimes referred to as a depth interval. In this case, depth region A (i.e., background 206) corresponds to depth interval 402, depth region B (i.e., circle 204) corresponds to depth interval 404, and depth region C (i.e., rod 406) corresponds to depth interval 406. The depth intervals can be calculated by visiting all pixels in each depth region and tracking the minimum and maximum depth values in the depth map.

[0072] The depth interval (i.e., depth value range) encodes an approximate segmentation, which can be used to reconstruct the class labels after decoding, as will be explained later.

[0073] One way to encode the depth intervals is to assume a predefined scan order over the grid cells and store, for each grid cell, the number of depth classes and the corresponding depth interval.

[0074] Note that depth intervals may overlap. During rendering, this would mean that the same pixel is rendered multiple times. However, this does not happen very often, so the encoder overhead to remove overlapping depth intervals is likely not worth the limited reduction in computational rendering cost.

[0075] A maximum limit can be set on the number of depth value ranges represented for a grid cell. For example, with 4 bits, it is possible to encode 16 depth classes. Note that it is necessary to be able to represent the situation where no depth classes exist (i.e., the entire grid cell is transparent), so in this case it is possible to represent no depth classes + 15 depth classes using 4 bits. Using 4-bit encoding, 128x128 pixel grid cells, and 1920x1080 image resolution, 15x8x4 = 480 bits are needed per frame to encode the number of depth classes.

[0076] The depth intervals themselves can be coded with, for example, 16-bit precision. Of course, the range of depth values can also be coded differentially across space and / or time.

[0077] In summary, including the range of depth values in the metadata for each cell adds low bitrate overhead and can be encoded at relatively low computational cost.

[0078] On the client device, the multiview texture is received along with the corresponding depth map, along with the depth interval for each grid cell (as metadata). The decoding itself typically occurs during rendering on a graphics processing unit (GPU). The depth interval can be used in both the vertex and fragment shaders.

[0079] In the vertex shader, a depth interval can be used to clip the depth of mesh vertices within a range of depth values, so that pixels within the depth region are warped correctly, and pixels within the same cell but outside the depth region are not drawn onto that depth region.

[0080] In the fragment shader, the depth interval can be used to determine the alpha channel (i.e. the transparency value of the texture): if the warped depth is outside the depth interval, the corresponding pixel can be set to transparent.

[0081] Each time a grid cell for a given depth layer and a given view is rendered, the corresponding depth interval can be passed to the GPU as a shader uniform parameter.

[0082] Starting with the vertex shader, low-resolution mesh vertices are drawn onto the depth map pattern. Figure 5 shows the low-resolution mesh drawn onto the cell 200 from Figures 2 and 3. It is important to understand that significant computational savings are achieved at this point in rendering. The mesh resolution per grid cell can be further reduced as long as the thin rod 204 is "hit" by at least one mesh vertex 502. Figure 5 shows that for the illustrated low-resolution mesh, at least three vertices 502 hit the thin object.

[0083] Note that rendering is done in back-to-front order. Thus, as shown in Figure 3 and Figure 4, the first depth region is depth region A. The vertex shader receives the range of depth values in depth region A and clips the depth values of vertices sampled within the range of depth values in depth region A in parallel for all vertices. Since there is little depth variation in depth region A, all vertices receive approximately the same depth value.

[0084] 6 shows a mesh 600 drawn for depth region A. Objects 202 and 204 are shown for context. The coordinates / depth values of vertex 602 are clipped to fit within the range of depth values for depth region A.

[0085] The second depth region that is rendered has the label B. Figure 7 shows the mesh 700 that is rendered for depth region B. Here, there is more depth variation due to the circular depth profile (shown in Figure 4). However, because the depth variation across the circle varies relatively smoothly in spatial coordinates, a lower resolution mesh is sufficient.

[0086] In this case, vertex 702 drawn on circle 202 is given the circle's corresponding depth value. Vertex 704 drawn across depth region B (i.e., across the background) is clipped to have the maximum depth value within the range of depth values corresponding to depth region B. Similarly, vertex 706 drawn across depth region C (i.e., across rod 204) is clipped to have the minimum depth value within the range of depth values corresponding to depth region B.

[0087] Finally, a third class C is drawn. Figure 8 shows the mesh 800 drawn for depth region C. As can be seen, the three samples that "hit" rod 204 set a relatively small range of depth values (as can be seen from range 406 in Figure 4), effectively clipping all vertices to the same depth value corresponding to the depth of rod 204.

[0088] In the vertex shader, the ray angle difference between the source and target rays can be used to pre-compute a weight that is passed to the fragment shader, which can be coded as transparency and later used for blending against the source view contribution.

[0089] The computation of texture transparency (alpha) on each mesh can be performed using the range of depth values corresponding to that mesh in combination with the (warped) higher resolution depth map. Note that within the cell, alpha is set to the ray angle weight. In later blending steps on the view, alpha (if above a threshold) is interpreted as a blend weight for blending the contribution of the source view, and (if below the threshold) as alpha, which is used for compositing against the previously rendered segment.

[0090] However, the alpha value may be modified depending on the depth value of the warped depth map compared to the range of corresponding depth values. In particular, within a cell, if the depth value of the warped depth map is outside the range of the corresponding depth value of that pixel, the alpha value may be set to 0 (i.e., transparent). This means that any pixel outside the range of depth values is set to transparent. Returning to the mesh of Figure 8, the texture of mesh 800 is transparent everywhere away from the location where the mesh is drawn on rod 204.

[0091] In conclusion, the vertex shader operates on the low-resolution mesh for each grid cell, decodes the segmentation information using the range of depth values, and warps the texture patch to the target viewpoint.

[0092] This is not the case for fragment shaders, where new foreground edges are composited pixel-accurately with the already drawn scene content. Pixels that are not within the depth interval are set to transparent, resulting in good segmentation of the foreground edges. To do this, a high-resolution depth map is warped towards the target viewpoint. The fragment shader thresholds the warped high-resolution depth map using a range of depth values and sets all pixels outside the depth interval to transparent.

[0093] Note that in the above examples we assumed that the mesh resolution is constant and lower than the depth map to save on expensive vertex shader computations. For example, with a quad size of 4 pixels, 4 vertices can represent a 4x4 = 16 pixel image patch, and this entire patch can be warped using 2 triangles. This is a clear complexity reduction over having 2 triangles per pixel, which is required for traditional depth image-based rendering.

[0094] However, computational complexity can be further reduced by allowing the mesh resolution to vary per grid cell. For example, the grid cell covering the table in Figure 1 requires only a single quad to perform an accurate viewpoint shift (because the table is planar). This means that many grid cells can use only four vertices to warp the entire 128x128 image patch. Grid cells whose depth smoothness magnitude (e.g., the first and / or second spatial derivatives of the depth map) exceeds a threshold may require more vertices for accurate warping. Therefore, mesh resolution can be calculated on the production side (i.e., in the encoder before transmission / broadcast) and transmitted as metadata.

[0095] This can be achieved, for example, by encoding the quad size (expressed as image pixels) per grid cell as an 8-bit number to represent quad sizes of 2, 4, 8, 16, 32, etc. Temporal differential encoding of this quad size parameter can also further reduce the bit rate, since on average, depth values within a grid cell do not change very quickly.

[0096] 9 shows a method for processing a depth map in an encoder. In step 902, a depth map of a scene is acquired. For example, a depth sensor can acquire the depth map. Alternatively, the depth map can be generated by calculating disparity vectors between two images of a scene taken from different viewpoints.

[0097] Depth regions are identified in the depth map in step 904. Depth regions can be identified using regions of similar and / or smooth depth. For example, depth values may cluster at a particular depth, indicating the presence of an object. A depth region generally indicates one or more objects in a scene. For example, a depth region may include pixels in the depth map with similar depth values and / or with relatively smooth depth variations (e.g., using first / second spatial derivatives of the depth map). Of course, conventional segmentation can also be used to identify depth regions. In this case, there is no need to transmit the segmentation mask obtained by segmentation.

[0098] In step 906, a range of depth values is determined for the depth region. For example, the maximum and minimum depth values in a depth region can be used to define the range of depth values for that depth region. Of course, other methods may be used to determine the range of depth values (e.g., a 5-point distribution of depth values). th percentiles and 95 th percentile).

[0099] Thus, the depth map and depth value ranges are transmitted in steps 908 and 910, respectively. The depth value ranges provide low bit-rate segmentation information for the depth map. The depth value ranges can be transmitted along with coded identification information that provides which pixels correspond to which depth value ranges.

[0100] 10 illustrates a method for synthesizing an image at a target viewpoint. Steps 1002, 1004, and 1006 involve receiving an image (texture patch) of a scene, a depth map of the scene, and a depth value range for the depth map, respectively. For example, a client decoder can receive these components from an encoder. The depth map is then processed using the depth value range and the image.

[0101] In a first example, the depth map and range of depth values are used in the vertex and / or fragment shader as described above.

[0102] In a second example, a depth map may be warped and blended to a target viewpoint, and the resulting depth map may then be modified based on the range of depth values (e.g., to avoid any errors due to the warping and blending).

[0103] In a third example, the decoder can identify which range of depth values each pixel corresponds to by using the depth value at that pixel before warping the depth map. Thus, a pseudo-segmentation map can be generated for each depth region, thereby providing segmentation of the depth region without having to perform any segmentation at the decoder and without the encoder having to transmit the complete segmentation map.

[0104] More generally, the range of depth values can be used as depth segmentation information in a graphics pipeline for use in synthesizing novel images at a target viewpoint.

[0105] The processed depth map can then be used together with the image to synthesize a new image at the target viewpoint.

[0106] Those skilled in the art can easily develop a processor to perform any of the methods described herein. Accordingly, each step in the flowchart may represent a respective operation performed by a processor, and may be performed by a respective module of the processor.

[0107] As described above, the system utilizes a processor to process data. The processor may be implemented in a variety of ways using software and / or hardware to perform the various functions required. The processor typically uses one or more microprocessors that are programmed using software (e.g., microcode) to perform the required functions. The processor may also be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0108] Examples of circuitry that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).

[0109] In various implementations, the processor may be associated with one or more storage media, e.g., volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. The various storage media may be mounted within the processor or controller, or may be transportable such that one or more programs stored on the storage media can be read by the processor.

[0110] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0111] The functions implemented by a processor may be implemented by a single processor or by multiple separate processing units, which may be considered to constitute a "processor". Such processing units may be remote from each other and may communicate with each other via wired or wireless means.

[0112] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0113] The computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0114] When the term "adapted for" is used in the claims or description, it is meant to be equivalent to the term "configured to." When the term "apparatus" is used in the claims or description, it is intended to be equivalent to the term "system," and vice versa.

[0115] Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. 1. A method for processing a depth map for subsequent transmission, comprising: obtaining a depth map of a scene comprising a plurality of depth values; defining a plurality of cells in the depth map, each cell corresponding to a portion of the depth map; identifying one or more depth regions in each of the cells; determining one or more ranges of depth values for the one or more depth regions; transmitting the depth map and the one or more ranges of depth values; A method having the following.

2. The method of claim 1 , wherein defining a plurality of cells in the depth map comprises dividing the depth map into a grid of cells.

3. The method of claim 1 or 2, further comprising transmitting a number of ranges of different depth values in each cell.

4. The method of claim 3 , wherein only a range of 16 or fewer different depth values is transmitted.

5. The method of claim 1 , wherein identifying one or more depth regions is based on determining a depth smoothness between depth values of the depth map.

6. determining a mesh resolution based on the depth smoothness; and transmitting said depth resolution; The method of claim 5 further comprising:

7. 1. A method for synthesizing an image at a target viewpoint, comprising: receiving at least one or more images and one or more depth maps of a scene; receiving one or more ranges of depth values for a plurality of cells in at least one of the depth maps, each range of depth values corresponding to one or more depth regions in one of the depth maps, each cell corresponding to a portion of the depth map; and processing at least one of the depth maps using the range of depth values for each cell to synthesize an image at the target viewpoint.

8. The method of claim 7 , wherein processing at least one depth map comprises rendering the depth regions in the one or more depth maps in a back-to-front depth order.

9. processing at least one of the depth maps, selecting a mesh for each range of depth values; applying the depth values of the one or more depth regions corresponding to each depth value range from the depth map to a corresponding mesh; modifying the positions of the vertices in each mesh to fall within a range of corresponding depth values; synthesizing an image at the target viewpoint using the image and the modified mesh(es); 9. The method of claim 7 or 8, comprising:

10. 10. The method of claim 9, further comprising obtaining a mesh resolution based on a depth smoothness of each of the one or more depth maps, wherein selecting a mesh for each range of depth values is based on the depth smoothness.

11. processing at least one of the depth maps, warping the one or more depth maps to a target viewpoint; modifying transparency values of a transparency map based on the one or more depth value ranges; and The method of claim 7 , wherein synthesizing an image at the target viewpoint is based on using one or more warped depth maps and the transparency map.

12. 11. The method of claim 7, further comprising receiving an indication of the number and shape of cells in a grid of cells defined for each of the one or more depth maps, and wherein rendering the depth region renders the depth region of each cell individually.

13. A computer program product which, when executed by a processor, causes the processor to carry out the method of any one of claims 1 to 12.

14. A processor configured to execute the program of claim 13.