Apparatus and method for processing depth maps
The method enhances depth map processing by selecting candidate depth values based on a cost function, prioritizing distant pixels, addressing accuracy and complexity issues in depth estimation for improved three-dimensional image rendering.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2026-03-03
AI Technical Summary
Existing depth estimation techniques are suboptimal in accuracy and computational efficiency, leading to artifacts and noise in three-dimensional image processing, particularly in applications requiring multiple views and dynamic virtual reality experiences.
A method for processing depth maps involves determining a set of candidate depth values for a pixel, calculating a cost value using a cost function, and selecting an updated depth value based on these values, with a focus on candidate values further away along a specific direction to enhance accuracy and reduce computational complexity.
This approach provides more accurate and consistent depth maps with reduced complexity, suitable for integration with disparity-based depth estimation, improving three-dimensional image quality and rendering performance.
Smart Images

Figure 0007823048000003 
Figure 0007823048000004 
Figure 0007823048000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for processing depth maps, particularly but not exclusively to processing depth maps for multi-view depth / disparity estimation. [Background technology]
[0002] Traditionally, the technological processing and use of images has been based on two-dimensional imaging, but increasingly the third dimension is explicitly considered in image processing.
[0003] For example, three-dimensional (3D) displays have been developed that add a third dimension to the viewing experience by providing different views of the scene being viewed to the viewer's two eyes. This can be achieved by the user wearing glasses to separate the two views being displayed. However, this is considered inconvenient for the user, so in many scenarios it is preferable to use an autostereoscopic display that uses a means in the display (such as a lenticular lens or a barrier) to separate the views, sending the views in different directions that reach the user's eyes individually. While a stereoscopic display requires two views, an autostereoscopic display typically requires more views (e.g., nine views).
[0004] Another example is the use of free viewpoints, which allows (within limits) spatial navigation of a scene captured by multiple cameras. This can be done, for example, on either a smartphone or a tablet, providing a game-like experience. Alternatively, the data can be viewed on an augmented reality (AR) or virtual reality (VR) headset.
[0005] In many applications, it is desirable to generate view images for new viewing directions. Various algorithms are known for generating such new view images based on image and depth information, but they tend to be highly dependent on the accuracy of the depth information provided (or derived).
[0006] In practice, three-dimensional image information is provided by multiple images corresponding to different viewing directions relative to the scene. Such information can be captured using dedicated 3D camera systems that capture two or more simultaneous images from offset camera positions.
[0007] However, in many applications, the images provided may not directly correspond to the desired orientation, or more images are required. For example, for auto-view displays, more than two images are required, and in practice, 9 to 26 view images are often used.
[0008] To generate images corresponding to different viewing directions, a viewing direction shift process may be employed. This is typically done by a viewing direction shift algorithm that uses an image from a single viewing direction together with associated depth information (or possibly multiple images and associated depth information). However, the provided depth information must be sufficiently accurate to generate new view images without significant artifacts.
[0009] Other exemplary applications include virtual reality experiences in which right-eye and left-eye views are continuously generated for a virtual reality headset to match movements and orientation changes by a user. Such generation of dynamic virtual reality views may often be based on light intensity images in combination with associated depth maps that provide relevant depth information.
[0010] The quality of one or more three-dimensional images presented from a new view depends on the quality of the received image and depth data; specifically, three-dimensional perception depends on the quality of the received depth information. Other algorithms or processes that rely on image depth information are known, and these also tend to be very sensitive to the accuracy and reliability of the depth information.
[0011] However, in many practical applications and scenarios, the depth information provided tends to be suboptimal, and indeed in many practical applications and usage scenarios, the depth information is not as accurate as desired, which may introduce errors, artifacts, and / or noise in the processing and in the generated images.
[0012] In many applications, depth information describing a real-world scene is inferred from depth cues determined from captured images. For example, depth information may be generated by estimating and extracting depth values by comparing view images for different view positions.
[0013] For example, in many applications, a three-dimensional scene is captured as a stereo image using two cameras at slightly different positions. A specific depth value is then generated by estimating the disparity between corresponding image objects in the two images. However, such depth extraction and estimation is problematic and prone to non-ideal depth values, which again results in artifacts and reduced three-dimensional image quality.
[0014] To improve depth information, several techniques have been proposed for post-processing and / or refining depth estimates and / or depth maps. However, these tend to be suboptimal, not optimally accurate and reliable, and / or may be difficult to implement, for example, due to the computational resources required. Examples of such algorithms are described in WO2020 / 178289A1 and EP3396949A1.
[0015] A specific approach has been proposed in which a depth map is initialized and then iteratively updated using a scanning approach, where the depth of the current pixel is updated based on a candidate set of candidate depth values, which are typically the depth values of neighboring pixels. The update of the depth value of the current pixel is dependent on a cost function. However, while such approaches improve the depth map in many scenarios, they tend to be suboptimal in all scenarios, including not always producing an optimally accurate depth map. They also tend to be computationally intensive due to the large number of candidate pixels that must be considered.
[0016] Therefore, improved approaches to generating / processing / modifying depth information would be advantageous, particularly approaches to processing depth maps that enable increased flexibility, facilitated implementation, reduced complexity, reduced resource requirements, improved depth information, more reliable and / or accurate depth information, an improved 3D experience, improved quality of rendered images based on the depth information, and / or improved performance. Summary of the Invention [Problem to be solved by the invention]
[0017] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0018] According to an aspect of the present invention, there is provided a method for processing a depth map, the method comprising: receiving a depth map; determining a set of candidate depth values for at least a first pixel of the depth map, the set of candidate depth values including depth values for other pixels of the depth map other than the first pixel; determining a cost value for each of the candidate depth values in the set of candidate depth values according to a cost function; selecting a first depth value from the set of candidate depth values according to the cost value of the set of candidate depth values; and determining an updated depth value for the first pixel according to the first depth value, the set of candidate depth values including a first candidate depth value along a first direction from the first pixel; a first intervening pixel set of at least one pixel along the first direction does not include a candidate depth value of the set of candidate depth values for which the cost function for the candidate depth value does not exceed the cost function for the first candidate depth value; and a distance from the first pixel to the first candidate depth value is greater than a distance from the first pixel to the first intervening pixel set.
[0019] The present invention refines depth maps to provide improved three-dimensional image processing and perceived rendering quality. In particular, the present approach provides more consistent and / or accurate depth maps in many embodiments and scenarios. The present process, in many embodiments, provides improved depth maps while maintaining sufficiently low complexity and / or resource requirements.
[0020] An advantage in many embodiments is that the approach is well suited for use and integration with depth estimation techniques, such as disparity-based depth estimation using stereo or multi-view images.
[0021] In particular, the present approach refines depth maps using a relatively low complexity and low resource demand approach, for example, allowing sequential bit scanning and processing with relatively few decisions per pixel, which is sufficient to increase overall accuracy.
[0022] A depth map indicates depth values for pixels of an image, which may be any value that indicates depth, including, for example, a disparity value, a z-coordinate, or a distance from viewpoint value.
[0023] The processing of the first pixel may be repeated with a new pixel of the depth map selected for each iteration. The selection of the first pixel depends on a scan sequence over the depth map. The pixel corresponds to a location / area in the depth map for which a depth value is provided. A pixel in the depth map corresponds to one or more pixels in the associated image for which the depth map indicates depth. A depth map is formed by a two-dimensional configuration of pixels with a depth value provided for each pixel. Thus, each pixel / depth value is provided for an area (of pixels) of the depth map. A reference to a pixel is a reference to a depth value (for the pixel) and vice versa. A reference to a pixel is a reference to a location in the depth map for a depth value. Each pixel of the depth map is associated with one depth value (and vice versa).
[0024] The cost function is implemented as a merit function, and the cost value is indicated by a merit value. The selection of the first depth value according to the cost value is implemented as the selection of the first depth value according to the merit value determined from the merit function. An increasing merit value / function is a decreasing cost value / function. The selection of the first depth value is the selection of the candidate depth value from the set of candidate depth values with the lowest cost value, which corresponds / equals the selection of the first depth value from the set of candidate depth values with the highest merit value.
[0025] The updated depth value has a value that is dependent on the first depth value. This results in an updated depth value that is the same as the depth value of the first pixel before processing, depending on the situation and possibly the pixel. In some embodiments, the updated depth value is determined as a function of the first depth value, and specifically as a function that is less dependent on other depth values in the depth map than the first depth value. In many embodiments, the updated depth value is set equal to the first depth value.
[0026] In some embodiments, the first direction may be the only direction in which the intervening set of pixels as described lies. In some embodiments, the first direction may be a direction at an angular interval of a direction from the first pixel in which the intervening set of pixels as described lies. The angular interval is or has a span / width / extent of not more than 1°, 2°, 3°, 5°, 10°, or 15°. The first direction may, in such embodiments, be replaced with a reference to a direction within such angular interval.
[0027] The set of candidate depth values includes a first candidate depth value along the first direction that is further away from the first pixel than at least one pixel along the first direction that is not included in the set of candidate depth values or has a higher cost function than the first candidate depth value.
[0028] According to an optional feature of the invention, the cost function along the first direction has a monotonically increasing cost gradient as a function of distance from the first pixel when the distance is below a distance threshold, and a decreasing cost gradient as a function of distance from the first pixel when at least one distance from the first pixel is above a threshold.
[0029] This provides improved performance and / or implementation in many embodiments by ensuring that there is a first intervening pixel set of at least one pixel that does not include a candidate depth value of the set of candidate depth values for which the cost function for the candidate depth value does not exceed the cost function for the first candidate depth value, and that the distance from the first pixel to the first candidate depth value is greater than the distance from the first pixel to the first intervening pixel set.
[0030] A monotonically increasing slope as a function of distance is a slope that always increases or remains constant for increasing distance.
[0031] According to an optional feature of the invention, the first set of intervening pixels is a set of pixels whose depth values are not included in the set of candidate values.
[0032] This, in many embodiments, provides improved performance and / or implementation by ensuring that there is a first intervening pixel set of at least one pixel that does not include a candidate depth value of the set of candidate depth values for which the cost function for that candidate depth value does not exceed the cost function for the first candidate depth value, and the distance from the first pixel to the first candidate depth value is greater than the distance from the first pixel to the first intervening pixel set.
[0033] According to an optional feature of the invention, the cost function includes a cost contribution that is dependent on the difference between image values of the multi-view images for pixels offset by a disparity that matches the candidate depth value to which the cost function is applied.
[0034] This approach provides an advantageous depth estimation approach based on multi-view images combined with consideration of multi-view disparity, whereby, for example, an initial depth map can be iteratively updated based on correspondences between different images of the multi-view image.
[0035] According to an optional feature of the invention, the method further comprises determining the first direction as a gravity direction of the depth map, the gravity direction being a direction in the depth map that matches a direction of gravity in a scene represented by the depth map.
[0036] This provides particularly efficient performance and improved depth maps, and exploits typical properties of the scene to provide improved and often more accurate depth maps.
[0037] According to an optional feature of the invention, the first direction is a vertical direction in the depth map.
[0038] This provides particularly efficient performance and improved depth maps, and exploits typical properties of the scene to provide improved and often more accurate depth maps.
[0039] According to an optional feature of the invention, the method further comprises determining a depth model for at least a portion of the scene represented by the depth map, the cost function for a depth value being dependent on a difference between the depth value and a model depth value determined from the depth model.
[0040] This provides particularly efficient performance and improved depth maps, and exploits typical properties of scene objects to provide improved and often more accurate depth maps. The cost function increases the cost for increasing differences between depth values and model depth values.
[0041] According to an optional feature of the invention, the cost function is asymmetric with respect to whether the depth value is above or below the model depth value.
[0042] This provides a particularly advantageous depth map in many embodiments and scenarios.
[0043] According to an optional feature of the invention, the depth model is a background model for the scene.
[0044] This provides a particularly advantageous depth map in many embodiments and scenarios.
[0045] According to an optional feature of the invention, the method further comprises including candidate depth values in the set of candidate depth values not from a depth map, and including at least one depth value from another depth map of a temporal sequence of depth maps, the sequence including a depth map, a depth value that is independent of a scene represented by the depth map, and a depth value determined according to an offset of a depth value relative to the first pixel.
[0046] This provides a particularly advantageous depth map in many embodiments and scenarios.
[0047] According to an optional feature of the invention, the cost function for the depth values depends on the type of the depth values, the type being one of a group of types including at least one of a depth value of the depth map, a depth value of the depth map that is closer than a distance threshold, a depth value of the depth map that is farther away than a distance threshold, a depth value from another depth map of the temporal sequence of depth maps that includes the depth map, a depth value having a scene-independent depth value that is offset with respect to the depth value of the first depth value, a depth value that is independent of the scene represented by the depth map, and a depth value determined according to an offset of the depth value for the first pixel.
[0048] This provides a particularly advantageous depth map in many embodiments and scenarios.
[0049] According to an optional feature of the invention, the method is configured to process a plurality of pixels of the depth map by iteratively selecting a new first pixel from the plurality of pixels and performing a step for each new first pixel.
[0050] According to an optional feature of the invention, the set of candidate depth values in the second direction from the first pixel includes a pixel set of at least one pixel along the second direction that does not include a second candidate depth value among the set of candidate depth values for which the cost function for the candidate depth value does not exceed the cost function for the second candidate depth value, and the distance from the first pixel to the second candidate depth value is greater than the distance from the first pixel to the pixel set.
[0051] This provides a particularly advantageous depth map in many embodiments and scenarios, and in some embodiments, the first direction may be the only direction in which there is an intervening set of pixels as described.
[0052] According to an aspect of the present invention, there is provided an apparatus for processing a depth map, the apparatus including: a receiver for receiving a depth map; and a processor for processing the depth map, the processing including: determining a set of candidate depth values for at least a first pixel of the depth map, the set of candidate depth values including depth values for other pixels of the depth map other than the first pixel; determining a cost value for each of the candidate depth values in the set of candidate depth values according to a cost function; selecting a first depth value from the set of candidate depth values according to the cost value for the set of candidate depth values; The method includes performing a step of determining an updated depth value for the first pixel according to the depth value, wherein the set of candidate depth values includes a first candidate depth value along a first direction from the first pixel, and along the first direction, a first intervening pixel set of at least one pixel does not include a candidate depth value from the set of candidate depth values for which a cost function for the candidate depth value does not exceed the cost function for the first candidate depth value, and the distance from the first pixel to the first candidate depth value is greater than the distance from the first pixel to the first intervening pixel set.
[0053] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.
[0054] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Brief explanation of the drawings]
[0055] [Figure 1] 1 illustrates an example of a method for processing a depth map according to some embodiments of the present invention. [Figure 2] 1 illustrates an example of an apparatus for processing depth maps according to some embodiments of the present invention. [Figure 3] Some example depth maps are shown. [Figure 4] 1 shows an example of some pixels and depth values of a depth map. [Figure 5] 1 shows an example of some pixels and depth values of a depth map. [Figure 6] 10 illustrates an example of a cost function along a direction for methods according to some embodiments of the present invention. [Figure 7] 1 shows examples of depth maps generated by different processes. DETAILED DESCRIPTION OF THE INVENTION
[0056] The following description focuses on embodiments of the invention that are applicable to processing depth maps of images, and in particular to processing such depth maps as part of a multi-view depth estimation method, however, it will be appreciated that the invention is not limited to this application and applies to many other scenarios.
[0057] Images representing scenes are nowadays supplemented by depth maps that provide information about the depth of image objects in the scene, i.e., provide additional depth data for the image. Such additional information provides many additional services, for example, by enabling viewpoint movement, 3D representation, etc. Depth maps tend to provide a depth value for each of a number of pixels, where the pixels are typically arranged in an array having a first number of horizontal rows and a second number of vertical columns. The depth value provides depth information for an associated image pixel. In many embodiments, the resolution of the depth map is the same as the resolution of the image, so that each pixel of the image has a one-to-one connection to one depth value in the depth map. However, in many embodiments, the resolution of the depth map is lower than that of the image, and in some embodiments, the depth values of the depth map are common to multiple pixels of the image (specifically, depth map pixels are larger than image pixels).
[0058] A depth value may be any value that indicates depth, including in particular a depth coordinate value (e.g., which directly provides the z-value of a pixel) or a disparity value. In many embodiments, a depth map is a rectangular array (with rows and columns) of pixels, where each pixel provides a depth ( / disparity) value.
[0059] The accuracy of the representation of the depth of a scene is an important parameter in the resulting quality of the image rendered and perceived by a user. Therefore, generating accurate depth information is important. For artificial scenes (e.g., computer games), it is relatively easy to achieve accurate values, but for applications that require, for example, the capture of real-world scenes, this is very difficult.
[0060] Several different approaches to estimating depth have been proposed. One approach is to estimate disparity between different images capturing a scene from different viewpoints. However, such disparity estimation is inherently imperfect. Furthermore, this approach requires that the scene be captured from multiple directions, which is not the case, for example, for legacy capture. Another option is to perform motion-based depth estimation, which exploits the fact that image object motion (in a sequence of images) tends to be higher for objects closer to the camera than for objects further away (e.g., in the case of a translational camera, perhaps after correcting for the actual motion of the corresponding object in the scene). A third approach is to exploit predetermined (assumed) information about the depth in the scene. For example, in outdoor scenes (and indeed most typical indoor scenes), lower objects in the image tend to be closer than higher objects in the image (e.g., the floor or ground increases in distance to the camera with increasing height, and the sky tends to be further back than, for example, the low ground). Therefore, a predetermined depth profile is used to estimate suitable depth map values.
[0061] However, most depth estimation techniques tend to fall short of perfect depth estimation, but improved and usually more reliable and / or accurate depth values are advantageous for many applications.
[0062] In the following, an approach for processing and updating depth maps is described. In some embodiments, the approach is used as part of depth estimation, in particular as a depth estimation algorithm that considers disparity between multiple images capturing a scene from different viewpoints; indeed, the process may be an integrated part of a depth estimation algorithm that determines depth from multi-view images. However, it will be understood that this is not required for the approach, and that in some embodiments the approach is applied, for example, as post-processing of estimated depth maps.
[0063] This approach improves depth maps and provides more accurate depth information in many scenarios. Furthermore, it is suitable for combining different depth estimation considerations and approaches and can be used to improve depth estimation.
[0064] The approach will be described with reference to FIG. 1, which shows a flow chart of a method for processing a depth map, and FIG. 2, which shows elements of a corresponding apparatus for carrying out the method.
[0065] The device of Figure 2 may specifically be a processing unit such as a computer or a processing module. As such, it may be implemented by a suitable processor such as a CPU, MPU, DSP, etc. The device may further comprise volatile and non-volatile memory coupled to the processor as known to those skilled in the art. Furthermore, suitable input and output circuits, such as a user interface, a network interface, etc., may be included.
[0066] The method of FIG. 1 begins at step 101, in which receiver 201 receives a depth map. In many embodiments, the depth map is received along with an associated image. Furthermore, in many embodiments, the depth map is received for an image along with one or more images that are further associated with a separate depth map. For example, other images depicting a scene from different viewpoints are received. In some embodiments, the image and depth map are part of a temporal sequence of images and depth maps, e.g., the image is an image or frame of a video sequence. Thus, in some embodiments, receiver 201 receives an image and / or depth map for another time. The depth map and image to be processed are hereinafter also referred to as a first or current depth map and a first or current image, respectively.
[0067] The received first depth map may be the first image that is processed to generate a more accurate depth map, and in particular, in some embodiments, is the initial input to the depth estimation process. The following description focuses on an example in which the method is used as part of a multi-view depth estimation process, where the depth estimation includes consideration of disparity between different images of the same scene from different viewpoints. The process is initialized with an initial depth map that provides a very rough indication of the possible depth of the image.
[0068] For example, an initial depth map may be generated simply by detecting image objects in an image and segmenting the image into image objects and background. The background section may be assigned a predetermined depth, and the image objects may be assigned a different predetermined depth indicating that they are further forward, or a rough disparity estimation may be performed, e.g., based on searching for corresponding image objects in another image, and the resulting estimated depth may be assigned to the entire image object. An example of a resulting depth map is shown in Figure 3(a).
[0069] As another example, a predetermined pattern of depth, such as increasing depth, is assigned to increase height in the image. An example of such a depth map is shown in Figure 3(b) and is suitable for landscape scenes and images, for example. These approaches may be combined, as shown in Figure 3(c).
[0070] In many embodiments, the left / right image pair is used to initialize a block-based 2D disparity vector field, for example, using a pre-adapted 3D depth model. In such embodiments, the depth values may be disparity values or depth values that directly indicate the distance from the viewpoint, or may be calculated from these. This approach takes into account knowledge of the scene's geometric structure, such as the ground and background. For example, based on the 3D model, a 2D disparity field is generated and used as an initial depth map.
[0071] Thus, in the described approach, an initial first image having a rough and usually inaccurate initial first depth map is received by receiver 201, in the example together with at least one other image depicting the scene from a different viewpoint, The approach then processes this depth map to generate a more accurate depth map that better reflects the actual depth of different pixels in the depth map.
[0072] The depth map(s) and (optionally) the image(s) are provided from the receiver 201 to a processor 203, which is arranged to perform the remaining method steps as described below with reference to FIG. 1.
[0073] Although the following approach could in principle be applied to only a subset of depth values / pixels in a depth map, or indeed in principle to only a single pixel in a depth map, the process is typically applied to all or nearly all pixels in a depth map. The application is typically sequential. For example, the process scans through the depth map, sequentially selecting pixels for processing. For example, the method might start from the upper left corner and scan first horizontally and then vertically, i.e., scanning from left to right and top to bottom, until the pixel in the lower right corner is processed.
[0074] Furthermore, in many embodiments, the approach may be iterative and applied repeatedly to the same depth map. In this manner, the resulting depth map from one process / update is used as the input depth map for a subsequent process / update. In many embodiments, the scanning of the depth map with resulting updates of depth values is repeated / iterated multiple times, for example, 5-10 iterations.
[0075] The method begins at step 101 where the receiver 201 receives and forwards a depth map to the processor 203 as described above.
[0076] In step 103, the next pixel is selected. When starting the depth map process, the next pixel is typically a predetermined pixel, such as the top right pixel. Otherwise, the next pixel may be the next pixel according to the sequence of operations applied, specifically a predetermined scanning sequence / order.
[0077] The identified pixel, hereinafter also referred to as the first or current pixel, the method then proceeds to determine a depth value for this pixel. The depth value of the first or current pixel is also referred to as the first or current depth value, and the terms initial and updated depth values are used to refer to the values of the pixel before and after processing, respectively, i.e., the depth values of the depth map before and after the current iteration, respectively.
[0078] In step 105, the processor 203 proceeds to determine / select a set of candidate depth values. The set of candidate depth values is selected to include depth values for a set of candidate pixels. The candidate set of a pixel of the current depth map includes several other pixels in the depth map. For example, the candidate set is selected to include depth values for pixels in a neighborhood around the current pixel, such as a set of pixels within a given distance of the current pixel or within a window / kernel around the current pixel.
[0079] In many embodiments, the candidate set of pixels also includes the depth value of the current pixel itself, ie, the first depth value is itself one of the candidate depth values in the set.
[0080] Furthermore, in many embodiments, the set of candidate depth values also includes depth values from other depth maps, for example, in many embodiments where the image is part of a video stream, one or more depth values from previous and / or subsequent frames / images are also included in the set of candidate depth values, or depth values from other views for which a depth map is simultaneously estimated are included.
[0081] In some embodiments, the set of candidate depth values further includes values that are not directly depth values of the depth map. For example, in some embodiments, the set of candidate depth values includes one or more fixed depth values or relative offset depth values, such as depth values that are greater than or less than the current initial depth value by a fixed offset. In another example, the set of candidate depth values includes one or more random or semi-random depth values.
[0082] Following step 105, in step 107, a cost value is determined for the set of candidate depth values, specifically, a cost value is determined for each candidate depth value in the set of candidate depth values.
[0083] The cost value is determined based on a cost function that depends on several different parameters, as will be described in more detail below. In many embodiments, the cost function of a candidate depth value for a pixel of the current depth map depends on the difference between the image values of the multi-view image offset by the disparity corresponding to the depth value. Thus, for a first depth value, or possibly for each candidate depth value of a set of candidate depth values belonging to the set of candidate depth values, the cost function monotonically decreases as a function of the difference between the two view images of the multi-view image in an image area having a disparity between the two image views that matches the candidate depth value. The image area may specifically be the image area that includes the current pixel and / or the pixel of the candidate depth value. The image area is typically relatively small, e.g., no more than 1%, 2%, 5%, or 10% of the image, and / or no more than 100, 1000, 2000, 5000, or 10,000 pixels.
[0084] In some embodiments, for a given candidate depth value, the processor 203 determines the disparity between the two images that matches the depth value. This disparity is then applied to identify an area in one of the two images that is offset to an area in the other image by that disparity. A measure of the difference between the image signal values, e.g., RGB values, in the two areas is determined. In this way, a difference measure is determined for the two images / image areas based on the assumption that the candidate depth value is correct. The smaller the difference, the more likely the candidate depth value is an accurate reflection of depth. Thus, the lower the difference measure, the lower the cost function.
[0085] The image area is typically a small area around the first / current pixel, and indeed in some embodiments includes only the first / current pixel.
[0086] For a candidate depth value corresponding to the current depth map, the cost function therefore includes a cost contribution that depends on the match between the two multi-view images for the disparity corresponding to the candidate depth value.
[0087] In many embodiments, the cost values for candidate depth values of other depth maps associated with the image, such as temporally offset depth maps and images, also include corresponding image match cost contributions.
[0088] In some embodiments, the costs of some of the candidate depth values do not include an image match cost contribution, e.g., a fixed cost value is assigned for a predetermined fixed depth offset that is not associated with a depth map or image.
[0089] The cost function is typically determined to indicate the likelihood that the depth value reflects the accurate or correct depth value for the current pixel.
[0090] It is understood that determining a merit value based on a merit function and / or selecting a candidate depth value based on a merit value is essentially also determining a cost value based on a cost function and / or selecting a candidate depth value based on a cost value. A merit value can be converted to a cost function simply by applying a function to the merit value, where the function is any monotonically decreasing function. A higher merit value corresponds to a lower cost value; for example, selecting a candidate depth value with the highest merit value is directly equivalent to selecting a cost value with the lowest value.
[0091] Following step 107 of determining a cost value for each candidate depth value in the set of candidate depth values, a depth value from the set of candidate depth values is selected as a function of the cost value for the set of candidate depth values, per step 109. The selected candidate depth value is hereinafter referred to as the selected depth value.
[0092] In many embodiments, the selection is the selection of the candidate depth value for which the lowest cost value has been determined, while in some embodiments, more complex criteria are evaluated that also take into account other parameters (equivalently, such considerations are typically considered part of the (modified) cost function).
[0093] Thus, for the current pixel, the approach selects the candidate depth value that is most likely to reflect the correct depth value for the current pixel as determined by the cost function.
[0094] Following step 109, in step 111, an updated depth value is determined for the current pixel based on the selected depth value. The exact update depends on the particular requirements and preferences of each individual embodiment. For example, in many embodiments, the previous depth value for the first pixel may simply be replaced by the selected depth value. In other embodiments, the update may take into account the initial depth value, e.g., the updated depth value is determined as a weighted combination of the initial depth value and the selected depth value, e.g., the weight is dependent on the absolute cost value for the selected depth value.
[0095] Therefore, in the following step 111, an updated or modified depth value for the current pixel has been determined. In step 113, it is evaluated whether all pixels of the current depth map to be processed have in fact been processed. Typically, this corresponds to determining whether all pixels in the image have been processed, and in particular detecting whether the scanning sequence has reached its end.
[0096] If not, the method returns to step 103 where the next pixel is selected and the process is repeated for this next pixel. Otherwise, the method proceeds to step 115, where it is evaluated whether further iterations should be applied to the depth map. In some embodiments, only one iteration is performed and step 115 is omitted. In other embodiments, the process is iteratively applied to the depth map, for example, until some stopping criterion is achieved (e.g., the overall amount of change that occurs in previous iterations falls below a threshold) or until a predetermined number of iterations have been performed.
[0097] If another iteration is required, the method returns to step 103, where the next pixel is determined as the first pixel in the new iteration. Specifically, the next pixel is the first pixel in the scan sequence. If no further iterations are required, the method proceeds to step 117, where it ends with an updated and typically improved depth map. The depth map can be output, for example, via output circuitry 205, to another function, such as a view synthesis processor.
[0098] While the above description focuses on application to a single pixel, it will be appreciated that a block process may also be applied, for example, where the determined updated depth value is applied to all depth values within the block containing the first pixel.
[0099] The specific selection of candidate depth values depends on the desired operation and performance for a particular application. Typically, the set of candidate depth values includes several pixels in the neighborhood of the current pixel. A kernel, area, or template may overlay the current pixel, and pixels within the kernel / area / template may be included in the candidate set. Additionally, the same pixel in a different time frame (typically just before or just after the current frame for which the depth map is provided) may be included, as well as potentially in a kernel of neighboring pixels that is typically significantly smaller than the kernel of the current depth map. Typically, at least two offset depth values (corresponding to increased and decreased depths, respectively) are also included.
[0100] However, to reduce computational complexity and resource requirements, the number of depth values included in the candidate set is typically significantly limited, particularly since every candidate depth value is evaluated for every new pixel, and typically for every pixel in the depth map for every iteration, resulting in many additional processing steps for each additional candidate value.
[0101] However, it has also been found that the depth map improvement that can be achieved tends to be highly dependent on the particular choice and weighting of the candidates, which in fact tends to have a large impact on the quality of the final depth map. Thus, the trade-off between quality / performance and computational resource requirements is very difficult and very sensitive to the decision on candidate depth values.
[0102] In many typical applications, it is often preferred to have no more than about 5-20 candidate depth values in the set of candidate depth values per pixel. In many practical scenarios, it is necessary to limit the number of candidate depth values to about 10 candidates to achieve real-time processing for video sequences. However, the relatively small number of candidate depth values makes the decision / selection of which candidate depth values to include in the set very critical.
[0103] When determining the set of candidate depth values of a depth map to consider for updating the current depth value, an intuitive approach is to include depth values close to the current pixel, with weighting such that selecting depth values further away is less likely (or at least not more likely), i.e., all other parameters being equal, candidate depth values closer to the current pixel are selected than those further away. Thus, an intuitive approach is to generate the set of candidate depth values as a kernel that includes pixels in a (usually very small) neighborhood around the current pixel, with a cost function that monotonically increases with distance from the current pixel. For example, the set of candidate depth values may be determined as all pixels within a predetermined distance of one or two pixels from the current pixel, and a cost function that increases with distance from the first pixel may be applied.
[0104] However, while such an intuitive approach provides advantageous performance in many embodiments, the inventors have recognized that in many applications advantageous performance can be achieved by taking a counterintuitive approach of increasing the bias for selecting pixels that are more distant along a direction over pixels that are closer in that direction, such that the probability of selecting a more distant pixel increases with respect to pixels that are closer along a direction.
[0105] In some embodiments, this is achieved by a set of candidate depth values along a direction being selected / generated / created / determined to include or include one or more pixels that are further away than one or more pixels that are not included in the set of candidate depth values, for example, the closest pixel or pixels along the direction are included in the set of candidate depth values, followed by one or more pixels that are not included in the set of candidate depth values, followed by one or more further away pixels that are included in the set of candidate depth values.
[0106] An example of such an approach is shown in FIG. 4. In this example, four neighboring depth values / pixels 401 surrounding a current depth value / pixel 403 are included in the set of candidate depth values. The set of candidate depth values may further include the current depth value. However, the set of candidate depth values is further arranged to include depth values / pixels 405 that are far away along a given direction 407. The far away depth values / pixels 405 are at a distance of Δ pixels from the current pixel 403, where Δ>2 and is typically much larger. In this example, the set of candidate depth values further includes a depth value for a pixel 409 that is co-located in the temporally neighboring depth map.
[0107] Thus, in this example, a set of candidate depth values is created that includes only seven depth values, and the system proceeds to determine a cost value by evaluating a cost function for each of these depth values. The system then selects the depth value with the lowest cost value and updates the current depth value, for example, by setting it to the value of the selected candidate depth value. The small number of candidate depth values allows for very fast and / or low resource demanding processing.
[0108] Furthermore, despite constraints on how many candidate depth values are evaluated, an approach that includes not only nearby depth values but also one or more distant depth values has been found to provide particularly advantageous performance in practice, for example, often resulting in a more consistent and / or accurate updated depth map being generated. For example, many depth values represent objects present in the scene and at substantially the same distance. Considering more distant depth values will, in many situations, include depth values that belong to the same object but provide better depth estimates, for example, because they have less local image noise. Consideration of a particular direction reflects likely characteristics of the object, such as, for example, geometric properties or a relationship with the capture orientation relative to the image. For example, for a background, depth tends to be relatively constant horizontally and vary vertically, so identifying pixels that are more distant horizontally increases the likelihood that both the current pixel and the more distant pixel represent the same (background) depth.
[0109] Thus, in some embodiments, the set of candidate depth values is determined to include a first candidate depth value along a direction from the current pixel that is further away than the set of intervening pixels, but along that direction the intervening set includes one or more pixels that are closer to the current pixel than the first candidate depth value but whose depth values are not included in the candidate depth values. In some embodiments, therefore, there is a gap along the direction between the pixel location of the first candidate depth value and the current pixel where there are one or more pixels that are not included in the set of candidate depth values. In many embodiments, the gap is between the first candidate depth value and one or more neighboring pixels that are included in the set of candidate depth values.
[0110] In some embodiments, the candidate set is adapted based on the image object or image object type to which the current pixel belongs. For example, the processor 203 is configured to perform an image object detection process to detect image objects in the depth map (e.g., by detecting them in related images). The candidate set is then adjusted according to the detected image objects. Also, in some embodiments, the first direction may be adapted according to the characteristics of the image object to which the current pixel belongs. For example, the pixel may be known to belong to a particular image object or object type, and the first direction may be determined according to the characteristics of this object, such as the longest direction for the image object. For example, boats on water tend to have a significantly longer extension in the horizontal direction than in the vertical direction, so the first direction is determined as the horizontal direction. Also, a candidate set that extends further horizontally than vertically may be selected.
[0111] In some embodiments, the system detects a certain type of object (car, plane, etc.) and then goes on to refine the candidate set based on the category of the classified pixel.
[0112] In some embodiments, increasing the weighting of at least one further away depth value over closer depth values along the direction is not achieved by excluding one or more pixels along the direction from being included in the set of candidate depth values, hi some embodiments, all pixels along a first direction may be included from the current pixel to the further away pixel, hereafter referred to as the first candidate depth value and the first candidate pixel, respectively.
[0113] In such an embodiment, increasing the bias of the first candidate depth value relative to depth values belonging to an intervening set of candidate depth values with lower bias is achieved by appropriately designing the cost function such that the cost function is lower for the first candidate depth value than the cost function for one or more depth values closer to the first pixel.
[0114] In some embodiments, the cost function along the direction has a cost gradient that increases monotonically with distance from the first pixel to the candidate depth value where the cost function is evaluated for distances below a distance threshold, and a cost gradient that decreases with distance for at least one distance from the first pixel that is above the threshold. Thus, up to the distance threshold, the cost function increases or is constant with increasing distance to the first pixel. However, for at least one distance above the distance threshold, the cost function instead decreases.
[0115] For example, the distance threshold corresponds to the distance to the last pixel in the intervening set. Thus, the cost function increases (or remains constant) for pixels up to and including the farthest pixel in the intervening set. However, the cost gradient decreases with distance between this farthest pixel and the first candidate pixel. Thus, the cost function for the first candidate depth value will be lower than for at least one pixel along a direction closer to the current pixel.
[0116] A cost function for one pixel being smaller or lower than another pixel means that the resulting cost value determined by the cost function is smaller / lower when all other parameters are the same, i.e., when considering two pixels where all parameters except position are the same. Similarly, a cost function for one pixel being larger or higher than another pixel means that the resulting cost value determined by the cost function is larger / higher when considering two pixels where all other parameters are the same. Also, a cost function for one pixel exceeding another pixel means that the resulting cost value determined by the cost function is higher when considering two pixels where all other parameters are the same.
[0117] For example, a cost function typically takes into account several different parameters, e.g., a cost value is determined as C=f(d,a,b,c,...), where d refers to the position (e.g., distance) of the pixel relative to the current pixel, and a,b,c,... reflect other parameters that are taken into account, such as image signal values of related images, values of other depth values, smoothness parameters, etc.
[0118] The cost function f(d,a,b,c,...) for pixel A is lower than for pixel B if C=f(d,a,b,c,...) is lower for pixel A than for pixel B when the parameters a,b,c,... are the same for the two pixels (and similarly for the other terms).
[0119] An example of using a cost function to bias further away pixels is shown in Figure 5. This example corresponds to the example of Figure 4, but the set of candidate depth values further includes all pixels along the direction, i.e., pixels 501-507. However, in this example, the cost function is arranged to bias the first candidate depth value higher than several intervening pixels, specifically pixels 501-507.
[0120] An example of a possible cost function and its dependence on distance d from the first pixel is illustrated in Figure 6. In this example, the cost function is very low for neighboring pixel 401 and then increases with distance for pixel 501. However, for the first candidate pixel, the cost function decreases for pixel 501 but still remains higher for neighboring pixel 401. Thus, applying this cost function biases the first candidate depth value higher than the intervening pixel 501, but not as much as neighboring pixel 401.
[0121] Thus, in this approach, the set of candidate depth values includes a first candidate depth value along a first direction from a first pixel that is a greater distance from the current pixel than an intervening set of pixels. The intervening set of pixels includes at least one pixel, all of which are either not included in the set of candidate depth values or have a higher cost function than the first candidate depth value. Thus, the set of intervening pixels does not include a candidate depth value whose cost function does not exceed the cost function for the first candidate depth value. Thus, the set of candidate depth values includes a first candidate depth value along a first direction that is further away from the first pixel than at least one pixel along the first direction that is not included in the set of candidate depth values or has a higher cost function than the first candidate depth value.
[0122] In some embodiments, this corresponds directly to an intervening pixel set, which is a set of pixels whose depth values are not included in the set of candidate values. In some embodiments, this corresponds directly to an intervening set that includes at least one candidate pixel / depth value for which the cost function for that candidate pixel / depth value exceeds the cost function of the first candidate depth value. In some embodiments, this corresponds directly to a cost function along a direction that has a monotonically increasing cost gradient with respect to distance from the first pixel to the candidate depth value (for which the cost value is determined) if the distance is below a distance threshold, and a decreasing cost gradient with respect to distance if at least one distance from the first pixel is above a threshold.
[0123] The exact cost function depends on the particular embodiment. In many embodiments, the cost function includes a cost contribution that depends on the difference between image values in the multi-view image for pixels offset by a disparity corresponding to the depth value. As previously mentioned, a depth map may be a map for an image (or set of images) of a multi-view image set capturing a scene from different viewpoints. Thus, there is disparity between the location of the same object in different images, and this disparity depends on the depth of the object. Thus, for a given depth value, disparities between two or more images of the multi-view image can be calculated. Thus, in some embodiments, for a given candidate depth value, disparities to other images are determined, so that, assuming the depth value is correct, the location of the first pixel location in the other images can be determined. Image values, such as color or brightness values, for one or more pixels at each location can be compared, and a suitable difference measure can be determined. If the depth value is indeed correct, it is more likely that the image values are the same and the difference measure is smaller than if the depth value is not the correct depth. The cost function may therefore include consideration of the difference between image values; in particular, the cost function reflects increasing costs for increasing difference measures.
[0124] By including such a matching criterion, this approach can be used as an integral component of disparity-based depth estimation between multi-view images. A depth map is initialized and then processed iteratively, with updates biased toward smaller image value differences. In this way, this approach effectively performs joint depth determination and search / matching between different images.
[0125] In this approach, depth values further away in a first direction are therefore biased / weighted higher than at least one depth value that is closer to the first pixel along the first direction. In some embodiments, this may be true for multiple directions, but in many embodiments, it is (only) true for one direction or directions within a small interval, e.g., an interval of 1°, 2°, 3°, 5°, 10°, or 15°. Equivalently, directions are considered to have degrees of angular interval of 1°, 2°, 3°, 5°, 10°, or 15° or less.
[0126] Therefore, in many embodiments, the approach is that for any candidate depth value along a second direction from a first pixel, all depth values along the second direction that are a shorter distance to the first pixel belong to the set of candidate depth values, and the cost function along the second direction increases monotonically with distance (all other parameters being equal).
[0127] For most embodiments, the set of candidate depth values in the second direction from the first pixel does not include a second candidate depth value for which the pixel set of at least one pixel along the second direction does not include a candidate depth value in the set of candidate depth values for which the cost function does not exceed the cost function for the second candidate depth value, and the distance from the first pixel to the second candidate depth value is greater than the distance from the first pixel to the pixel set.
[0128] In practice, typically only one direction contains candidate depth values that are not included in the set of candidate depth values, or that are included but are further away than one or more pixels with a higher cost function.
[0129] Considering further apart depth values / pixels restricted to one direction (including potentially small angular intervals) allows particularly advantageous performance in many embodiments, as the approach can be adapted to specific characteristics of the scene that constrain consideration of further apart pixels to situations where the further apart pixels are particularly likely to reflect the correct depth.
[0130] In particular, in many embodiments, the first direction corresponds to the direction of gravity in the depth map / image / scene. The inventors have recognized that by considering pixels further away along the direction of gravity, advantageous operation is achieved as the likelihood that such depth values are correct is greatly increased.
[0131] In particular, the inventors recognized that in many practical scenes, objects are positioned or standing on the ground, and in such scenes, the depth of the entire object is typically comparable to the depth of the portion of the object that is furthest away in the direction of gravity. The inventors further recognized that this typically translates into a corresponding relationship in the depth map, where depth values for an object are often more similar to depth values in the direction of gravity in the depth map than depth values in the local neighborhood. For example, the head of a person standing on a flat surface will have a depth approximately equal to the depth of the feet. However, the depth of the neighborhood around the head will be significantly different because it includes pixels corresponding to the distant background. Therefore, depth processing based solely on the depth of the neighborhood may be less reliable and accurate for the head than for the feet. However, with the described approach, the method can include not only the neighborhood, but also depth values and pixels further away in the direction of gravity. For example, when processing depth values for a person's head, the described approach results in one candidate depth value being the depth value from the person's feet (which more accurately reflects the true depth, especially after several iterations).
[0132] 7 illustrates an example of the improvement that can be achieved. The figure illustrates an image, the corresponding depth map after application of a corresponding process that does not consider the far gravity direction candidates, and finally, the depth map after application of a process that does not consider the far gravity direction candidates. Comparing portions 701 and 703, it can be seen that the head of the player on the far left has an incorrect depth value in the first example, but the correct depth value in the second example.
[0133] In some embodiments, the direction of gravity in the depth map may be predetermined, and the first direction may be predetermined. In particular, for many typical depth maps and images, horizontal capture is performed (or post-processing is performed to align the horizontal direction of the image with the scene), and the direction may be predetermined as the vertical direction of the depth map / image. In particular, the direction may be the up-down direction in the image.
[0134] In some embodiments, the processor 203 is arranged to determine a first direction as the direction of gravity in the depth map, the direction of gravity in the depth map being a direction corresponding to the direction of gravity in the scene represented by the depth map.
[0135] In some embodiments, such a determination may be based on an evaluation of input, such as from a level indicator of a camera capturing the image for which the depth map is to be updated. For example, if data is received indicating that the stereoscopic camera is at an angle of, say, 30° with respect to the horizontal, the first direction may be determined as a direction offset by 30° with respect to the vertical in the depth map and image.
[0136] In many embodiments, the direction of gravity in the depth map may be based on an analysis of the depth map and / or the image. For example, the direction of gravity may be chosen to be opposite to the vector pointing from the center of the image to the pixel location in the average weighted image, with a per-pixel weighting proportional to the amount of blue in the pixel's color. This is a simple way to determine the direction of gravity using a blue sky in a photograph. Approaches are known that rectify stereo image pairs (or multi-view images) so that the so-called epipolar lines are horizontal. Gravity may always be assumed to be orthogonal to the epipolar lines.
[0137] In some embodiments, the cost function may include consideration of a depth model for at least a portion of the scene. In such embodiments, the processor 203 is arranged to evaluate the depth model to determine an expected depth model. The cost function then depends on the difference between the depth value and a model depth value determined from the depth model.
[0138] The depth model may be a model that imposes depth constraints on at least some depth values of a depth map, and the depth constraints may be absolute or relative. For example, the depth model may be a 3D model of a scene object that, when projected onto the depth map, results in a corresponding depth relationship between depth values for pixels corresponding to the image object. In this way, the absolute depth of the scene object is not known, but a depth relationship can be implied if it is known what type of object is represented by the scene object. As another example, the depth model may be a disparity model estimated relative to a static ground plane and / or a static background, or may be a disparity model of, for example, a set of dynamically moving planar or cylindrical objects (e.g., representing athletes on a stadium).
[0139] The cost function thus evaluates the model to determine an expected depth value for a candidate depth value according to the depth model, then compares the actual depth value with the expected value and determines a cost contribution that increases monotonically as the difference increases (at least for some depth values).
[0140] In some embodiments, the cost contribution may be asymmetric, so that it differs depending on whether the depth value is higher or lower than the expected value. For example, a different function may be applied so that depth values further away from the model result in a significantly higher cost contribution than depth values closer to the model, thereby biasing the update toward depths further forward than the model. Such an approach is particularly advantageous when the model is a background model that provides an indication / estimation of background depth. In such cases, the cost contribution makes it less likely that the depth map will be updated to reflect the depth, resulting in perceptually significant artifacts / errors where objects are rendered further behind the depth background.
[0141] Indeed, in some cases, the cost contribution for a depth value indicating a depth higher than the background depth is so high that this depth value is very unlikely to be selected; for example, the cost contribution from model comparison is set to a very high value (in principle even to infinity) in such cases.
[0142] As an example, the cost contribution to the model evaluation is given by:
number
[0143] In some embodiments, only candidate depth values from the depth map itself are considered for the set of candidate depth values. However, in some embodiments, the set of candidate depth values is generated to include other candidate depth values. As previously mentioned, the set of candidate depth values includes one or more depth values from another depth map in the temporal sequence of depth maps that includes the depth map. Specifically, depth values from depth maps of previous and / or subsequent frames in the video sequence are included. In some embodiments, the set of candidate depth values may include depth values determined according to an offset of the depth value for the first pixel. For example, the set of candidate depth values may include depth values generated by adding a predetermined offset to the depth value for the current pixel and / or depth values generated by subtracting a predetermined offset from the depth value for the current pixel.
[0144] Including different types of depth values provides improved performance in many applications and scenarios, and in particular often allows for larger updates with fewer constraints. Furthermore, different types of depth values may be included by designing a cost function that considers different types of potential probabilities of indicating the correct depth for the current pixel. In particular, the cost function may depend on the type of depth value and therefore take into account the type of depth value to which the cost function is applied. More specifically, the cost function considers whether the depth value is a depth value from the depth map, a depth value from the depth map that is closer than a distance threshold (e.g., in the immediate vicinity), a depth value from the depth map that is farther away than a distance threshold (e.g., pixels further away along the gravity direction), a depth value from another depth map in the temporal sequence of depth maps that includes the depth map, a depth value that is independent of the scene represented by the depth map, or a depth value determined according to a depth value offset relative to the first pixel. Of course, in many embodiments, only a subset of these is considered.
[0145] As an example, the following cost function may be evaluated for each candidate depth value in the set of candidate depth values: Ctotal =w1C match +w2C smoothness +w3C model +w4C candidate where C match is a cost that depends on the match error between the current view and one or more other views, and C smoothness The cost component C weights both spatial smoothness and penalizes depth transitions within regions of constant color intensity. Several different approaches to determining such cost values / contributions are known to those skilled in the art, and for the sake of brevity, these will not be described further. model is the stated model cost contribution, which may reflect the deviation of the disparity from a prior known or estimated disparity model. candidate may introduce a cost contribution that depends on the type of depth values, for example, whether they are from the same depth map, such as temporally adjacent depth maps.
[0146] As an example, C candidate is given by:
number
[0147] The cost of local neighborhood candidates is usually small, since such neighborhoods are very likely to be good predictors. The same applies to temporal neighbor candidates, but the cost is a bit higher to avoid errors for fast-moving objects. The cost of offset updates must be higher to avoid introducing noise. Finally, the cost of far-away (gravitational) candidates is usually higher than the cost of regular local neighborhood candidates, since the spatial distance is large. Multiple such candidates at different low positions (different values for Δ) may be used. In this case, the cost can be increased as a function of increasing distance Δ from the pixel being processed.
[0148] It will be appreciated that, for clarity, the above description has described embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without detracting from the invention. For example, functionality illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Hence, references to specific functional units or circuits should be seen only as references to suitable means for providing the described functionality, rather than indicative of a strict logical or physical structure or organization.
[0149] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units or as part of other functional units. As such, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
[0150] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined according to the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0151] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, these may conceivably be advantageously combined, and their inclusion in different claims does not imply that such combinations are not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must be performed, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, singular references do not exclude a plurality. Thus, references to "a," "first," "second," etc. do not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and are not to be construed as limiting the scope of the claims in any way.
Claims
1. 1. A method of operating an apparatus for processing a depth map, the method comprising: receiving the depth map with a receiver of the device; For at least a first pixel of the depth map: a processor of the device determining a set of candidate depth values, the set of candidate depth values including depth values for other pixels of the depth map other than the first pixel; the processor determining a cost value for each of the candidate depth values in the set of candidate depth values in response to a cost function; the processor selecting a first depth value from the set of candidate depth values according to the cost value for the set of candidate depth values; the processor determining an updated depth value for the first pixel in response to the first depth value; The set of candidate depth values includes a first candidate depth value along a first direction from the first pixel, and a first intervening pixel set of at least one pixel along the first direction does not include a candidate depth value from the set of candidate depth values for which the cost function for the candidate depth value does not exceed the cost function for the first candidate depth value, and the distance from the first pixel to the first candidate depth value is greater than the distance from the first pixel to the first intervening pixel set, and the method of operating the device further includes a step in which the processor determines the first direction as a gravity direction of the depth map, and the gravity direction is a direction in the depth map that matches the direction of gravity in a scene represented by the depth map.
2. 2. The method of claim 1, wherein the cost function along the first direction has a monotonically increasing cost gradient as a function of distance from the first pixel when the distance is below a distance threshold, and a decreasing cost gradient as a function of distance from the first pixel when at least one distance from the first pixel is above a threshold.
3. The method of claim 1 , wherein the first set of intervening pixels is a set of pixels whose depth values are not included in the set of candidate depth values.
4. 4. A method of operating an apparatus according to claim 1, wherein the cost function includes a cost contribution that depends on the difference between image values of multi-view images for pixels offset by a disparity that matches the candidate depth value to which the cost function is applied.
5. The method of claim 1 , wherein the first direction is a vertical direction in the depth map.
6. A method of operating an apparatus as described in any one of claims 1 to 5, wherein the processor further comprises a step of determining a depth model for at least a portion of the scene represented by the depth map, and the cost function for a depth value depends on the difference between the depth value and a model depth value determined from the depth model.
7. The method of claim 6 , wherein the cost function is asymmetric with respect to whether the depth value is above or below the model depth value.
8. A method according to claim 6 or 7, wherein the depth model is a background model for the scene.
9. The method of claim 8, further comprising: including candidate depth values in the set of candidate depth values that are not from the depth map; the processor receives depth values from another depth map of a temporal sequence of depth maps, the temporal sequence including the depth map; a scene-independent depth value represented by the depth map; and a depth value determined according to a depth value offset for the first pixel; 9. A method of operating an apparatus according to claim 1, further comprising the step of:
10. The cost function for the depth value depends on the type of the depth value, the type being: depth values of the depth map; depth values in the depth map that are closer than a distance threshold; depth values in the depth map that are farther apart than a distance threshold; a depth value from another depth map of a temporal sequence of depth maps that includes the depth map; a depth value having a scene independent depth value offset with respect to the depth value of the first depth value; a scene-independent depth value represented by the depth map; and a depth value determined according to a depth value offset for the first pixel; 10. A method of operating a device according to any one of claims 1 to 9, wherein the method is one of a group of types comprising at least one of:
11. A method of operating an apparatus described in any one of claims 1 to 10, wherein the processor processes multiple pixels of the depth map by iteratively selecting a new first pixel from a plurality of pixels and performing each step for each new first pixel.
12. 2. The method of claim 1, wherein the set of candidate depth values in a second direction from the first pixel does not include a second candidate depth value for which a pixel set of at least one pixel along the second direction for a second candidate depth value does not include the candidate depth value among the set of candidate depth values for which a cost function for the candidate depth value does not exceed the cost function for the second candidate depth value, and the distance from the first pixel to the second candidate depth value is greater than the distance from the first pixel to the pixel set.
13. 1. An apparatus for processing a depth map, the apparatus comprising: a receiver for receiving a depth map; a processor for processing the depth map, the processing comprising: For at least a first pixel of the depth map: determining a set of candidate depth values, the set of candidate depth values including depth values for other pixels of the depth map other than the first pixel; determining a cost value for each of the candidate depth values in the set of candidate depth values in response to a cost function; selecting a first depth value from the set of candidate depth values according to the cost value for the set of candidate depth values; determining an updated depth value for the first pixel in response to the first depth value; the set of candidate depth values includes a first candidate depth value along a first direction from the first pixel, a first intervening pixel set of at least one pixel along the first direction does not include a candidate depth value in the set of candidate depth values for which the cost function for the candidate depth value does not exceed the cost function for the first candidate depth value, a distance from the first pixel to the first candidate depth value is greater than a distance from the first pixel to the first intervening pixel set, and the processing further includes determining the first direction as a gravity direction of the depth map, the gravity direction being a direction in the depth map that matches a direction of gravity in a scene represented by the depth map.
14. A computer program comprising computer program code means adapted to perform all the steps of the apparatus according to claim 13 when said computer program is run on a computer.
Citation Information
Patent Citations
Apparatus and method for processing depth maps
JP2020518058A
Processing of depth maps for images
WO2020178289A1