Generate depth maps for images
Patent Information
- Application Number
- JP2025514737
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-21
- Filing Date
- 2023-09-19
- Publication Date
- 2026-09-08
AI Technical Summary
Existing depth map generation methods suffer from inaccuracies and errors, particularly near depth transitions and due to misalignment between depth sensors and optical cameras, leading to suboptimal view synthesis and rendering in immersive video applications.
An apparatus and method for generating a depth map by combining depth maps from multiple locations, designating uncertain pixels, and using image values to determine depth values for uncertain pixels, ensuring accurate alignment and reducing errors.
This approach improves depth map accuracy, especially around depth transitions, enabling efficient and high-quality view shifting and synthesis, reducing computational complexity, and enhancing immersive video experiences.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to depth map generation, particularly but not exclusively to depth map generation that supports view synthesis / shifting based on captured images of a scene. [Background technology]
[0002] The variety and range of image and video applications has increased dramatically in recent years, and new services and methods for using and consuming video are continually being developed and introduced.
[0003] For example, one service that is gaining popularity is the presentation of image sequences in such a way that the viewer can actively interact with the system to change the parameters of the rendering.A very attractive feature in many applications is the ability to change the effective viewing position and direction of the viewer, for example, allowing the viewer to move and look around the displayed scene.
[0004] An example of a proposed video service or application is immersive video, where a video is played, for example, on a VR headset, to provide a three-dimensional experience. With immersive video, a viewer has the freedom to view and move around a displayed scene, which can be perceived as being seen from different viewpoints. However, in many typical approaches, the amount of movement is limited to a relatively small region around a nominal viewpoint, which typically corresponds, for example, to the viewpoint from which video capture of the scene was performed. In such applications, three-dimensional scene information is often provided that enables high-quality viewpoint image synthesis for viewpoints relatively close to the reference viewpoint, but degrades when the viewpoint deviates too much from the reference viewpoint.
[0005] Immersive video is often referred to as six degrees of freedom (6DoF) or 3DoF+ video. MPEG Immersive Video (MIV) is a new standard in which metadata to enable and standardize immersive video is used on top of existing video codecs.
[0006] In many applications, the capture of a real-world scene is based on multiple spatially distinct cameras, such as cameras arranged in a line, which provide images of the scene from different view positions. Additionally, the captured images can be provided along with a depth map that provides information about the distance from the camera to objects in the scene being captured. For example, immersive video data may be provided in the form of a multi-view, possibly accompanied by a depth data (MVD) representation of the scene.
[0007] In some applications, depth information can be estimated based on information in the captured images, for example, by estimating disparity values of corresponding objects in different images. However, because such estimations tend to be suboptimal and can result in inaccuracies and errors, many capture systems include dedicated depth sensors that can capture distances to objects. For example, depth sensors based on indirect time of flight (iToF) are popular, and many cameras that combine optical and depth sensors have been developed. Many capture systems can use several such cameras to provide sets of images and depths for different view poses.
[0008] However, while the inclusion of one or more depth sensors in a capture device tends to add robustness and more reliable depth information, the depth sensors may also suffer from errors or inaccuracies. Most notably, accurate depth measurements near depth transitions (steps) tend to be difficult due to ambiguous / mixed signals received by the depth sensors. Note that similar problems are known in stereo or multi-view depth estimation based on image matching (correlation) due to the inability of pixels to match due to occlusion.
[0009] Another problem is that due to practical implementation constraints, the depth sensor and the light sensor are generally physically offset relative to each other. Therefore, the measured depth information is for an offset pose relative to the captured image, and therefore the image and the depth map do not perfectly match. This can introduce errors and inaccuracies, for example, when performing view shifting and rendering based on the depth map. Summary of the Invention [Problem to be solved by the invention]
[0010] Therefore, improved approaches would be advantageous, particularly approaches that allow for improved operation, increased flexibility, reduced complexity, easier implementation, improved synthesized image quality, improved rendering, improved depth information, improved consistency between depth maps and linked images, improved and / or easier view synthesis for different view poses, reduced data requirements, and / or improved performance and / or operation. [Means for solving the problem]
[0011] Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination.
[0012] According to one aspect of the present invention, there is provided an apparatus for generating a depth map for an image representing a view of a scene, the apparatus comprising: a position receiver configured to receive a capture location of an image, the capture location being a location of capture of the image; a depth receiver configured to receive a first depth map providing depth values from a first location and a second depth map providing depth values from a second location, the first depth map including at least some pixels designated as uncertain depth pixels and the second depth map including at least some pixels designated as uncertain depth pixels; a view shift processor configured to perform a first view shift of the first depth map from a first position to a capture position to generate a first view-shift depth map, and to perform a second view shift of the second depth map from the second position to the capture position to generate a second view-shift depth map, the view shift processor further configured to designate as uncertain pixels of the first view-shift depth map to which the first view shift does not project depth values of pixels not designated as uncertain pixels in the first depth map, and pixels of the second view-shift depth map to which the second view shift does not project depth values of pixels not designated as uncertain pixels in the second depth map; a combiner configured to generate a combined depth map for the capture location by combining depth values of co-located pixels of the first view-shifting depth map and the second view-shifting depth map, the combiner configured to designate a pixel of the combined depth map as an uncertain depth pixel if any of the co-located pixels is designated as an uncertain pixel; and a depth map generator configured to generate an output depth map for the capture location by determining depth values in the combined depth map for pixels designated as uncertain pixels using at least one of image values of the image and depth values of pixels in the combined depth map that are not designated as uncertain pixels.
[0013] The approach allows for an improved depth map to be generated for an image based on depth maps from depth sensors at other locations. This approach allows for improved image synthesis for different view poses. In many embodiments, depth data is generated that allows for efficient, high-quality view shifting and synthesis.
[0014] The terms pose or configuration are commonly used in the art to refer to position and / or orientation.
[0015] This approach allows for efficient combination of depth information from different depth sensors and depth maps. In particular, this approach can prevent artifacts, errors, and inaccuracies that result from using incorrect or inaccurate depth values in many scenarios. This can provide improved consistency in many scenarios.
[0016] This approach can allow for efficient and easy implementation, and in many scenarios can offer reduced computational resource usage and complexity.
[0017] The position can be a scene position. The position can be referenced to a scene coordinate system. The same position of pixels (representing depth from the same position) in different depth maps are pixels that represent the same view direction from a given position.
[0018] Uncertain pixels are often referred to as unknown, unreliable, invalid and / or untrusted pixels.
[0019] In many embodiments, the device may include a view combiner configured to generate, from the images and the output depth map, view images for view poses having positions different from the capture position.
[0020] In accordance with an optional feature of the invention, the angle between the direction from the capture position to the first position and the direction from the capture position to the second position is greater than or equal to 90°.
[0021] This can provide particularly advantageous operation and / or implementation in many embodiments, in which the angle can be 100°, 120°, 140°, or 160° or greater.
[0022] In some embodiments, the angle between the line connecting the capture location to the first location and the line connecting the capture location to the second location is 90°, 100°, 120°, 140°, 160° or more.
[0023] In some embodiments, the angle between a line passing through the capture position and the first position and a line passing through the capture position and the second position is 90°, 100°, 120°, 140°, 160° or more.
[0024] In accordance with an optional feature of the invention, the first location, the second location and the capture location are arranged in a linear configuration with the capture location being between the first location and the second location.
[0025] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0026] According to an optional feature of the invention, the first location, the second location and the capture location are arranged such that the movement directions of pixels at the same depth are in opposite directions for the first view shift and the second view shift.
[0027] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0028] In some embodiments, the first position, the second position and the capture position are arranged such that the direction of movement of the image coordinates for a given depth is in opposite directions relative to the first view shift and the second view shift.
[0029] According to an optional feature of the invention, the combiner is configured to determine a depth value of the combined depth map for a given pixel as a weighted combination of a first depth value for a pixel in a first view shift map that is not designated as an uncertain depth pixel and has the same location as the given pixel, and a second depth value for a pixel in a second view shift map that is not designated as an uncertain depth pixel and has the same location as the given pixel.
[0030] This can provide particularly advantageous operation and / or implementation in many embodiments, which can typically provide depth noise or error suppression.
[0031] In some embodiments, the combiner is configured to determine a depth value of the combined depth map for a given view direction as a weighted combination of a first depth value for the given view direction in the first view shift map and a second depth value for the given view direction in the second view shift map, where the first depth value is for a pixel in the first depth map that is not designated as an uncertain depth pixel and the second depth value is for a pixel in the second depth map that is not designated as an uncertain depth pixel.
[0032] According to an optional feature of the invention, the depth map generator is configured to generate depth values for pixels designated as uncertain in the combined depth map by estimating them from depth values of pixels not designated as uncertain in the combined depth map.
[0033] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0034] According to one aspect of the present invention, there is provided an image capture system comprising the aforementioned apparatus, and further comprising a first image camera at a capture location configured to provide an image to an image receiver, a first depth sensor at a first location configured to provide depth data for a first depth map to the depth receiver, and a second depth sensor at a second location configured to provide depth data for a second depth map to the depth receiver.
[0035] This approach can provide an image capture system that can generate particularly accurate and / or reliable 3D information, including both image and depth information.
[0036] According to an optional feature of the invention, the capture system further comprises at least one further image camera at a further capture location, wherein an angle between a direction from the further capture location to the first location and a direction from the further capture location to the second location is greater than or equal to 90°.
[0037] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0038] The approach described with reference to the first camera can also be applied to each of the at least one further image camera.
[0039] In many embodiments, the angle may be 100°, 120°, 140°, or 160° or more.
[0040] According to an optional feature of the invention, the capture system further comprises a plurality of image cameras including the first image camera and a plurality of depth sensors including the first depth sensor and the second depth sensor, wherein for every image camera of the plurality of image cameras, two depth sensors of the plurality of depth sensors are positioned such that an angle between directions from the position of the image camera to the two depth sensors is 120° or greater. This can provide particularly advantageous operation and / or implementation in many embodiments. In many embodiments, this angle can be 140° or 160° or greater.
[0041] In accordance with an optional feature of the invention, the plurality of image cameras and the plurality of depth sensors are arranged in a line.
[0042] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0043] In accordance with an optional feature of the invention, at least one depth sensor of the plurality of depth sensors is positioned between each pair of adjacent imaging cameras of the plurality of imaging cameras.
[0044] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0045] In accordance with an optional feature of the invention, the plurality of image cameras and the plurality of depth sensors are arranged in a two-dimensional configuration.
[0046] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0047] In accordance with an optional feature of the invention, the number of depth sensors in the plurality of depth sensors exceeds the number of imaging cameras in the plurality of imaging cameras.
[0048] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0049] In accordance with an optional feature of the invention, the number of imaging cameras in the plurality of imaging cameras exceeds the number of depth sensors in the plurality of depth sensors.
[0050] This can provide particularly advantageous operation and / or implementation in many embodiments.
[0051] According to one aspect of the present invention, there is provided a method for generating a depth map for an image representing a view of a scene, the method comprising: receiving a capture location of an image, the capture location being a location of capture of the image; receiving a first depth map providing depth values from a first location and a second depth map providing depth values from a second location, the first depth map including at least some pixels designated as uncertain depth pixels and the second depth map including at least some pixels designated as uncertain depth pixels; performing a first view shift of the first depth map from the first position to the capture position to generate a first view-shifted depth map; performing a second view-shift of the second depth map from the second position to the capture position to generate a second view-shift depth map; designating as uncertain pixels of the first view-shift depth map to which the depth values of pixels not designated as uncertain pixels in the first depth map are not projected due to the first view shift and pixels of the second view-shift depth map to which the depth values of pixels not designated as uncertain pixels in the second depth map are not projected due to the second view shift; generating a combined depth map for the capture location by combining depth values of co-located pixels of the first view-shift depth map and the second view-shift depth map, and designating a pixel of the combined depth map as an uncertain depth pixel if the co-located pixel is designated as an uncertain pixel; generating an output depth map for the capture location by determining depth values in the combined depth map for pixels designated as uncertain pixels using at least one of image values of the image and depth values of pixels not designated as uncertain.
[0052] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0053] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Figure 1] FIG. 1 is a diagram showing an example of elements of an image distribution system. [Figure 2] FIG. 1 illustrates an example image capture scenario. [Figure 3] 1 illustrates example elements of an apparatus according to some embodiments of the present invention. [Figure 4] FIG. 1 illustrates an example of a scene capture scenario. [Figure 5] FIG. 1 illustrates an example of a scene capture scenario. [Figure 6] FIG. 1 illustrates an example of a depth map for a scene capture scenario. [Figure 7] FIG. 1 illustrates an example of processing a depth map for a scene capture scenario. [Figure 8] A diagram showing example elements of a depth and image capture scenario. [Figure 9] A diagram showing example elements of a depth and image capture scenario. [Figure 10]A diagram showing example elements of a depth and image capture scenario. [Figure 11] A diagram showing example elements of a depth and image capture scenario. [Figure 12] A diagram showing example elements of a depth and image capture scenario. [Figure 13] FIG. 1 illustrates some elements of a possible configuration of a processor for implementing elements of an apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] The capture, distribution, and presentation of three-dimensional images and videos are becoming increasingly prevalent and desirable in several applications and services. A particular approach, known as immersive video, typically involves providing views of real-world scenes, and often real-time events, while allowing for small observer movements, such as relatively small head movements and rotations. For example, real-time video broadcasts of sporting events can provide a user with the impression of sitting in the stands watching the sporting event. The user can, for example, look around and have a natural experience similar to that of a spectator present in that position within the stands. In recent years, display devices with positional tracking and 3D interaction that support applications based on 3D capture of real-world scenes have become increasingly popular. Such display devices are well suited for immersive video applications, providing an enhanced three-dimensional user experience.
[0055] Although the following description focuses on immersive video applications, it will be understood that the principles and concepts described can be used in many other applications and embodiments, including embodiments in which a single image is captured.
[0056] In many approaches, immersive video can be provided locally to the viewer, for example, by a standalone device that does not use or even have access to a remote video server. However, in other applications, the immersive application can be based on data received from a remote or central server. For example, video data can be provided to a video rendering device from a remote central server and processed locally to generate the desired immersive video experience.
[0057] 1 shows an example of an immersive video system in which a video rendering client 101 cooperates with a remote immersive video server 103 over a network 105, such as the Internet. The server 103 may be configured to potentially support a large number of client devices 101 simultaneously.
[0058] Server 103 may support immersive video experiences, for example, by transmitting three-dimensional video data describing a real-world scene. The data may specifically describe the visual features and geometric properties of the scene generated from real-time capture of the real world by a set of (possibly three-dimensional) cameras.
[0059] To provide such services for real-world scenes, the scene is typically captured from different positions, using different camera capture poses. As a result, multi-camera capture and, for example, 6DoF (six degrees of freedom) processing are rapidly gaining in relevance and importance. Applications include live concerts, live sports, and telepresence. The freedom to choose one's own viewpoint enriches these applications by enhancing the sense of presence beyond that of regular video. Furthermore, immersive scenarios can be envisioned in which observers can move through and interact with the live-captured scene. For broadcast applications, this would require real-time view synthesis on the client device. View synthesis introduces errors, and these errors vary depending on the implementation details of the algorithm.
[0060] To provide an adequate representation of the scene, the captured image / video is typically provided with a depth map that indicates the depth / distance of objects represented by the pixels of the image. In many embodiments, the depth map linked to the image contains one depth value and depth pixel for each pixel of the image, and therefore the depth map and the image have the same resolution. In other embodiments, the depth map may have a lower resolution, and the depth map may provide some depth values that are common to multiple image pixels.
[0061] In some capture systems, a depth map can be estimated or derived from the captured image. However, improved depth information is typically achievable by using a dedicated depth sensor. For example, depth sensors have been developed that transmit a signal and estimate the time delay of the echo arriving from each direction. A depth map can then be generated using depth values reflecting this measured time-of-flight in each direction. 3D cameras have been developed that include both an optical camera and one or more depth sensors. Multiple such cameras are often used together in an appropriate capture configuration. For example, multiple 3D cameras, along with associated depth cameras, can be arranged in a line to provide MVD representations of the image for different capture poses.
[0062] However, while such approaches can provide excellent results in many scenarios, they do not perform optimally in all situations. In particular, in some cases, it may be difficult for the depth sensor to accurately determine the time of flight and therefore the distance to the object. In general, depth estimation is often difficult or inaccurate around depth transitions, such as the edges of objects.
[0063] In many applications, it has been found to be advantageous to separately treat depth values that are considered reliable and depth values that are considered not sufficiently reliable. Accordingly, in many capture systems, the generated depth map can be provided with data indicating whether the pixel / depth value is considered an uncertain pixel or a certain pixel. For example, for each pixel and depth value of the depth map, a binary map can be provided that indicates whether the corresponding pixel of the depth map is an uncertain depth pixel, i.e., whether the depth value provided by the depth map for that pixel is designated as an uncertain pixel.
[0064] It will be understood that any suitable approach or algorithm for determining which depth and depth pixels (pixels of the depth map) are designated as uncertain pixels and which are not can be used, and different approaches can be used in different embodiments. Typically, a depth sensor can be equipped with built-in functionality to generate a confidence estimate or confidence level for each depth value, and if the confidence estimate or confidence level indicates a confidence level below a given threshold, the corresponding pixel is designated as uncertain. For example, if echoes in a given direction are spread over a relatively long time interval, or if, for example, two or more significant peaks are detected, the confidence is considered low by the depth sensor, and the pixel is designated as an uncertain pixel.
[0065] It will be understood that uncertain pixels correspond to unreliable / uncertain pixels or unknown depth pixels, and in many embodiments, this term is replaced by any such term. Depth pixels or values that are not designated as uncertain / unreliable / unknown depth pixels may be referred to as certain / reliable / known depths.
[0066] Thus, the output from the depth sensor may be a depth map that includes multiple pixels that are considered to be uncertain, specifically unknown depth values and pixels. Processing such a depth map involves addressing such pixels and generating depth values for the uncertain and unknown pixels using other means. For example, the depth values of the uncertain pixels can be determined by disparity estimation between different corresponding pixels in different images for different views. Alternatively or additionally, the depth values of the uncertain depth pixels can be determined by extrapolation of depth pixels that are considered to be certain or known depth pixels.
[0067] Another problem with many capture devices is that the depth sensor does not perfectly align with the optical camera sensor, i.e., the depth data in the depth map is provided relative to a different position than the capture position of the image, i.e., the depth data is provided relative to a different position than the position of the image. Although capture devices are designed to try to reduce such position offsets as much as possible, they can still introduce inaccuracies and errors, for example, when performing view shifting and image synthesis based on the image and depth information.
[0068] Below, we describe an approach and apparatus that can generate improved depth maps for many scenarios and many depth maps, particularly by generating a depth map for a capture location by considering depth values of at least two depth maps from different locations, warping these depth maps to the capture location, and using a specific approach for combining the shifted depth maps, where the combining includes specific consideration of whether pixels are designated as certain or uncertain.
[0069] This approach will be described with reference to a capture device as shown in Figure 2. In this example, an optical camera is placed at an image capture location 201, a first depth sensor is placed at a first depth sensing location 203, and a second depth sensor is placed at a second depth sensing location 205. The three locations are positioned in a straight line with the first depth sensing location 203 and the second depth sensing location 205 on either side of the capture location 201.
[0070] In some embodiments, the optical camera and the depth sensor can each be, for example, separate devices / items. In other embodiments, the camera and the depth sensor can be implemented, for example, within the same device, e.g., they can all be implemented within a single 3D camera device. In other embodiments, the depth sensor can be part of, for example, a different device or camera. For example, multiple 3D cameras can each include one optical camera and one depth sensor positioned adjacent to each other. Positioning these cameras next to each other results in an arrangement such as that shown in FIG. 2, where the two depth sensors are depth sensors of adjacent 3D cameras. Thus, in such a scenario, each camera can provide an image for capture position 201 and depth maps for depth sensing positions 201 and 203. Thus, depth maps from the two devices can be used to provide depth maps for first depth sensing position 203 and second depth sensing position 205.
[0071] Elements of an apparatus for generating a depth map for a capture location are shown in Figure 3. The operation of the apparatus will be described with particular reference to the configuration of Figure 2, although it will be understood that the principles and approaches described can be used with many different capture configurations.
[0072] The device of Figure 3 comprises a position receiver 301 configured to receive an image capture position, where the capture position is a (scene) location for the capture of an image. The position receiver 301 may be specifically coupled to one or more cameras and receive position indications therefrom (e.g., the cameras may include GPS functionality for determining position and provide this to the position receiver 301). In other embodiments, the position receiver 301 may receive the capture position from a separate device, dedicated position measurement, or other means including, for example, direct user input of the capture position.
[0073] In a particular example, the position receiver 301 is coupled to the camera in the capture configuration of FIG. 2 and receives data describing the capture position 201 directly from the camera.
[0074] The apparatus further includes a depth receiver 303 configured to receive a depth map representing a scene. Specifically, the depth receiver 303 can be configured to receive the depth map from a depth sensor of the capture configuration. The depth receiver 303 can also receive location information indicating the location at which the depth map was captured. Specifically, the depth receiver 303 can receive a first depth map providing depth values from a first location (also referred to as a first depth sensing location), such as a depth map from the sensor of FIG. 2 at the first depth sensing location 203. The depth receiver 303 can further receive a second depth map providing depth values from a second location (also referred to as a second depth sensing location), such as a depth map from the sensor of FIG. 2 at the second depth sensing location 205. The following description focuses on an embodiment in which depth maps are received from only two locations, but it will be understood that depth maps from more locations can be received in other embodiments.
[0075] The depth receiver 303 also receives data indicating whether pixels of the first and second depth maps are designated as uncertain or certain pixels. The indication of whether a pixel is considered uncertain reflects a degree of certainty / reliability / certainty / plausibility (or existence) that the depth value of that pixel is correct / an accurate indication of the depth of the object represented by the pixel.
[0076] In some embodiments, the depth receiver 303 may receive data comprising, for example, a list or other indicator of all pixels in the first and / or second depth maps that are considered uncertain, or an indicator of all pixels that are considered certain. In other embodiments, the depth receiver 303 may receive, for example, a confidence map for each depth map, where each pixel in the confidence map indicates a confidence level for the depth value of that pixel in the depth map. In some embodiments, an indicator of whether a pixel is certain may be included in the depth map itself. For example, a depth value may be assigned to each pixel, where, for example, odd values indicate certain values and even values indicate uncertain pixels. As another example, uncertain values may be indicated by a particular depth value, such as, for example, a depth value of zero that is reserved to indicate an uncertain depth value.
[0077] It will also be understood that although the uncertain / certain pixel indicators can be provided by binary values, in other embodiments they can be provided as non-binary values. For example, the confidence value indicator can be indicated by a range that includes multiple possible values. For example, the confidence value can be given by an 8-bit value. In such a case, the pixel can be designated as uncertain if the confidence value meets the uncertainty criterion, and can be designated as certain otherwise. Specifically, if the confidence value does not exceed a threshold, the pixel is designated as uncertain, and otherwise the pixel is designated as certain.
[0078] It will be appreciated that any suitable approach and criteria for designating depth values / depth value pixels as certain / uncertain may be used depending on the particular preferences and requirements of individual embodiments.
[0079] The depth receiver 303 and the position receiver 301 are coupled to a view shift processor 305 configured to perform a view shift of the depth map. The view shift of the depth map shifts the depth map from representing depth from a depth sensing position to representing depth from a capture position of the image. Thus, the view shift processor 305 is configured to view shift the first depth map from the first depth sensing position 203 to the capture position 201 and to view shift the second depth map from the second depth sensing position 205 to the capture position 201.
[0080] Therefore, the view shift processor 305 is configured to generate a view-shifted depth map for a capture position from the depth maps of the different depth sensor positions, so that after the view shift of the view shift processor 305, all view-shifted depth maps represent depths from the same position, and in fact represent depths from the same position as the position at which the image is captured.
[0081] It will be appreciated that many different algorithms for view-shifting of images and depth maps are known to those skilled in the art, and these will not be described herein for the sake of brevity and clarity. It will be understood that any suitable view-shifting approach and algorithm may be used without detracting from the invention (and note that view-shifting of images may typically be applied directly to view-shifting of depth maps).
[0082] The view-shift processor (305) is further configured to designate pixels of the view-shifted depth map as uncertain or uncertain depending on the designation of the pixel in the original view-shifted depth map.
[0083] The view shift processor 305 is particularly configured to designate as uncertain pixels in the view-shifted depth map those pixels whose depth values are not projected by the view shift that are not designated as uncertain pixels in the original depth map.
[0084] A view shift moves the position of a pixel in the original depth map to a modified position in the view-shifted depth map that changes depending on the view shift. Thus, each pixel in the view-shifted depth map may receive no contribution from any pixel in the original depth map, or may receive contributions from one or more pixels. If a given pixel in the view-shifted depth map does not receive any contribution from a pixel in the original depth map that is considered a certain pixel, the pixel is designated as an uncertain pixel. If a pixel does not receive any contribution from any pixel, it is designated as an uncertain pixel. If a pixel receives contributions from one or more original pixels, all of which are designated as uncertain pixels, the pixel is also designated as an uncertain pixel. Such contributions may arise, for example, from depth values used for the uncertain pixel, such as default values or estimated values that are considered unreliable. However, in many embodiments, the depth values associated with uncertain pixels are not considered, and therefore, no view shift is performed on the uncertain pixel. Thus, in such a scenario, the uncertain pixel does not contribute to another pixel as a result of the view shift.
[0085] However, a pixel of the view-shifted depth map is designated as a certain pixel if the pixel receives contributions from one or more pixels in the original depth map that are designated as certain. If two or more certain pixels shift to a given pixel as a result of the view shift, the depth value can typically be selected as the one most forward.
[0086] Thus, in many embodiments, if there is no certain pixel projected onto a given pixel by view shift, the pixel can be determined to be uncertain, hi some embodiments, all other pixels can be designated as certain (non-uncertain) pixels.
[0087] In some embodiments, a pixel of a view-shifted depth map may be considered an uncertain pixel if one pixel of the original depth map projects onto it as a result of the view shift. Thus, in some embodiments, a pixel may be designated as a certain pixel if at least one certain pixel of the original depth map projects onto it, and in some embodiments, subject to the further requirement that no uncertain pixels project onto it. In some embodiments, all other pixels may be designated as uncertain pixels.
[0088] The depth value of a view-shifted pixel can be determined according to an appropriate combination criterion when multiple depth values are projected onto that pixel as a result of the view shift. In many embodiments, any depth values from uncertain pixels are discarded, and the combination can include only depth values from certain pixels. Typically, the combination of multiple depth values from certain pixels can be a selective combination, such as specifically selecting the most forward pixel. In other embodiments, other approaches can be used, including, for example, weighted combination or combination or selection based on confidence level.
[0089] The view-shift processor 305 is coupled to a combiner 307, which generates a combined depth map for the capture position by combining depth values of co-located pixels in the view-shifted depth maps. Co-located pixels in different depth maps (from the capture position) are pixels that represent the same view direction from the capture position. The co-located pixels can be pixels in the view-shifted depth map that represent the same view direction from the capture position. For a given pixel in the combined depth map, the depth values for the same pixel (at the same location) in the view-shifted depth map are combined according to an appropriate formula to generate a combined depth value for that pixel. This process can be repeated for all pixels in the combined depth map.
[0090] The combiner 307 is further configured to designate a pixel of the combined depth map as uncertain or not depending on the designation of the pixel of the view-shifted depth map to be combined. The combiner 307 is particularly configured to designate a pixel as uncertain if any of the co-located pixels from the view-shifted depth maps are designated as uncertain. Thus, if any pixel used in the combination to determine a combined depth value for the combined depth map is designated as an uncertain pixel, the corresponding pixel of the combined depth map is also designated as an uncertain pixel (even if one or more pixels are designated as certain). Specifically, the combiner 307 may designate a pixel of the combined depth map as certain only if all pixels of the co-located view-shifted depth map are designated as certain, and may designate the combined pixel as uncertain if any of the pixels are designated as uncertain.
[0091] In many embodiments, combiner 307 can be configured to generate the depth values of the combined depth map as a weighted combination of the depth values in the view-shifted depth maps, where the combination for a given pixel representing a given view direction can be a weighted combination of the depth values in the view-shifted depth maps representing the same direction that are not designated as uncertain depth pixels.
[0092] As a specific example, each view-shifted depth map can be initialized with a value that indicates that the pixel / depth value is uncertain. In the case of 16-bit integer depth encoding, the value 0 can generally be used. Depth map pixel value z at depth map / sensor / location i i (x, y) is view-shifted to capture position k, where Let TIFF2025529398000002.tif878 represent this view-shift depth value. Note that TIFF2025529398000003.tif976 also contains uncertain pixels, indicated by the reserved value u.
[0093] Now, if we have two predictions from opposite directions, the combined depth value z k (x, y) is It can be determined as TIFF2025529398000004.tif1890.
[0094] The above equation provides two functions: (1) uncertain pixels are handled; (2) if both predictions are certain, they are averaged to suppress remaining depth measurement noise. Such weighting will be more accurate in the presence of depth noise, depth quantization, and / or camera / sensor calibration errors.
[0095] The apparatus further comprises a depth map generator 309 configured to generate an output depth map for the capture location by determining depth values of the combined depth map for pixels designated as uncertain. Thus, the depth map generator 309 can generate (new) depth values for the uncertain pixels. The generation of these new values is typically based on at least one of image values of the image and depth values of the combined depth map that are not designated as uncertain.
[0096] In many embodiments, the depth map generator 309 can perform image-based extrapolation of certain depth values. For example, the depth value of an uncertain pixel can be determined by a weighted combination of the depth values of certain pixels within that pixel's kernel / neighborhood. The weight of the combination can depend on the image similarity between corresponding pixels in the image. Such an approach can typically provide very efficient generation of depth values. Alternatively, matching-based multi-view correspondence estimation can be used to estimate the depth of uncertain pixels. Certain pixels can help locally constrain the solution.
[0097] In many embodiments, the depth map generator 309 can be configured to determine the depth value of an uncertain pixel as a function of one or more depth values of certain pixels in a region determined for the location of the uncertain pixel, which function can depend on the image values for the location of the uncertain pixel and the location of the certain pixel, and in particular, the difference between the image values.
[0098] In many embodiments, the depth map generator 309 can be configured to determine the depth value of an uncertain pixel as a weighted combination (particularly often a weighted sum) of the depth values of a set of certain pixels (pixels not designated as uncertain pixels). The set of certain pixels can be pixels having locations that satisfy a geometric criterion relative to the location of the uncertain pixel, such as being within a given distance. In many embodiments, the weights can also depend on the image values of the certain pixels, typically the image values of the uncertain pixels. For example, the weights for certain pixels can be a monotonically decreasing function of the difference between the image values for the pixel locations of the certain pixels and the uncertain pixels.
[0099] In many embodiments, the depth map generator 309 can be configured to determine the depth value of an uncertain pixel by extrapolating the depth values of certain pixels, which may typically be, or at least include, neighboring pixels.
[0100] In some embodiments, the depth map generator 309 can be configured to determine a depth value for an uncertain pixel based solely on the depth values for one or more certain pixels. For example, in some cases, the depth map generator 309 can simply set the depth value of the uncertain pixel equal to the depth value of the nearest certain pixel, or, for example, as the average (or other weighted sum / kernel) of a set of certain pixels whose positions meet a criterion relative to the position of the uncertain pixel.
[0101] Indeed, in some cases, the depth value of an uncertain pixel can be determined by considering only its image value, rather than the depth values of certain pixels. For example, the depth map generator 309 can determine the color that is dominant for certain pixels and consider this to correspond to the background. If the image values of the uncertain pixel differ by less than a threshold, the depth map generator 309 can assign the uncertain pixel a nominal depth value of the background; otherwise, it can assign a nominal foreground depth value, such as a value corresponding to the display surface (often considered zero depth).
[0102] The described apparatus enables an improved depth map to be generated for an image from depth information provided by an (offset) depth sensor. In many embodiments, it can provide more accurate depth values, particularly around depth transitions within a scene when observed from a capture position. The improved depth map further enables improved view shifting and image synthesis based on the image and the combined depth map, thereby enabling improved services such as improved immersive video, augmented reality applications, and the like.
[0103] This approach is particularly based on the inventor's recognition that conventional approaches can introduce artifacts and errors.
[0104] 4 shows an example of a scene being captured, including a background object B and a foreground object F. This example further illustrates how the scene is observed / captured from a first depth sensing position 203. In this example, an uncertain depth sensing area U is captured around a depth transition. Due to the depth transition, the depth sensor cannot provide an accurate depth measurement, and therefore the depth values around the depth transition are considered uncertain, and therefore the pixels around this region are designated as uncertain.
[0105] When this first depth map is shifted to the capture position 201 (image camera position) view, the foreground and background depth values move separately, changing the view of the depth transition and the region corresponding to the uncertain depth transition, as shown in FIG. 4. Specifically, two separate regions U of uncertain pixels are created. Furthermore, a gap G is introduced where the depth value cannot be determined based on the first depth map (i.e., disocclusion occurs). In such an example, the view shift processor 305 can set the entire region of uncertain and disoccluded pixels (U+G) as uncertain pixels U'.
[0106] FIG. 5 shows the corresponding situation captured from the second depth sensing position 205. In this case, no gaps exist because no unocclusion occurs. Furthermore, the uncertainty region of the foreground object not only overlaps with the uncertainty region U′ of the background object, but also overlaps with part of the certain region E of the background object. The corresponding situation when considering the depth map is shown in FIG. 6, which shows the depth map captured at the second depth sensing position 205 and the view-shifted depth map obtained for the capture position 201, respectively. This figure also shows the location of the true depth step 601 for reference. As shown, region 603 corresponds to the overlap of the uncertainty region of background object B.
[0107] Therefore, warping (view-shifting) the depth image to a new viewpoint can cause valid / certain depth pixels of the background to cover new uncertain regions (of foreground objects). Therefore, special care must preferably be taken when combining the view-shift and depth map to avoid introducing undesired effects.
[0108] The above-described approach is configured to mitigate or even prevent the occurrence of such a scenario. This reflects the inventor's recognition that such a problem occurs particularly when a foreground object occludes a background object, but not when the foreground object unoccludes the background object. Furthermore, the inventor recognized that an improved approach can be achieved by positioning a depth sensor at a different orientation relative to the capture location, view-shifting the images, and combining the view-shifted images as described above. In particular, this approach can designate a pixel as uncertain if the pixel is uncertain when resulting from a view shift from a depth sensor in a different orientation, which can prevent the pixel from being designated as certain unless it is guaranteed that no occlusion of the background object occurs. While this approach may typically result in a larger uncertainty region, it ensures that a depth value is not designated as certain while resulting from occlusion of a background object by a foreground object. Therefore, by combining and considering more depth maps from different orientations, the risk of designating a flawed depth value as certain due to a warping operation can be reduced.
[0109] In the above approach, the device is configured to designate a pixel as uncertain if any of the view-shifted depth maps actually have a corresponding pixel designated as uncertain. Figure 7 shows this applied to the scenarios of Figures 4-6, specifically, Figure 7 shows the designation applied to each horizontal row of a first view-shifted depth map 701, a second view-shifted depth map 703, and the resulting combined depth map 705.
[0110] This approach can automatically adapt to scene characteristics, and typically adaptively determine pixel uncertainty regions that reflect the actual scene scenario. Thus, for each capture scenario, this approach can automatically adapt the size of the uncertainty region to correspond to the actual scene characteristics. For example, this approach is automatically suitable for depth transitions in opposite directions, and no information about the direction or presence of any depth steps is required.
[0111] While this approach may in some scenarios designate larger regions of pixels as uncertain than is perhaps strictly necessary, this is typically a perfectly acceptable trade-off since depth values in such areas can be determined image-based, based on extrapolation, multi-view matching, etc.
[0112] In some embodiments, the device may further comprise a view synthesis processor (not shown) configured to generate view images for a view pose (at a location different from the capture location) from the image and the combined depth map. It will be appreciated that many algorithms for such view synthesis are known to those skilled in the art.
[0113] The described approach is particularly suitable for some capture configurations that include one or more image cameras and two or more depth sensors. In particular, the capture configuration can be advantageously configured such that occlusion by a foreground object does not (and ideally cannot) occur simultaneously for all depth sensors.
[0114] In many embodiments, such as the example of Figure 2, the image cameras and depth sensors can be advantageously arranged in a linear configuration, advantageously with a depth sensor on either side of the image camera. Thus, in many embodiments, they can be advantageously arranged in a linear configuration with capture locations between depth sensing locations. Such a configuration can be extended to multiple image cameras and more depth sensors, as shown in Figure 8 (filled triangles indicate depth sensors, and unfilled triangles indicate image cameras).
[0115] In these examples, the depth sensors and image cameras are arranged along a line, and the direction from the capture location 201 to at least two different depth sensors is 180°. Thus, in these examples, the direction from the capture location to one depth sensing location and the direction from the capture location to another depth sensing location form an angle of 180°. However, this angle need not be strictly required to be 180°, and in some embodiments, it may be narrower. While many embodiments benefit from an angle as close to 180° as possible, in many embodiments, it is sufficient for this angle to be 90°, 100°, 120°, 140°, 160°, or greater, depending on the specific requirements of each individual embodiment. Thus, in many embodiments, the angle between the line connecting the capture location to a first location and the line connecting the capture location to a second location is 90°, 100°, 120°, 140°, 160°, or greater. In many embodiments, the angle between a line passing through the capture position and the first position and a line passing through the capture position and the second position is 90°, 100°, 120°, 140°, 160° or more.
[0116] The apparatus typically arranges at least two depth sensing locations such that when a view shift to a capture location is performed, the movement directions of pixels at the same depth caused by the view shift are opposite. For a given pixel in the combined depth map, one contribution can be provided from the pixel to the right of the given pixel, and another contribution can be provided from the pixel to the left of the given pixel. In some embodiments, the first location, the second location, and the capture location are arranged such that the movement directions of image coordinates for a given depth are opposite for view shifts of at least two different depth maps / at least two depth sensing locations.
[0117] Such an effect is typically achieved, for example, by configurations such as those previously described (and described below), and can provide particularly advantageous performance by ensuring that potentially occluded pixels are appropriately designated as uncertain.
[0118] In many embodiments, the capture arrangement may comprise multiple image cameras positioned at different locations. Such an approach allows for improved capture of a scene, and in particular, allows for 3D processing and view synthesis for different view poses (e.g., reducing the risk that parts of the scene will be occluded in all captured images).
[0119] In many cases, the placement of at least the second imaging camera will be such that the characteristics described above for the first imaging camera / capture location are also satisfied for the second imaging camera / capture location. In particular, the angle between a line passing through the direction from the second capture location to the first depth sensing location and a line passing through the direction from that further capture location to the second location is 90°, 100°, 120°, 140°, 160°, or greater. In such a configuration, the operations described for the first imaging camera / capture location can also be performed for the second imaging camera / capture location based on the same depth sensor and the same depth map. It will be understood that this approach may be repeated for more cameras. For example, FIGS. 9 and 10 show examples of linear configurations that can process multiple cameras as described based on a common depth sensor.
[0120] In many embodiments, a capture configuration can be adopted in which multiple image cameras and multiple depth sensors are configured such that for all image cameras, the angle between a line passing through / in the direction from the corresponding capture position to at least one depth sensing position and a line passing through / in the direction from the corresponding capture position to at least another depth sensing position is 90°, 100°, 120°, 140°, 160° or more.
[0121] In many such embodiments, the image camera and image sensor are advantageously arranged in a linear configuration, examples of suitable configurations are shown in Figures 8, 9 and 10.
[0122] In some such systems, at least one depth sensor can be advantageously positioned between each pair of adjacent image cameras of the multiple image cameras. In particular, in some embodiments, this configuration can include an alternating arrangement of depth sensors and image cameras. Such an approach can, for example, enable improved performance, for example, reducing the area of uncertain pixels because a shorter view shift from the nearest depth sensor in either direction is required. Furthermore, many embodiments can enable easy implementation. For example, this configuration can be achieved using multiple integrated cameras, each including one image camera and one depth sensor. In many cases, such integrated cameras can simply be arranged in a row.
[0123] In the preceding discussion, the capture configurations considered were primarily one-dimensional configurations of cameras and depth sensors arranged in a line. However, in many embodiments, the image cameras and depth sensors can be advantageously arranged in a two-dimensional configuration. In such examples, the depth sensor may not be directly beside the depth camera, but may be offset, for example, in a different direction. Such an example is shown in FIG. 11.
[0124] The foregoing properties and relationships can also be applied to two-dimensional configurations, typically having a relatively large number of image cameras and / or depth sensors. This approach can specifically involve selecting an appropriate set (or subset) of depth sensors for each camera and performing the described operations for these depth sensors. In many embodiments, the subset of depth sensors for a given depth camera can consist of two depth sensors.
[0125] In many embodiments, (at least some) depth sensors can be positioned asymmetrically around the image camera. For example, in many embodiments, the distance to the next-closest depth sensor / depth sensor location can be more than twice, or in some cases more than five or ten times, the distance to the closest depth sensor / depth sensing location. For example, in many embodiments, depth sensors can be positioned very close to one, some, or often all, of the image cameras. Such nearby sensors tend to reduce the view shift required to change from the depth sensing location to the camera location and provide a more accurate depth map for the image. Another depth sensor can be positioned significantly further away, for example, on the opposite side of the image camera, and is particularly well-suited to help determine uncertain pixels as described above. Due to this additional distance, depth values may be less accurate; therefore, depth values can be based primarily or exclusively on the nearby depth sensor, while the remote depth sensor can be used primarily or exclusively to determine the uncertainty region. Such an approach is highly advantageous in many embodiments, as it enables or facilitates the use of several depth sensors for multiple image cameras. For example, in the example of FIG. 12, each camera can be approximately co-located with one image sensor (indeed, they can be part of the same integrated camera device). This depth sensor can be primarily the source of depth values for the depth map generated for the image cameras. However, in addition, a single depth sensor can be placed in the center of the multiple image cameras such that it is opposite the local depth sensor. This single depth sensor can allow the described approach to be used for all image cameras, so that only a single central depth sensor is needed.
[0126] In some capture configurations, the number of depth sensors can exceed the number of image cameras. This can facilitate operation and / or improve performance in many embodiments. For example, it is well suited to scenarios where depth sensors are significantly cheaper than image cameras. Typically, a larger number of depth sensors can provide more accurate depth information for an image and allow for more flexible configurations. It also typically reduces the size of the region of uncertain pixels. The described approach can provide very advantageous operation in scenarios where the number of depth sensors exceeds the number of image cameras.
[0127] In some capture devices, the number of image cameras can exceed the number of depth sensors. This can facilitate operation and / or improve performance in many embodiments. It is well suited, for example, to scenarios where image cameras are cheaper or more readily available than depth sensors. The described approach can provide highly advantageous operation in scenarios where the number of image cameras exceeds the number of depth sensors.
[0128] 13 is a block diagram illustrating an exemplary processor 1300 according to an embodiment of the disclosure. The processor 1300 can be used to implement one or more processors that implement devices or elements thereof as described above. The processor 1300 can be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) / (where the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) / (where the ASIC is designed to form a processor), or a combination thereof.
[0129] Processor 1300 may include one or more cores 1302. Core 1302 may include one or more arithmetic logic units (ALUs) 1304. In some embodiments, core 1302 may include a floating point logic unit (FPLU) 1306 and / or a digital signal processing unit (DSPU) 1308 in addition to or instead of ALU 1304.
[0130] The processor 1300 may include one or more registers 1312 communicatively coupled to the core 1302. The registers 1312 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 1312 may be implemented using static memory. The registers may provide data, instructions, and addresses to the core 1302.
[0131] In some embodiments, processor 1300 may include one or more levels of cache memory 1310 communicatively coupled to cores 1302. Cache memory 1310 may provide computer-readable instructions to cores 1302 for execution. Cache memory 1310 may provide data for processing by cores 1302. In some embodiments, computer-readable instructions may be provided to cache memory 1310 by local memory, for example, local memory attached to external bus 1316. Cache memory 1310 may be implemented using any suitable cache memory type, such as, for example, static random access memory, dynamic random access memory, and / or any other suitable memory technology.
[0132] Processor 1300 may include a controller 1314 that can control input to processor 1300 from other processors and / or components included in the system and / or output from processor 1300 to other processors and / or components included in the system. Controller 1314 may control data paths within ALU 1304, FPLU 1306, and / or DSPU 1308. Controller 1314 may be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates in controller 1314 may be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0133] Registers 1312 and cache memory 1310 may communicate with controller 1314 and core 1302 via internal connections 1320A, 1320B, 1320C, and 1320D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology. Input and output for processor 1300 may be provided via bus 1316, which may include one or more conductive lines. Bus 1316 may be communicatively coupled to one or more components of processor 1300, such as controller 1314, cache memory 1310, and / or registers 1312. Bus 1316 may be coupled to one or more components of the system.
[0134] The bus 1316 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 1332. The ROM 1332 may be masked ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory 1333. The RAM 1333 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 1335. The external memory may include flash memory 1334. The external memory may include a magnetic storage device such as a disk 1336. In some embodiments, external memory may be included in the system. All positions can be referenced to the scene coordinate system.
[0135] The invention can be implemented in any suitable form including hardware, software, firmware, or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.
[0136] In accordance with standard terminology in the field, the term pixel may be used to refer to pixel-related properties such as light intensity, depth, position, etc. of the portion / element of a scene represented by the pixel. For example, the depth of a pixel, i.e., pixel depth, may be understood to refer to the depth of the object represented by that pixel. Similarly, the brightness of a pixel, i.e., pixel luminosity, may be understood to refer to the brightness of the object represented by that pixel. Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0137] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, these may be advantageously combined in some cases, and their inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must operate, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, a reference to the singular does not exclude a plurality. Thus, references to "a," "an," "first," "second," etc., do not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.
Claims
1. A device for generating a depth map for an image representing a view of a scene, A position receiver configured to receive the capture position of an image, wherein the capture position is the capture position of the image, A depth receiver configured to receive a first depth map providing depth values from a first position and a second depth map providing depth values from a second position, wherein the first depth map includes at least some pixels designated as uncertain depth pixels, and the second depth map includes at least some pixels designated as uncertain depth pixels, A view shift processor configured to perform a first view shift of the first depth map from the first position to the capture position in order to generate a first view shift depth map, and to perform a second view shift of the second depth map from the second position to the capture position in order to generate a second view shift depth map, further configured to designate as uncertain pixels among the pixels of the first view shift depth map, pixels on which the depth values of pixels not designated as uncertain pixels in the first depth map by the first view shift are not projected, and pixels on the second view shift depth map, on which the depth values of pixels not designated as uncertain pixels in the second depth map by the second view shift are not projected, A coupler configured to generate a combined depth map for the capture position by combining the depth values of pixels at the same position in the first viewshift depth map and the second viewshift depth map, wherein if any of the pixels at the same position are designated as uncertain pixels, the coupler is configured to designate the pixels in the combined depth map as uncertain depth pixels, A device comprising: a depth map generator configured to generate an output depth map for the capture location by determining the depth values of the combined depth map for pixels designated as uncertain pixels using at least one of the image values of the image and the depth values of pixels in the combined depth map that are not designated as uncertain pixels.
2. The apparatus according to claim 1, wherein the angle between the direction from the capture position to the first position and the direction from the capture position to the second position is 90° or more.
3. The apparatus according to claim 1, wherein the first position, the second position and the capture position are arranged in a linear configuration, and the capture position is located between the first position and the second position.
4. The apparatus according to claim 1, wherein the first position, the second position and the capture position are arranged such that the direction of pixel movement at the same depth is opposite to that of the first view shift and the second view shift.
5. The apparatus according to claim 1, wherein the coupler is configured to determine the depth value of the combined depth map for a given pixel as a weighted combination of a first depth value of a pixel in the first viewshift depth map that is not designated as an uncertain depth pixel at the same position as the given pixel, and a second depth value of a pixel in the second viewshift depth map that is not designated as an uncertain depth pixel at the same position as the given pixel.
6. The apparatus according to claim 1, wherein the depth map generator is configured to generate depth values for pixels designated as uncertain in the combined depth map by estimation from depth values for pixels not designated as uncertain in the combined depth map.
7. A capture system having the device described in any one of claims 1 to 6, A first image camera at the capture position is configured to provide the aforementioned image to the image receiver, A first depth sensor at the first position, configured to provide depth data for the first depth map to the depth receiver, A second depth sensor at the second position, configured to provide depth data for the second depth map to the depth receiver, A capture system that further includes [the following features].
8. Having at least one additional image camera at a further capture position, The capture system according to claim 7, wherein the angle between the direction from the further capture position to the first position and the direction from the further capture position to the second position is 90° or more.
9. The capture system according to claim 7, comprising a plurality of image cameras including the first image camera, and a plurality of depth sensors including the first depth sensor and the second depth sensor, wherein for all of the plurality of image cameras, two of the plurality of depth sensors are arranged such that the angle between the direction from the position of the image camera to the two depth sensors is 120° or more.
10. The capture system according to claim 9, wherein the plurality of image cameras and the plurality of depth sensors are arranged in a linear configuration.
11. The capture system according to claim 9, wherein at least one of the plurality of depth sensors is positioned between each pair of adjacent image cameras among the plurality of image cameras.
12. The capture system according to claim 11, wherein the plurality of image cameras and the plurality of depth sensors are arranged in a two-dimensional configuration.
13. The capture system according to claim 9, wherein the number of depth sensors among the plurality of depth sensors exceeds the number of image cameras among the plurality of image cameras.
14. The capture system according to claim 9, wherein the number of image cameras among the plurality of image cameras exceeds the number of depth sensors among the plurality of depth sensors.
15. A method for generating a depth map for an image representing a view of a scene, A step of receiving the capture position of the aforementioned image, wherein the capture position is the capture position of the aforementioned image, A step of receiving a first depth map that provides depth values from a first position and a second depth map that provides depth values from a second position, wherein the first depth map includes at least some pixels designated as uncertain depth pixels, and the second depth map includes at least some pixels designated as uncertain depth pixels. The steps include: performing a first view shift of the first depth map from the first position to the capture position in order to generate a first view shift depth map; The steps include performing a second view shift of the second depth map from the second position to the capture position in order to generate a second view shift depth map, The steps of designating pixels in the first viewshift depth map that do not have depth values projected for pixels that are not designated as uncertain pixels in the first depth map by the first viewshift, and pixels in the second viewshift depth map that do not have depth values projected for pixels that are not designated as uncertain pixels in the second depth map by the second viewshift, as uncertain. A step of generating a combined depth map for the capture position by combining the depth values of pixels at the same position in the first viewshift depth map and the second viewshift depth map, wherein if any of the pixels at the same position are designated as uncertain pixels, the pixels in the combined depth map are designated as uncertain depth pixels. The steps of generating an output depth map for the capture location by determining the depth values of the combined depth map for pixels designated as uncertain using at least one of the image values of the image and depth values not designated as uncertain, A method of having.
16. A computer program that is executed by a computer and causes the computer to perform the method of claim 15.