Processing of depth maps for images
By receiving and updating the depth map of multiple images, using weighted combination and projection error indication to improve depth map consistency, solving the problem of depth map inconsistency in virtual reality applications, and improving view consistency and user experience.
Patent Information
- Application Number
- CN202080033592.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-05
- Filing Date
- 2020-03-03
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-03-03
AI Technical Summary
The depth map generated by the prior art has inconsistency and error in virtual reality applications, resulting in a downgrade of viewer experience, and the existing methods have failed to effectively solve the error and inconsistency problems between depth maps.
By receiving a plurality of images and corresponding depth maps, the depth value of the first depth map is updated based on the depth value of at least the second depth map, and the weight is determined using weighting combination and projection error indication to improve the consistency of the depth map.
Provides a more consistent and high-quality depth map, improving view consistency and user experience in virtual reality applications, and reducing complexity and resource requirements.
Smart Images

Figure CN113795863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the processing of depth maps for images, and in particular, but not exclusively, to the processing of depth maps to support view synthesis for virtual reality applications. Background Art
[0002] The variety and range of image and video applications have increased dramatically in recent years, with the continuous development and introduction of new services and ways to utilize and consume video.
[0003] For example, an increasingly popular service is one that provides image sequences in a way that allows the viewer to actively and dynamically interact with the system to change the parameters of the presentation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and viewing direction, for example allowing the viewer to move and "look around" within the presented scene.
[0004] This feature can specifically allow users to be provided with a virtual reality experience. This can allow the user, for example, to move (relatively) freely within the virtual environment and dynamically change their position and where they are looking. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide the specific requested view. This approach is well known, for example, from gaming applications for computers and game consoles, such as first-person shooter games.
[0005] It is also desirable, particularly for virtual reality applications, that the presented image be three-dimensional. Indeed, to optimize the viewer's immersion, it is often more preferable for the user to experience the presented scene as a three-dimensional scene. Indeed, a virtual reality experience should preferably allow the user to select their own position, camera viewpoint, and moment in time relative to the virtual world.
[0006] Many virtual reality applications are based on models of predetermined scenes, and often on artificial models of the virtual world. It is often desirable to provide a virtual reality experience based on real-world capture.
[0007] In many systems, such as those based on real-world scenes, an image representation of the scene is provided, wherein the image representation includes an image and depth for one or more capture points / viewpoints in the scene. The image-plus-depth representation provides a very effective representation of, in particular, real-world scenes, wherein the representation is not only relatively easy to generate by capturing the real-world scene, but is also well-suited for a renderer to synthesize views from viewpoints other than those captured. For example, a renderer can be arranged to dynamically generate a view that matches the current local viewer pose. For example, the viewer pose can be dynamically determined, and views can be dynamically generated to match the viewer pose based on the image and, for example, a provided depth map.
[0008] In many practical systems, calibrated multi-view camera rigs can be used to allow users to take different perspectives on the captured scene during playback. Applications include selecting a personal viewpoint during a sports game or replaying a captured 3D scene on an augmented reality or virtual reality headset.
[0009] You Yang et al., "Cross-View Multi-Lateral Filter for Compressed MultiView Depth Video" (IEEE TRANSACTIONS ON IMAGE PROCESSING., January 1, 2019 (2019-01-01), Vol. 28, Part 1, pp. 302-315, XP055614403, US ISSN: 1057-7149, DOI: 10.1109 / TIP.2018.2867740) disclose a cross-view multi-lateral filtering scheme to improve the quality of compressed depth maps / videos within an asymmetric multi-view video framework with depth compression. With this scheme, the distorted depth map is enhanced via non-local candidates selected from the current and neighboring viewpoints at different time slots. Specifically, these candidates are clustered into a macrosuper pixel that represents the cross-view, spatial, and temporal priority of the physical and semantic cross-relationship.
[0010] WOLFF KATJA et al., “Point Cloud Noise and Outlier Removal for Image-Based 3D Reconstruction,” 2016 FOURTH INTERNATIONAL CONFERENCE ON 3DVISION (3DV), IEEE, October 25, 2016, pp. 1178–25, XP033027617, DOI: 10.1109 / 3DV.2016.20, disclose an algorithm that uses an input image and a corresponding depth map to remove pixels that are geometrically or photometrically inconsistent with the color surface implied by the input. This allows standard surface reconstruction methods to perform less smoothing and, therefore, achieve higher quality.
[0011] In order to provide smooth transitions between discrete captured viewpoints as well as some extrapolation beyond the captured viewpoints, a depth map is usually provided and used to predict / synthesize views from these other viewpoints.
[0012] Depth maps are typically generated using (multi-view) stereo matching between capture cameras or more directly using depth sensors (based on structured light or time-of-flight). However, such depth maps obtained from depth sensors or disparity estimation processes inherently have errors and can lead to errors and inaccuracies in the synthesized view. This degrades the viewer's experience.
[0013] Therefore, improved methods for generating and processing depth maps would be advantageous. In particular, systems and / or methods that allow for improved operation, increased flexibility, improved virtual reality experience, reduced complexity, facilitated implementation, improved depth maps, improved synthesized image quality, improved rendering, improved user experience, and / or improved performance and / or operation would be advantageous. Summary of the Invention
[0014] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.
[0015] According to one aspect of the present invention, a method for processing a depth map is provided, the method comprising: receiving a plurality of images representing a scene from different viewing postures and corresponding depth maps; updating a depth value of a first depth map of the corresponding depth map based on a depth value of at least a second depth map of the corresponding depth map, the first depth map being for the first image, and the second depth map being for the second image; the updating comprising: determining a first candidate depth value for a first depth pixel of the first depth map at a first depth map position in the first depth map, the first candidate depth value being determined in response to at least one second depth value of a second depth pixel of the second depth map at a second depth map position in the second depth map; determining a first depth value for a first depth pixel based on a weighted combination of multiple candidate depth values for a depth map position, the weighted combination including the first candidate depth value weighted by a first weight; wherein determining the first depth value includes: determining a first image position in a first image for the first depth map position, determining a third image position in a third image of the plurality of images, the third image position corresponding to a projection of the first image position to the third image based on the first candidate depth value; determining a first match error indication indicating a difference between an image pixel value in the third image for the third image position and an image pixel value in the first image for the first image position, and determining the first weight in response to the first match error indication.
[0016] The method may provide an improved depth map in many embodiments, and may specifically provide a set of depth maps with increased consistency. The method may allow for improved visual quality when synthesizing an image based on the image and the updated depth map. Figure 1 Consistency.
[0017] The inventors have recognized that inconsistencies between depth maps may often be more perceptible than consistent errors or noise between depth maps, and that certain methods can provide updated depth maps that are more consistent. The methods can be used as depth refinement algorithms that improve the quality of depth maps for a set of multi-view images of a scene.
[0018] The described method may be convenient to implement in many embodiments and may be implemented with relatively low complexity and resource requirements.
[0019] A location in the image can directly correspond to a location in the corresponding depth map, and vice versa. There can be a one-to-one correspondence between a location in the image and a location in the corresponding depth map. In many embodiments, the pixel location in the image and the corresponding depth map can be the same, and the corresponding depth map can include one pixel for each pixel in the image.
[0020] In some embodiments, the weights may be binary (eg, one or zero), and the weighted combination may be a selection.
[0021] It should be understood that the term projection generally refers to the projection of three-dimensional spatial coordinates in a scene to two-dimensional image coordinates (u, v) in an image or depth map. However, projection can also refer to the mapping between dimensional image coordinates (u, v) for scene points from one image or depth map to another, i.e., from a set of image coordinates (u1, v1) for one pose to another set of image coordinates (u2, v2) for another pose. Such projections between image coordinates for images corresponding to different viewing poses / positions generally take into account the corresponding spatial scene points, and are specifically performed by taking into account the depth of the scene points.
[0022] In some embodiments, determining the first depth value includes: projecting the first depth map position to a third depth map position in a third depth map in the corresponding depth map, the third depth map being for a third image, and the projection is based on the first candidate depth value, determining a first match error indication, the first match error indication indicating a difference between an image pixel value in the third image for the third depth map position and an image pixel value in the first image for the first depth map position, and determining a first weight in response to the first match error indication.
[0023] According to an optional feature of the invention, determining the first candidate depth value comprises determining, based on the second value of the first depth map and at least one of the first original depth value, a second depth map position relative to the first depth map position by a projection between a first viewing pose of the first image and a second viewing pose of the second image.
[0024] This may provide particularly advantageous properties in many embodiments, and may in particular allow for improved depth maps with improved consistency in many scenes.
[0025] The projection can be based on the first original depth value, from the second depth map position to the first depth map position, and thus from the second viewing pose to the first viewing pose.
[0026] The projection can be based on the second depth value, from the second depth map position to the first depth map position, and thus from the second viewing pose to the first viewing pose.
[0027] The original depth value may be a non-updated depth value of the first depth map.
[0028] The original depth values may be depth values of the first depth map as received by the receiver.
[0029] In some embodiments, determining the first depth value includes projecting the first depth map position to a third depth map position in a third depth map of the corresponding depth map, the third depth map being for the third image and the projection being based on the first candidate depth value, determining a first match error indication indicating a difference between an image pixel value in the third image for the third depth map position and an image pixel value in the first image for the first depth map position, and determining a first weight in response to the first match error indication.
[0030] According to an optional feature of the invention, the weighted combination comprises candidate depth values determined from regions of the second depth map determined responsive to positions of the first depth map.
[0031] This can provide increased depth in many embodiments. Figure 1 The first candidate depth value may be derived from one or more depth values of the region.
[0032] According to an optional feature of the invention, the region of the second depth map is determined as a region around the second depth map position, and the second depth map position is determined as a depth map position in the second depth map equal to the first depth map position in the first depth map.
[0033] This may allow for a low complexity and low resource but efficient determination of a suitable depth value.
[0034] According to an optional feature of the invention, the region of the second depth map is determined as a region around a position in the second depth map at the first depth position determined by a projection from the first depth map position based on the original depth value in the first depth map.
[0035] In many embodiments, this can provide increased depth Figure 1The original depth value may be a depth value of a first depth map received by the receiver.
[0036] According to an optional feature of the invention, the method further comprises determining a second matching error indication, the second matching error indication indicating a difference between an image pixel value in the second image for the second depth map position and an image pixel value in the first image for the first depth map position; and wherein determining the first weight is also responsive to the second matching error indication.
[0037] This may provide an improved depth map in many embodiments.
[0038] According to an optional feature of the invention, the method further comprises determining an additional matching error indication, wherein the additional matching error indication indicates a difference between an image pixel value in the other image for the depth map position and an image pixel value in the first image for the first depth map position corresponding to the first depth map position; and wherein determining the first weight is also responsive to the additional matching error indication.
[0039] This may provide an improved depth map in many embodiments.
[0040] According to an optional feature of the invention, the weighted combination comprises depth values of the first depth map in an area surrounding the position of the first depth map.
[0041] This may provide an improved depth map in many embodiments.
[0042] According to an optional feature of the invention, the first weight depends on a confidence value of the first candidate depth value.
[0043] This can provide improved depth maps in many scenarios.
[0044] According to an optional feature of the invention, only the depth values of the first depth map having a confidence value below a threshold are updated.
[0045] This may provide improved depth maps in many scenarios, and may specifically reduce the risk of updating accurate depth values to less accurate depth values.
[0046] According to an optional feature of the invention, the method further comprises selecting a set of depth values of the second depth map for inclusion in the weighted combination which satisfy a requirement that depth values of the set of depth values must have confidence values above a threshold.
[0047] This can provide improved depth maps in many scenarios.
[0048] According to an optional feature of the invention, the method further comprises: projecting a given depth map position for a given depth value in a given depth map to a corresponding position in a plurality of corresponding depth maps; determining a variation metric for a set of depth values, the set of depth values comprising the given depth value and depth values at corresponding positions in the plurality of corresponding depth maps; and determining a confidence value for the given depth map position in response to the variation metric.
[0049] This may provide a particularly advantageous determination of confidence values which may result in an improved depth map.
[0050] According to an optional feature of the present invention, the method further comprises: projecting a given depth map position for a given depth value in a given depth map to a corresponding position in another depth map, wherein the projection is based on the given depth value; projecting the corresponding position in the other depth map to a test position in the given depth map, wherein the projection is based on the depth value at the corresponding position in the other depth map; and determining a confidence value for the given depth map position in response to a distance between the given depth map position and the test position.
[0051] This may provide a particularly advantageous determination of confidence values which may result in an improved depth map.
[0052] According to one aspect of the present invention, there is provided an apparatus for processing a depth map, the apparatus comprising: a receiver for receiving a plurality of images representing a scene from different viewing postures and corresponding depth maps; an updater for updating a depth value of a first depth map of a corresponding depth map based on a depth value of at least a second depth map of the corresponding depth map, the first depth map being for a first image and the second depth map being for a second image; the updating comprising: determining a first candidate depth value for a first depth pixel of the first depth map at a first depth map position in the first depth map, the first candidate depth value being determined in response to at least one second depth value of a second depth pixel of the second depth map at a second depth map position in the second depth map; determining a first depth value for a first depth pixel by a weighted combination of a plurality of candidate depth values for a first depth map position, the weighted combination comprising the first candidate depth value weighted by a first weight; wherein determining the first depth value comprises: determining a first image position in a first image for the first depth map position, determining a third image position in a third image of the plurality of images, the third image position corresponding to a projection of the first image onto the third image based on the first candidate depth value; determining a first match error indication, the first match error indication indicating a difference between an image pixel value in the third image for the third image position and an image pixel value in the first image for the first image position, and determining the first weight in response to the first match error indication.
[0053] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which
[0055] Figure 1 An example of an apparatus for providing a virtual reality experience is shown;
[0056] Figure 2 shows an example of elements of an apparatus for processing a depth map according to some embodiments of the present invention;
[0057] Figure 3 shows an example of elements of a method of processing a depth map according to some embodiments of the present invention;
[0058] Figure 4 shows an example of a camera setup for capturing a scene;
[0059] Figure 5 shows an example of elements of a method of updating a depth map according to some embodiments of the present invention;
[0060] Figure 6 shows an example of elements of a method of determining weights according to some embodiments of the present invention;
[0061] Figure 7 An example of processing of depth maps and images according to some embodiments of the present invention is shown. DETAILED DESCRIPTION
[0062] The following description focuses on an embodiment of the invention applicable to virtual reality experiences, but it will be appreciated that the invention is not limited to this application, but may be applied to many other systems and applications, such as specific applications involving view synthesis.
[0063] Virtual experiences that allow users to walk around in a virtual world are becoming increasingly popular, and services are being developed to meet this demand. However, providing efficient virtual reality services is very challenging, particularly if the experience is based on a capture of the real-world environment rather than a fully virtual generated artificial world.
[0064] In many virtual reality applications, viewer gesture input is determined, reflecting the gesture of a virtual viewer in a scene. The virtual reality device / system / application then generates one or more images of the viewer's view and viewport corresponding to the scene corresponding to the viewer gesture.
[0065] Typically, virtual reality applications generate three-dimensional output in the form of separate view images for the left and right eyes. These can then be presented to the user by suitable means, such as separate left-eye and right-eye displays, typically of a VR headset. In other embodiments, the images may be presented, for example, on an autostereoscopic display (in which case a large number of view images may be generated for the viewer's posture), or indeed in some embodiments, only a single two-dimensional image may be generated (e.g., using a conventional two-dimensional display).
[0066] Viewer gesture input can be determined in different ways in different applications. In many embodiments, the user's body movements can be tracked directly. For example, a camera surveying the user's area can detect and track the user's head (or even eyes). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal means. For example, the headset can include accelerometers and gyroscopes that provide information about the movement and rotation of the headset and, thereby, the head. In some examples, the VR headset can transmit a signal or include a (e.g., visual) identifier that enables external sensors to determine movement of the VR headset.
[0067] In some systems, the viewer gesture may be provided manually, such as by the user manually controlling a joystick or similar manual input. For example, the user may manually move the virtual viewer around the scene by controlling a first analog joystick with one hand and manually controlling the direction the virtual viewer is looking by manually moving a second analog joystick with the other hand.
[0068] In some applications, a combination of manual and automatic methods can be used to generate the input viewer pose. For example, a headset can track the orientation of the head, and a joystick can be used by the user to control the viewer's movement / position in the scene.
[0069] The generation of the image is based on a suitable representation of the virtual world / environment / scene. In some applications, a complete three-dimensional model of the scene may be provided, and the view of the scene from a particular viewer pose can be determined by evaluating this model.
[0070] In many practical systems, a scene can be represented by an image representation comprising image data. The image data can typically include one or more images associated with one or more capture or anchor poses, and in particular can include images for one or more viewports, where each viewport corresponds to a particular pose. An image representation can be used that includes one or more images, where each image represents a view of a given viewport for a given viewing pose. Such viewing poses or positions for which image data is provided are also typically referred to as anchor poses or positions, or capture poses or positions (because the image data can typically correspond to an image captured or to be captured by a camera positioned in the scene at a position and orientation corresponding to the capture pose).
[0071] Images are often associated with depth information, and in particular, a depth image or depth map is often provided. Such a depth map can provide a depth value for each pixel in the corresponding image, where the depth value indicates the distance from the camera / anchor / capture position to the object / scene point depicted by the pixel. Thus, the pixel value can be considered to represent a ray from the object / point in the scene to the camera's capture device, and the depth value for the pixel can reflect the length of this ray.
[0072] In many embodiments, the resolution of the image and the corresponding depth map may be the same, and thus an individual depth value may be included for each pixel in the image, i.e., the depth map may include one depth value for each pixel of the image. In other embodiments, the resolution may be different, and for example, the depth map may have a lower resolution, such that one depth value may apply to multiple image pixels. The following description will focus on embodiments where the resolution of the image and the corresponding depth map are the same, and thus for each image pixel (pixel of the image), there is a separate depth map pixel (pixel of the depth map).
[0073] A depth value may be any value that indicates the depth for a pixel, and thus it may be any value that indicates the distance from the camera position to an object of the scene depicted by a given pixel. A depth value may be, for example, a disparity value, a z-coordinate, a distance metric, or the like.
[0074] Many typical VR applications can continue to provide, based on this image plus depth representation, a view image corresponding to the viewport of the scene for the current viewer pose, and an image that is dynamically updated to reflect changes in the viewer pose, as well as an image that is being generated based on image data representing the (possible) virtual scene / environment / world. The application can do this by executing view synthesis and view shifting algorithms known to those skilled in the art.
[0075] In the art, the terms placement and pose are used as general terms for position and / or direction / orientation. For example, the combination of the position and direction / orientation of an object, camera, head, or view may be referred to as a pose or placement. Thus, a placement or pose indication may include six values / components / degrees of freedom, each value / component typically describing an individual property of the position / position or orientation / orientation of the corresponding object. Of course, in many cases, a placement or pose may be considered to have fewer components or be represented by fewer components, for example, if one or more components are considered to be fixed or unrelated (for example, if all objects are considered to be at the same height and have a horizontal orientation, four components may provide a complete representation of the object's pose). In the following, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).
[0076] Many VR applications are based on gestures with the greatest degrees of freedom, i.e., three degrees of freedom for each position and orientation, resulting in a total of six degrees of freedom. A gesture can therefore be represented by a set or vector of six values representing the six degrees of freedom, and thus, a gesture vector can provide a three-dimensional position and / or three-dimensional orientation indication. However, it should be understood that in other embodiments, a gesture can be represented by fewer values.
[0077] The posture may be at least one of an orientation and a position. The posture value may indicate at least one of an orientation value and a position value.
[0078] Systems or entities based on providing the most freedom to the viewer are often referred to as having 6 degrees of freedom (6DoF).Many systems and entities only provide orientation or position, and these are often referred to as having 3 degrees of freedom (3DoF).
[0079] In some systems, VR applications can be provided locally to viewers via a standalone device that does not use, or even access, any remote VR data or processing. For example, a device such as a game console may include memory for storing scene data, input for receiving / generating viewer gestures, and a processor for generating corresponding images based on the scene data.
[0080] In other systems, VR applications can be implemented and executed remotely, away from the viewer. For example, a device local to the user can detect / receive motion / posture data, which is transmitted to a remote device that processes the data to generate the viewer's posture. The remote device can then generate an appropriate view image for the viewer's posture based on scene data describing the scene. The view image is then transmitted to a device local to the viewer, which presents it. For example, the remote device can directly generate a video stream (typically a stereo / 3D video stream) that is directly presented by the local device. Thus, in such a paradigm, the local device may not perform any VR processing other than transmitting the motion data and presenting the received video data.
[0081] In many systems, functionality may be distributed across the local device and the remote device. For example, the local device may process received input and sensor data to generate a viewer pose that is continuously transmitted to the remote VR device. The remote VR device may then generate corresponding view images and transmit those view images to the local device for presentation. In other systems, the remote VR device may not generate the view images directly, but may select relevant scene data and transmit it to the local device, which may then generate the presented view images. For example, the remote VR device may identify the nearest capture point and extract the corresponding scene data (e.g., a spherical image and depth data from the capture point) and transmit it to the local device. The local device may then process the received scene data to generate an image for the specific current viewing pose. A viewing pose typically corresponds to a head pose, and references to a viewing pose can typically be considered equivalent to references to a head pose.
[0082] In many applications, particularly for broadcast services, a source may transmit scene data in the form of an image (including video) representation of the scene that is independent of the viewer's pose. For example, an image representation of a single view sphere for a single capture position may be transmitted to multiple clients. Individual clients may then locally synthesize a view image corresponding to the current viewer's pose.
[0083] One particularly interesting application is supporting a limited amount of movement, so that the rendered view is updated to follow the small movements and rotations of a largely static viewer who makes only small head movements and rotations. For example, a seated viewer can turn their head and move it slightly, and the rendered view / image is adjusted to follow such posture changes. This approach can provide a highly immersive (e.g., video) experience. For example, a viewer watching a sporting event can feel as if they are present at a specific location in the arena.
[0084] This limited degree of freedom in applications has the advantage of providing an improved experience while not requiring accurate scene representations from many different positions, significantly reducing capture requirements. Similarly, the amount of data that needs to be provided to the renderer can be greatly reduced. Indeed, in many scenarios, an image and often depth data for a single viewpoint needs to be provided to a local renderer capable of generating the desired view from it.
[0085] The method may be particularly well-suited for applications requiring the transmission of data from a source to a destination over a bandwidth-limited communication channel, such as broadcast or client-server applications.
[0086] Figure 1 An example of a VR system is shown in which a remote VR client device 101 communicates with a VR server 103 via a network 105, such as the Internet. The server 103 may be arranged to support a potentially large number of client devices 101 simultaneously.
[0087] The VR server 103 may support the broadcast experience, for example, by transmitting an image signal including an image representation in the form of image data that the client device may use to locally synthesize a view image corresponding to the appropriate gesture.
[0088] Figure 2 Example elements of an exemplary implementation of an apparatus for processing a depth map are shown. The apparatus may be specifically implemented in the VR server 103 and will be described in this manner. Figure 3 Shown is the Figure 2 Flowchart of a method for processing a depth map performed by a device.
[0089] The device / VR server 103 includes a receiver 201 that performs step 301 in which a plurality of images representing a scene from different viewing postures and corresponding depth maps are received.
[0090] The image includes light intensity information, and the pixel values of the image reflect the light intensity values. In some examples, the pixel value can be a single value, such as brightness for a grayscale image, but in many embodiments, the pixel value can be a set or vector of (sub) values, such as color channel values for a color image (e.g., RGB or YUV values can be provided).
[0091] A depth map for an image may include depth values for the same viewport. For example, for each pixel of an image for a given view / capture / anchor pose, the corresponding depth map includes a pixel with a depth value. Thus, the same location in the image and its corresponding depth map provide the ray intensity and depth of the ray corresponding to the pixel, respectively. In some embodiments, the depth map may have a lower resolution, and, for example, one depth map pixel may correspond to multiple image pixels. However, in this case, there may still be a direct one-to-one correspondence between a location in the depth map and a location (including sub-pixel locations) in the depth map.
[0092] For the sake of brevity and to avoid complexity, the following description will focus on an example where only three images and corresponding depth maps are provided. It is also assumed that these images are obtained by capturing the scene from three different viewing positions and having Figure 4 A linear arrangement of cameras with the same orientation as shown in is provided.
[0093] It should be understood that in many embodiments, a large number of images are typically received, and the scene is typically captured from an even larger number of capture poses.
[0094] The receiver is fed into a depth map updater, which for brevity is referred to hereinafter as updater 203. Updater 203 performs step 303, in which one or more (typically all) received depth maps are updated. The updating includes updating the depth values of the first received depth map based on the depth values of at least the second received depth map. Thus, cross-depth map and cross-looking pose updating is performed to generate an improved depth map.
[0095] In the example, the updater 203 is coupled to the image signal generator 205, which performs step 305 in which the image signal generator 205 generates an image signal comprising the received image and the updated depth map. The image signal can then be transmitted, for example, to the VR client device 101, where it can be used as a basis for synthesizing a view image for the current viewer pose.
[0096] In the example, the depth map update is thus performed in the VR server 103, and the updated depth map is distributed to the VR client device 101. However, in other embodiments, the depth map update may be performed, for example, in the VR client device 101. For example, the receiver 201 may be part of the VR client device 101 and receive the image and the corresponding depth map from the VR server 103. The received depth map may then be updated by the updater 203, and instead of the image signal generator 205, the apparatus may include a renderer or a view image synthesizer arranged to generate a new view based on the image and the updated depth map.
[0097] In other embodiments, all processing can be performed in a single device. For example, the same device can receive direct capture information and generate an initial depth map, such as through disparity estimation. The generated depth map can be updated, and the device's compositor can dynamically generate new views.
[0098] Therefore, the location of the described functionality and the specific use of the updated depth map will depend on the preferences and requirements of each embodiment.
[0099] The updating of the depth map is accordingly based on one or more other depth maps representing depth from different spatial locations and for different images. The method exploits the recognition that for depth maps, not only the absolute accuracy or reliability of individual depth values or depth maps is important for generating perceptual quality, but also the consistency between different depth maps is very important.
[0100] In fact, the heuristic insight gained is that when the errors or inaccuracies between depth maps are inconsistent, i.e., they vary with the source view, they are considered particularly harmful because they effectively cause the virtual scene to be perceived as shaking violently when the viewer changes position.
[0101] This kind of vision Figure 1 Consistency is not always fully enforced in the depth map estimation process. This is the case, for example, when a separate depth sensor is used to obtain a depth map for each view. In this case, the depth data is captured completely independently. At the other extreme, where all views are used to estimate depth (e.g., using a plane sweep algorithm), the results may still be inconsistent, as the results will depend on the specific multi-view disparity algorithm used and its parameter settings. The specific method described below can alleviate such problems in many scenarios and can update the depth maps to produce improved consistency between depth maps and, therefore, improved perceived image quality. The method can improve the quality of depth maps for multi-view images of a set of scenes.
[0102] Figure 5 A flowchart for performing an update on a pixel of a depth map is shown. The method can be repeated for some or all depth map pixels to generate an updated first depth map. The process can then be repeated for other depth maps.
[0103] The updating of a pixel (hereinafter referred to as a first depth pixel) in a depth map (hereinafter referred to as a first depth map) begins in step 501, where a first candidate depth value is determined for the first depth pixel. The position of the first depth pixel in the first depth map is referred to as the first depth map position. Corresponding terms are used for other views where only the digital label is changed.
[0104] A first candidate depth value is determined in response to at least one second depth value, the second depth value being a depth value of a second depth pixel at a second depth map position in the second depth map. Thus, the first candidate depth value is determined based on one or more depth values of another depth map. The first candidate depth value can specifically be an estimate of a correct depth value for the first depth pixel based on information contained in the second depth map.
[0105] Step 501 is followed by step 503, in which an updated first depth value is determined for the first depth pixel by a weighted combination of a plurality of candidate depth values for the first depth map position.The first candidate depth value determined in step 503 is included in the weighted combination.
[0106] Therefore, in step 501, one of a plurality of candidate depth values for subsequent combination is determined. In most embodiments, the plurality of candidate depth values may be determined in step 501 by repeating the process described for the first candidate depth value for other depth values in the second depth map and / or for depth values in other depth maps.
[0107] In many embodiments, one or more candidate depth values may be determined in other ways or from other sources. In many embodiments, one or more of the candidate depth values may be depth values from the first depth map, such as depth values in the neighborhood of the first depth pixel. In many embodiments, the original first depth value, i.e., the depth value for the first depth pixel in the first depth map received by the receiver 201, may be included as one of the candidate depth values.
[0108] Thus, updater 205 may perform a weighted combination of candidate depth values including at least one candidate depth value determined as described above.The number, attributes, origin, etc. of any other candidate depth values will depend on the preferences and requirements of various embodiments and the exact depth update operation required.
[0109] For example, in some embodiments, the weighted combination may include only the first candidate depth value and the original depth value determined in step 501. In this case, for example, only a single weight may be determined for the first candidate depth value, and the weight for the original depth value may be constant.
[0110] As another example, in some embodiments, the weighted combination can be a combination of a large number of candidate depth values, including values determined from other depth maps and / or positions, the original depth value, depth values in a neighborhood in the first depth map, or indeed even depth values based on alternative depth maps (e.g., depth maps using different depth estimation algorithms.) In such more complex embodiments, a weight can be determined, for example, for each candidate depth value.
[0111] It will be appreciated that any suitable form of weighted combination may be used, including, for example, a non-linear combination or a selective combination (where one candidate depth value is given a weight of 1 and all other candidate depth values are given a weight of 0). However, in many embodiments, a linear combination may be used, in particular a weighted average.
[0112] Thus, as a specific example, the updated depth value for image coordinate (u, v) in depth map / view k is It can be at least one of the i∈{1,…,n} candidate depth values z generated as described in step 501 i In this case, the weighted combination can correspond to a filter function given by:
[0113]
[0114] in, is the updated depth value at pixel position (u, v) for view k, z i is the i-th input candidate depth value, w i is the weight of the i-th input candidate depth value.
[0115] The method uses a specific way to determine the weight of the first candidate depth value (ie, the first weight). Figure 6 Flowchart and Figure 7 The image and depth map are described in the manner described.
[0116] Figure 7 An example is shown in which three images and three corresponding depth maps are provided / considered. A first image 701 is provided together with a first depth map 703. Similarly, a second image 705 is provided together with a second depth map 707, and a third image 709 is provided together with a third depth map 711. The following description will focus on determining a first weight for a first depth value of the first depth map 703 based on the depth value from the second depth map 707 and further considering the third image 709.
[0117] The determination of the first weight (for the first candidate depth value) is therefore determined for the first depth pixel / first depth map position based on one or more second depth values for the second depth pixel at the second depth map position in the second depth map 707. Specifically, Figure 7 As indicated by the middle arrow 713 , the first candidate depth value may be determined as the second depth value at the corresponding position in the second depth map 707 .
[0118] The determination of the first weight begins at step 601, where the updater determines a first image position in the first image 701 corresponding to a first depth map position, as indicated by arrow 715. Typically, this can simply be the same position and image coordinates. The pixel in the first image 701 corresponding to the first image position is referred to as a first image pixel.
[0119] The updater 203 then continues in step 603 to determine a third image position in a third image 709 of the plurality of images based on the first candidate depth value, wherein the third image position corresponds to a projection of the first image position onto the third image. The third image position may be determined from a direct projection of the image coordinates of the first image 701 indicated by arrow 717.
[0120] The updater 203 accordingly proceeds to project the first image position to a third image position in the third image 709. The projection is based on the first candidate depth value. Thus, the projection of the first image position to the third image 709 is based on a depth value that can be considered an estimate of the first depth value determined based on the second depth map 707.
[0121] In some embodiments, the determination of the third image position can be based on a projection of the depth map position. For example, the updater 203 can continue to project the first depth map position (the position of the first depth pixel) to a third depth map position in the third depth map 711 as indicated by arrow 719. The projection is based on the first candidate depth value. Thus, the projection of the first depth map position to the third depth map 711 is based on a depth value that can be considered an estimate of the first depth value determined based on the second depth map 707.
[0122] The third image position may then be determined as the image position in the third image 709 corresponding to the third depth map position as indicated by arrow 721 .
[0123] It should be understood that these two approaches are equivalent.
[0124] Projecting one depth map / image onto a different depth map / image can be the determination of a depth map / image position in a different depth map / image that represents the same scene point as in one depth map / image. Because the depth maps / images represent different viewing / capture poses, parallax effects will result in an offset in the image position for a given point in the scene. The offset will depend on the change in viewing pose and the depth of the point in the scene. Projecting from one image / depth map onto another can also be referred to as an image / depth map position offset or determination.
[0125] As an example, let the image coordinates (u, v) in a view (l) be l and its depth value z l(u, v) is projected to the corresponding image coordinates (u, v) of the adjacent view (k) k This can be done, for example, for a perspective camera by the following steps:
[0126] 1. Image coordinates (u, v) l Not used l Projected in 3d space (x,y,z) of camera intrinsic parameters (focal length and principal point) for camera (l) l middle.
[0127] 2. Using their relative extrinsic parameters (camera rotation matrix R and translation vector t), the unprojected point (x, y, z) in the coordinate system of the camera (l) l is transformed into the coordinate system (x, y, z)k of camera (k).
[0128] 3. Final point (x, y, z) k Projected onto the image plane of camera (k) (using the camera intrinsics of k), resulting in image coordinates (u, v) k .
[0129] A similar mechanism can be used for other camera projection types, such as equirectangular projection (ERP).
[0130] In the manner described, the projection based on the first candidate depth value can be considered to correspond to determining a third depth map / image position of a scene point for the first depth map / image position having a depth of the first candidate depth value (and for changes in viewing pose between the first and third viewing poses).
[0131] Different depths will produce different offsets, and in the present case, the offset in image and depth map positions between the first viewing pose for the first depth map 703 and the first image 701 and the third viewing pose for the third depth map 711 and the third image 709 is based on at least one depth value in the second depth map 707.
[0132] In step 603, the updater 203 determines a location in the third depth map 711 and the third image 709, respectively, which, if the first candidate depth value is indeed the correct value for the first depth value and the first image pixel, will reflect the same scene point as the first image pixel in the first image 701. Any deviation of the first candidate depth value from the correct value may result in an incorrect location being determined in the third image 709. It should be noted that the scene points here refer to scene points on the ray associated with the pixel, but they may not necessarily be the foremost scene points for both viewing poses. For example, if a scene point seen from the first viewing pose is occluded by a (more) foreground object seen from the second viewing pose, the depth map and the depth values of the image may represent different scene points and therefore have potentially very different values.
[0133] Step 603 is followed by step 605, in which a first match error indication is generated based on the content of the first and third images 701, 709 at the first image location and the third image location, respectively. Specifically, an image pixel value of the third image at the third image location is retrieved. In some embodiments, the image pixel value can be determined as an image pixel value in the third image 709 for which the third depth map location in the third depth map 711 provides the determined depth value. It will be appreciated that in many embodiments, i.e., where the same resolution is used for the third depth map 711 and the third image 709, directly determining the location in the third image 709 corresponding to the first depth map location (arrow 719) is equivalent to determining the location in the third depth map 711 and retrieving the corresponding image pixel.
[0134] Similarly, updater 203 continues to extract the pixel value in first image 701 at the first image position. It then continues to determine a first matching error indication indicating the difference between the pixel values of the two images. It will be appreciated that any suitable difference metric may be used, such as a simple absolute difference, a root sum square difference applied to, for example, pixel value components of multiple color channels, and the like.
[0135] Hence, the updater 203 determines 605 a first matching error indication indicating a difference between an image pixel value in the third image for the third image location and an image pixel value in the first image for the first image location.
[0136] Updater 203 then proceeds to step 607, where a first weight is determined in response to the first matching error indication. It should be understood that the specific manner in which the first weight is determined based on the first matching error indication may depend on various embodiments. In many embodiments, complex considerations including, for example, other matching error indications may be used, and more examples will be provided later.
[0137] As a low complexity example, in some embodiments the first weight may be determined as a monotonically decreasing function of the first matching error indication, and in many embodiments without considering any other parameters.
[0138] For example, in an example where the weighted combination includes only the first candidate depth value and the original depth value of the first depth pixel, the combination may apply a fixed weight to the original depth value, with the first weight increasing (typically also including weight normalization) the lower the first matching error indication.
[0139] The first match error indication can be considered to reflect how closely the first and third images match in representing a given scene point. If there are no occlusion differences between the first and third images, and if the first candidate depth value is correct, the image pixel values should be identical and the first match error indication should be zero. If the first candidate depth value deviates from the correct value, the image pixels in the third image may not directly correspond to the same scene point, and the first match error indication may increase. If occlusion changes, the error may be significantly higher. Therefore, the first match error indication can provide a good indication of the accuracy and suitability of the first candidate depth value for the first depth pixel.
[0140] In different embodiments, different approaches can be used to determine the first candidate depth value from one or more depth values of the second depth map. Similarly, different approaches can be used to determine which candidate values to generate for the weighted combination. Specifically, multiple candidate values can be generated from the depth values of the second depth map, and the candidate values can be selected based on the depth information of the second depth map. Figure 6 The described approach calculates weights individually for each of these candidate values.
[0141] In many embodiments, determining which second depth values to use to derive the first candidate depth value depends on a projection between the first depth map and the second depth map, thereby determining corresponding positions in the two depth maps. Specifically, in many embodiments, the first candidate depth value can be determined as the second depth value at a second depth map position that is considered to correspond to the first depth map position, i.e., the second depth value is selected as the depth value that is considered to represent the same scene point.
[0142] Determining the corresponding first and second depth map positions can be based on a projection from the first depth map to the second depth map, i.e., can be based on the original first depth value, or it can be based on a projection from the second depth map to the first depth map, i.e., can be based on the second depth value. In some embodiments, projections in two directions can be performed and, for example, an average of these projections can be used.
[0143] Thus, determining the first candidate depth value may include determining a second depth map position relative to the first depth map position by projecting between a first viewing pose of the first image and a second viewing pose of the second image based on the second value and at least one of the first original depth value of the first depth map.
[0144] For example, for a first pixel in a given first depth map, updater 203 may extract a depth value and use it to project the corresponding first depth map position to a corresponding second depth map position in the second depth map. It may then extract a second depth value at that position and use it as the first candidate depth value.
[0145] As another example, for a second pixel in a given second depth map, updater 203 may extract a depth value and use it to project the corresponding second depth map position to the corresponding first depth map position in the first depth map. It may then extract the second depth value and use it as a first candidate depth value for the first depth pixel at the first depth map position.
[0146] In such an embodiment, the depth value in the second depth map is used directly as the first candidate depth value. However, since the two depth map pixels represent (in the absence of occlusions) the distance to the same scene point but from different viewpoints, the depth values may be different. In many practical embodiments, this difference in distance to the same scene point from cameras at different positions / viewing poses is negligible and can be ignored. Therefore, in many embodiments, it can be assumed that the cameras are perfectly aligned and looking in the same direction and have the same position. In that case, if the object is flat and parallel to the image sensor, the depth may indeed be exactly the same in the two corresponding depth maps. Deviations from this situation are typically small enough to be negligible.
[0147] However, in some embodiments, determining the first candidate depth value from the second depth value may include modifying the projection of the depth value. This may be based on a more detailed geometric calculation, including taking into account the projective geometry of the two views.
[0148] In some embodiments, more than a single second depth value may be used to generate the first candidate depth value. For example, spatial interpolation may be performed between different depth values to compensate for projections that are not aligned with the center of a pixel.
[0149] As another example, in some embodiments, the first candidate depth value may be determined as a result of spatial filtering, where a kernel centered at a second depth map location is applied to the second depth map.
[0150] The following description will focus on embodiments in which each candidate depth value depends only on a single second depth value and is also equal to the second depth value.
[0151] In many embodiments, the weighted combination may also include multiple candidate depth values determined from different second depth values.
[0152] Specifically, in many embodiments, the weighted combination may include candidate depth values for a region of the second depth map. The region may generally be determined based on the first depth map position. Specifically, the second depth map position may be determined by projection (in either or both directions) as described above, and the region may be determined as a region around the second depth map position (e.g., having a predetermined contour).
[0153] This approach can provide a set of candidate depth values for the first depth pixel in the first depth map. For each candidate depth value, the updater 203 can execute Figure 6 Method to determine the weights for a weighted combination.
[0154] A particular advantage of this approach is that the selection of the second depth value for the candidate depth values is not overly important, as the subsequent weight determination will appropriately weigh good and bad candidates.Thus, in many embodiments, a relatively low complexity approach can be used to select candidate values.
[0155] In many embodiments, the region can be determined simply as a predetermined region around a position in the second depth map determined by a projection from the first depth map to the second depth map based on the original first depth value. In fact, in many embodiments, the projection can even be replaced by simply selecting the region around the same depth map position in the second depth map as in the first depth map. Thus, this approach allows for the candidate set of depth values to be selected simply by selecting the second depth value in the region around the same position in the second depth map as the first pixel in the first depth map.
[0156] This approach may in practice reduce resource usage while providing efficient operation.This approach may be particularly suitable when the size of the region is relatively large compared to the position / disparity offset that occurs between the depth maps.
[0157] As previously mentioned, many different methods may be used to determine the weights for the various candidate depth values in the weighted combination.
[0158] In many embodiments, the first weight can also be determined in response to additional match error indicators determined for other images other than the third image. In many embodiments, the described method can be used to generate match error indicators for all other images other than the first image. A combined match error indicator can then be generated, for example, as an average of these match error indicators, and the first weight can be determined based on this.
[0159] Specifically, the first weight may depend on a matching error indicator that is a function of the independent matching errors from the view filtered to all other views l≠k. i An example indicator of the weight is:
[0160] w i (z i )=min l≠k (e kl (z i )),
[0161] Among them, e kl (z i ) is a given candidate Z i The matching error between views k and l is the matching error between l and l. The matching error can, for example, depend on the color difference for a single pixel, or can be calculated as a spatial average around the pixel position (u, v). Instead of calculating the minimum matching error for views l≠k, the average or median value can be used, for example. In many embodiments, the evaluation function can preferably be robust to match error outliers caused by occlusions.
[0162] In many embodiments, a second match error indication may be determined for the second image (i.e., for the view from which the first candidate depth value was generated). The second match error indication may be determined using the same method as described for the first match error indication, and the second match error indication may be generated to indicate a difference between an image pixel value in the second image for the second depth map position and an image pixel value in the first image for the first depth map position.
[0163] A first weight may then be determined in response to the first match error indication and the second match error indication (and possibly other match error indications or parameters).
[0164] In some embodiments, such weight determination may take into account not only, for example, the average match error indication, but also the relative differences between the match indications. For example, if the first match error indication is relatively low and the second match error indication is relatively high, this may be due to occlusion occurring in the second image relative to the first image (but not in the third image). Therefore, the first weight may be reduced or even set to zero.
[0165] Other examples of weighting considerations could be using statistical measures such as the median matching error or other quantiles. Similar reasoning applies here. For example, if we have a linear camera array of nine cameras, all looking in the same direction, we can assume that the central camera will always be looking at unoccluded areas around the edge of the object, four anchors to the left, or four anchors to the right. In this case, the overall weight for a candidate's well-being could simply be a function of the lowest four of the eight total matching errors.
[0166] In many embodiments, the weighted combination can include other depth values from the first depth map itself. Specifically, a set of depth pixels in the first depth map surrounding the first depth position can be included in the weighted combination. For example, a predetermined spatial kernel can be applied to the first depth map, resulting in a low-pass filter of the first depth map. The weighting of the spatially low-pass filtered first depth map values and the candidate depth values from other views can then be adjusted, for example by applying a fixed weight to the low-pass filtered depth values and a variable first weight to the first candidate depth values.
[0167] In many embodiments, the determination of the weights, in particular the determination of the first weights, also depends on the confidence value for the depth value.
[0168] Depth estimation and measurement are inherently noisy and subject to various errors and variations. In addition to depth estimation, many depth estimation and measurement algorithms can also generate a confidence value that indicates how reliable the provided depth estimate is. For example, disparity estimation can be based on detecting matching regions in different images, and a confidence value can be generated to reflect how similar the matching regions are.
[0169] The confidence values can be used in different ways. For example, in many embodiments, a first weight for a first candidate depth value can depend on the confidence value for the first candidate depth value, and specifically the confidence value of the second depth value used to generate the first candidate depth value. The first weight can be a monotonically increasing function of the confidence value for the second depth value, and thus the first weight can increase in order to increase the confidence of the base depth value used to generate the first candidate depth value. Thus, the weighted combination can be biased towards depth values that are considered reliable and accurate.
[0170] In some embodiments, the confidence value for the depth map may be used to select which depth values / pixels to update and for which depth pixels to keep the depth values unchanged. Specifically, the updater 203 may be arranged to select only the depth values / pixels of the first depth map whose confidence value is below a threshold for updating.
[0171] Therefore, instead of updating all pixels in the first depth map, the updater 203 specifically identifies depth values that are considered unreliable and updates only those values. This can result in an improved overall depth map in many embodiments because, for example, a very accurate and reliable depth estimate can be prevented from being replaced by more uncertain values generated from depth values from other viewpoints.
[0172] In some embodiments, the set of depth values of the second depth map included in the weighted combination, either by contribution to different candidate depth values or contribution to the same candidate depth value, may depend on a confidence value for the depth value. Specifically, only depth values having a confidence value above a given threshold may be included, and all other depth values may be discarded from processing.
[0173] For example, the updater 203 may initially generate a modified second depth map by scanning the second depth map and removing all depth values with confidence values below a threshold. The previously described processing may then be performed using the modified second depth map, wherein all operations requiring the second depth value are bypassed if such a second depth value does not exist in the second depth map. For example, if the second depth value does not exist, a candidate depth value for the second depth value is not generated.
[0174] In some embodiments, the updater 203 may also be arranged to generate a confidence value for the depth value.
[0175] In some embodiments, a confidence value for a given depth value in a given depth map may be determined in response to changes in depth values in other depth maps for corresponding positions in those depth maps.
[0176] The updater 203 may first project the depth map position for a given depth value for which a confidence value was determined to a corresponding position in a plurality of other depth maps, and typically to all of these positions.
[0177] Specifically, for the image coordinate (u,v) in the depth map k k Given a depth value at , a set L of other depth maps (usually for neighboring views) is determined. For each of these depth maps (l∈L), the corresponding image coordinates (u,v) for l∈L are computed by reprojecting l .
[0178] The updater 203 may then consider the depth values in these other depth maps at these corresponding positions. A variation measure for these depth values at the corresponding positions may be determined. Any suitable variation measure may be used, such as a variance measure.
[0179] The updater 203 may then proceed to determine a confidence value for a given depth map position from this measure of variation, and in particular, an increasing degree of variation may indicate a decreasing confidence value.Thus, the confidence value may be a monotonically decreasing function of the measure of variation.
[0180] Specifically, for l∈L, given a depth value z k and at (u, v) l The set of adjacent depth values corresponding to z l , a confidence metric can be calculated based on the consistency of these depth values. For example, the variance of these depth values can be used as a confidence metric. Low variance means high confidence.
[0181] It is often desirable to make this determination accurate for corresponding image coordinates (u, v) that may be potentially occluded by objects in the scene or by the camera boundaries. k The resulting outliers are more robust. One specific way to achieve this is to select two adjacent views l0 and l1 on opposite sides of the camera view (k) and use the minimum of the depth difference
[0182]
[0183] In some embodiments, a confidence value for a given depth value in a given depth map can be determined by evaluating the error resulting from projecting the corresponding given depth position into another depth map and then projecting it back using the two depth values at the two depths.
[0184] Therefore, the updater 203 can first project a given depth map position to another depth map based on a given depth value. The depth value at the projected position is then retrieved and the position in the other depth map is projected back to the original depth map based on the other depth value. This produces a test position that is indeed the same as the original depth map position if the two depth values used for the projection match perfectly (e.g., taking into account camera and capture properties and geometry). However, any noise or error will produce a difference between the two positions.
[0185] Updater 203 may correspondingly continue to determine a confidence value for a given depth map position in response to the distance between the given depth map position and the test position. The smaller the distance, the higher the confidence value, and thus the confidence value may be determined as a monotonically decreasing function of the distance. In many embodiments, multiple other depth maps may be considered to account for the distance.
[0186] Therefore, in some embodiments, the confidence value may be determined based on the geometric consistency of the motion vector. kl represents a 2D motion vector, which will give its depth z k Pixel (u, v) kBring to the adjacent view l. Each corresponding pixel position (u, v) in the adjacent view l l Each has its own depth z l , which produces a vector d returned to view k lk In the ideal case of zero error, all of these vectors map exactly back to the original point (u, v) k However, this is not the case in general, and certainly not for regions with insufficient confidence. Therefore, a good metric for lack of confidence is the average error in the back-projected positions. This error metric can be expressed as:
[0187]
[0188] Among them, f((u,v) l ;z l ) indicates the use of depth value z l The image coordinates in view k back-projected from the adjacent view l. The norm ‖·‖ can be L1 or L2 or any other norm. The confidence value can be determined as a monotonically decreasing function of this value. It should be understood that the term "candidate" does not imply any limitation on depth values, and the term candidate depth value can refer to any depth value included in the weighted combination.
[0189] It should be understood that for clarity, the above description has described embodiments of the present invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without departing from the present invention. For example, functions illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Therefore, references to specific functional units or circuits are merely to be considered as references to suitable means for providing the described functionality, and do not indicate a strict logical or physical structure or organization.
[0190] The present invention can be implemented in any suitable form, including hardware, software, firmware or any combination of these. The present invention can alternatively be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be implemented physically, functionally and logically in any suitable manner. In fact, the function can be implemented in a single unit, in multiple units or as a part for other functional units. Therefore, the present invention can be implemented in a single unit, or can be distributed between different units, circuits and processors physically and functionally.
[0191] Although the present invention has been described in conjunction with certain embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, although features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0192] Furthermore, although listed separately, multiple devices, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, these may be advantageously combined, and inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one claim category does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories, where appropriate. Furthermore, the order of features in the claims does not imply that the features must be performed in any specific order. Specifically, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps may be performed in any suitable order. Furthermore, singular references do not exclude plural references. Thus, references to "a," "an," "first," "second," etc. do not exclude a plurality. The terms "first," "second," "third," etc., are used as labels and therefore do not imply any other limitation other than providing clear identification of the corresponding features, and should not be construed as limiting the scope of the claims in any way. Reference numerals in the claims are provided merely as illustrative examples and should not be construed as limiting the scope of the claims in any way.
Claims
1. A method for processing a depth map, the method comprising: receiving (301) from a receiving device a plurality of images representing a scene from different viewing postures and corresponding depth maps; Updating (303) depth values of a first depth map of the corresponding depth maps based on depth values of at least a second depth map of the corresponding depth maps, the first depth map being for a first image and the second depth map being for a second image; the updating (303) comprising: determining (501) a first candidate depth value for a first depth pixel of the first depth map at a first depth map position in the first depth map, the first candidate depth value being determined in response to at least one second depth value for a second depth pixel of the second depth map at a second depth map position in the second depth map; determining (503) a first depth value for the first depth pixel by weighted combining a plurality of candidate depth values for the first depth map location, the weighted combination comprising the first candidate depth value weighted by a first weight; Wherein, determining (503) the first depth value comprises: determining (601) a first image position in the first image for the first depth map position, determining (603) a third image position in a third image of the plurality of images, the third image position corresponding to a projection of the first image position onto the third image; determining (605) a first matching error indication indicating a difference between an image pixel value in the third image for the third image location and an image pixel value in the first image for the first image location, and The first weight is determined (607) in response to the first match error indication.
2. The method according to claim 1, wherein Determining (501) a first candidate depth value includes determining, based on a second depth value and at least one of a first original depth value of the first depth map, a second depth map position relative to the first depth map position by projection between a first viewing pose of the first image and a second viewing pose of the second image.
3. The method according to claim 1 or 2, wherein The weighted combination includes candidate depth values determined from a region of the second depth map determined responsive to the first depth map location.
4. The method according to claim 3, wherein: The area of the second depth map is determined to be an area around the second depth map position, and the second depth map position is determined to be a depth map position in the second depth map that is equal to the first depth map position in the first depth map.
5. The method according to claim 3, wherein The area of the second depth map is determined as an area around a location in the second depth map determined by projection from the first depth map location based on an original depth value in the first depth map at the first depth map location.
6. The method of claim 1 , further comprising determining a second match error indication, the second match error indication indicating a difference between an image pixel value in the second image for the second depth map position and the image pixel value in the first image for the first depth map position; and wherein Determining the first weight is also responsive to the second match error indication.
7. The method of claim 1 , further comprising determining an additional matching error indication indicating a difference between image pixel values in other images for depth map positions corresponding to the first depth map position and the image pixel values in the first image for the first depth map position; and wherein Determining the first weight is also responsive to the additional match error indication.
8. The method according to claim 1 or 2, wherein: The weighted combination includes depth values of the first depth map in an area surrounding the first depth map location.
9. The method according to claim 1 or 2, wherein: The first weight depends on a confidence value of the first candidate depth value.
10. The method according to claim 9, wherein: Only the depth values of the first depth map having confidence values lower than a threshold are updated.
11. The method of claim 9, further comprising selecting a set of depth values of the second depth map to include in the weighted combination to satisfy a requirement that depth values of the set of depth values must have a confidence value above a threshold.
12. The method according to claim 9, further comprising: projecting a given depth map position for a given depth value in a given depth map to a corresponding position in a plurality of corresponding depth maps; determining a change metric for a set of depth values, the set of depth values comprising the given depth value and depth values at the corresponding location in the plurality of corresponding depth maps; and A confidence value for the given depth map position is determined responsive to the variation metric.
13. The method according to claim 9, further comprising: projecting a given depth map position for a given depth value in a given depth map to a corresponding position in another depth map, the projection being based on the given depth value; projecting the corresponding position in the other depth map to a test position in the given depth map, the projection being based on a depth value at the corresponding position in the other depth map; A confidence value for the given depth map position is determined responsive to a distance between the given depth map position and the test position.
14. An apparatus for processing a depth map, the apparatus comprising: a receiver (201) for receiving (301) a plurality of images representing a scene from different viewing poses and corresponding depth maps; an updater (203) for updating (303) depth values of a first depth map of the corresponding depth map based on depth values of at least a second depth map of the corresponding depth map, the first depth map being for a first image and the second depth map being for a second image; the updating (303) comprising: determining (501) a first candidate depth value for a first depth pixel of the first depth map at a first depth map position in the first depth map, the first candidate depth value being determined in response to at least one second depth value for a second depth pixel of the second depth map at a second depth map position in the second depth map; determining (503) a first depth value for the first depth pixel by weighted combining a plurality of candidate depth values for the first depth map location, the weighted combination comprising the first candidate depth value weighted by a first weight; Wherein, determining (503) the first depth value comprises: determining (601) a first image position in the first image for the first depth map position, determining (603) a third image position in a third image of the plurality of images, the third image position corresponding to a projection of the first image position onto the third image; determining (605) a first matching error indication indicating a difference between an image pixel value in the third image for the third image location and an image pixel value in the first image for the first image location, and The first weight is determined (607) in response to the first match error indication.
15. A computer program product comprising computer program code means adapted to perform all the steps according to claims 1-13, when said program is run on a computer.
Citation Information
Patent Citations
Fast general multipath correction in time-of-flight imaging
CN105899969A
Depth estimation for an image
EP3418975A1