Dealing with blur in multi-view imaging

JP2024538038A5Pending Publication Date: 2025-10-14KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024521326
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-20
Filing Date
2022-10-12
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Motion blur in multi-view imaging causes inaccuracies in depth estimation and mispositioning of fast-moving objects, leading to erroneous depth and texture data when combining images from different cameras.

Method used

A method involving determining a sharpness metric for each image, generating a confidence score based on sharpness, and using these scores to weight pixel blending during image synthesis, thereby addressing motion blur and improving depth estimation accuracy.

Benefits of technology

The method effectively reduces the impact of motion blur by weighting images based on their sharpness, resulting in more accurate depth estimation and clearer synthesis of new virtual images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for processing multi-view data of a scene, the method including obtaining at least two images of the scene from different cameras, determining a sharpness metric for each image, and determining a confidence score for each image based on the sharpness metric, the confidence score for use in determining weights when blending the images to synthesize a new virtual image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of multi-view imaging, in particular to processing multi-view data of a scene. [Background technology]

[0002] Multi-camera video capture, depth estimation, and view interpolation enable applications such as augmented and virtual reality playback.

[0003] Shooting fast-moving objects results in motion blur, which can cause multiple problems when synthesizing images from various source views obtained from different cameras. For example, when performing multi-view depth estimation, the algorithm has difficulty determining the correct depth of blurred foreground objects due to their semi-transparent appearance and lack of texture. Then, when the estimated depth map is used to synthesize new images from source views, a single fast-moving object may be mispositioned and potentially visible in multiple image locations due to the fact that one or more source views estimated an incorrect depth for the object. Summary of the Invention [Problem to be solved by the invention]

[0004] Therefore, there is a need for a method of dealing with motion blur when processing multi-view data, and more generally, a need to deal with blur in multi-view data.

[0005] US Patent Application Publication No. 2012 / 212651 discloses an image processing apparatus comprising an acquisition unit for capturing an object from different viewpoints, an identification unit configured to identify defect images, and a determination unit configured to determine a weight for each captured image for image synthesis based on the defect images. [Means for solving the problem]

[0006] The invention is defined by the claims.

[0007] According to an example embodiment of the present invention, there is provided a method for processing multi-view data of a scene, the method comprising: The method comprises the steps of acquiring at least two images of a scene from different cameras, determining a sharpness index for each image, and determining a confidence score for each image based on the sharpness index, the confidence score being for use in determining weights when blending the images to synthesize a new virtual image.

[0008] The method can solve the problem of erroneous or inaccurate depth and / or texture data caused by differences between cameras used to capture images. In particular, differences between cameras can cause different degrees of blur in the images. A major difference between cameras is the pose of each camera. For example, different poses can cause motion blur that affects different images more or less. The method is suitable for handling motion blur.

[0009] Motion blur is the apparent streaking or blurring of moving objects in an image. It results from an object moving from a first position to a second position during a single exposure time. Faster object movements or longer exposure times can result in significant blurring.

[0010] Multi-view data of a scene typically requires data of the scene (e.g., depth maps and images) acquired by several different cameras. If there are moving objects in the scene, the amount of motion blur depends on the position of the camera in the scene (and therefore the viewpoint of the image / depth map). For example, a camera capturing an object moving away from the camera will not have much motion blur, but an object moving horizontally across the camera's viewpoint may cause a significant amount of motion blur.

[0011] Therefore, we propose to use a sharpness index for images from different viewpoints to generate a confidence score for the image. The sharpness index provides an indication of the amount of blur in the image, e.g., the blur is motion blur. The sharper the image, the less blur there is in the image.

[0012] The sharpness index may include a measure of sharpness, which is a measure of the clarity of an image with respect to both focus and contrast. The sharpness index may include a determination of the contrast of a corresponding image.

[0013] A sharpness index may be determined for each image, and the sharpness indexes may be compared for each image. Based on the sharpness index, a confidence score may be generated for each image. The confidence score may be a value between 0 and 1, where 0 is unreliable and 1 is reliable. In general, an image having a high sharpness index (i.e., exhibiting a small amount of blur) has a higher confidence score compared to an image having a low sharpness index (i.e., exhibiting a large amount of blur).

[0014] The confidence score may be directly proportional to the sharpness. The confidence score may be 0 for values ​​of the sharpness index below a cutoff threshold (i.e., confidence=0 for particularly blurry objects). The confidence score can be limited to a number of discrete confidence values, each of which is defined by a different sharpness index (e.g., dividing the sharpness index into eight ranges for eight discrete confidence values).

[0015] The confidence score may be determined by comparing the sharpness indexes at each comparison viewpoint and determining the confidence score based on the corresponding sharpness indexes and the comparison. For example, for two images with relative sharpness indexes of "medium" and "low" for the first and second images, respectively, the confidence score for the first image is higher than the confidence score for the second image. However, if the first image had a relative sharpness index of "high", the confidence score for the first image may be higher than the previous example.

[0016] The confidence scores may also be used to identify (and further reduce the effect of) potential sources of artifacts that may occur when combining images with corresponding images, such as those that may arise due to images being acquired with different sharpness or different focal lengths, as well as due to motion blur (as discussed above).

[0017] In general, the difference in sharpness of each image is used to determine a confidence score for each image. Each image (obtained at a different viewpoint) may have a different level of sharpness for a variety of different reasons (e.g., different focal lengths, accidental defocus, motion blur, post-processing blur, etc.). Thus, the present invention generally provides a method for handling blur for multi-view data processing.

[0018] The sharpness index may be a sharpness map that includes a number of sharpness values ​​each corresponding to one or more pixels of a corresponding image.

[0019] The confidence score may include a map of confidence values, each confidence value corresponding to one or more pixels and / or one or more sharpness values ​​of the image. The confidence values ​​may be determined based on the sharpness values ​​of the corresponding pixels. The higher the sharpness value, the higher the confidence value. In other words, pixels with high sharpness are given a high confidence value and pixel groups with low sharpness are given a low confidence value. Each pixel in the image may have a corresponding sharpness value and confidence value. In other words, the confidence score may be a confidence map.

[0020] The determination of the sharpness value for one or more pixels may depend on one or more neighboring sharpness values.

[0021] For example, motion blur typically affects areas larger than just one pixel. Therefore, in some cases, it may be advantageous to set a low confidence value for the pixel being directly evaluated, and also for neighboring pixels. This may result in a gradient of confidence values ​​in the confidence map, and a reduction in outlier confidence values. This is advantageous because a single pixel (or small group of pixels) is unlikely to be blurred if the neighboring pixels are not blurred.

[0022] The method may further include the steps of obtaining at least one depth map of the scene, warping at least two images to a target viewpoint based on the at least one depth map, and blending the images at the target viewpoint to generate a composite image, wherein during blending, each pixel in the images is weighted based on a corresponding confidence score.

[0023] The method may further comprise the steps of obtaining at least one depth map of the scene, warping the at least one image to at least one image comparison viewpoint using the at least one depth map such that there are at least two images at each image comparison viewpoint, and comparing pixel color values ​​of the images at each comparison viewpoint, wherein determining a confidence score for each image is further based on the comparison of the pixel color values.

[0024] The sharpness metric of an image also corresponds to any warped image obtained by warping an image. The warped image contains the same pixel color values ​​as the original image warped to a different viewpoint (i.e., a comparison viewpoint). Similarly, the sharpness metric of a warped image also corresponds to the image warped to generate the warped image.

[0025] In a multi-view camera setup, the sensors (e.g., cameras, depth sensors, etc.) are placed at different locations in the scene and may therefore have different amounts of motion blur. This also means that some images may have the expected color values ​​of an object while some may not (i.e., the object appears blurry). Therefore, the images can be warped to a comparison viewpoint and compared. The images are warped to a comparison viewpoint so that the color values ​​of the images can be compared from a common viewpoint (i.e., the texture data of one image can be compared to the corresponding texture data of another image).

[0026] Each image is warped by using a depth map at the viewpoint of that image. A depth map can be acquired at each image viewpoint, or the method can acquire at least one depth map and warp it to each image viewpoint. Each pixel of the image can be given a 3D coordinate in the virtual scene by using the pose (position and orientation) of the viewpoint and the depth for each pixel (obtained from the depth map at the image viewpoint). Then, the pixel is warped to different viewpoints by projecting the 3D coordinates onto the imaging plane from different poses in the virtual scene.

[0027] To further determine the confidence score, corresponding pixel color values ​​are compared in the image comparison viewpoints. One way to compare pixel colors is to determine the average color of a location for all views (e.g., averaging all pixels corresponding to a location), and the greater the difference between the pixel color and the average from all views, the lower the corresponding confidence value.

[0028] Thus, if one of the images is blurred and has background pixel values ​​where an object is supposed to be, the color values ​​of the blurred areas will have a low confidence value in the confidence score.

[0029] The image comparison viewpoints may include or consist of all viewpoints of the images.

[0030] Alternatively, the comparison viewpoints may consist of a subset of one or more of these viewpoints.

[0031] The images, depth maps and corresponding confidence scores can be transmitted, for example, to a client device that wishes to synthesize a new image based on the images.

[0032] In this case, the confidence score determination can be done before any rendering / compositing of the new image, e.g., it can be performed on a server (or encoder) where there is more available processing power, and then sent to the client to render the new image using the transmitted images, depth map and confidence scores.

[0033] Synthesis of an image from multi-view data can be based on blending pixel values ​​of different images of the multi-view data. If different images show a fast moving object with different amounts of motion blur, the synthesized image may show the fast moving object at the wrong position in the scene, or at multiple positions in the scene.

[0034] Therefore, it is proposed to use a confidence score to weight each pixel in an image during blending of the images to generate a composite image. For pixels relating to regions in an image that have a large amount of motion blur, the corresponding depth map is likely to have a low confidence score (e.g., for that particular region). Thus, pixels that show significant motion blur are weighted lower compared to other corresponding pixels that do not have motion blur during blending, and the composite view is less likely to show fast moving objects at different positions (or multiple positions).

[0035] The color value of a pixel in the composite image can be the following formula, where n is the number of depth maps, C is the color value of the corresponding pixel in the warped depth map, and W is the corresponding numerical confidence score:

number

[0036] The image comparison viewpoint may be a target viewpoint, and the method may further include blending the images at the target viewpoint to generate a composite image, where during blending each pixel in the images is weighted based on a corresponding confidence score.

[0037] The determination of the confidence score can be done during the synthesis of the new image. The images may need to be warped during synthesis, and also need to be warped for comparison. Therefore, both warping steps can be combined, and when the images are warped to synthesize the new image, confidence scores can be determined for use in the blending (i.e., for use as an index of pixel weights) before they are blended.

[0038] The method may further include the steps of obtaining at least two depth maps obtained from different sensors or generated from at least one different image of the scene, warping the at least one depth map to at least one depth comparison viewpoint such that there are at least two depth maps at each image comparison viewpoint, comparing the depth maps at each depth comparison viewpoint, and determining a confidence score for each depth map based on the comparison of the depth maps.

[0039] The depth comparison viewpoint can be one or more of the viewpoints of the depth map. In this case, each depth map can be warped to one or more viewpoints of the other depth map(s) and compared at each of the depth comparison viewpoints. Alternatively, the depth comparison viewpoint can be one or more of the image comparison viewpoints. The depth comparison viewpoint may be only the target viewpoint.

[0040] Depth maps from different viewpoints have different depth values ​​for objects based on the amount of motion blur. For example, a depth map may be generated based on performing depth estimation on two images. If the two images have a significant amount of motion blur, the object (or a portion of the object) may appear semi-transparent in the image, and the texture values ​​of pixels corresponding to the object may be based in part on background texture values. Thus, the estimated depth map may have background depth values ​​for locations corresponding to the object.

[0041] Similarly, a depth map acquired using a depth sensor may have erroneous depth values ​​for objects due to the movement of the objects during the capture / exposure time of the depth sensor.

[0042] Thus, the confidence score may further indicate the likelihood that the depth values ​​for a particular region of the depth map are correct. For example, in the case of a depth map estimated from two images with a significant amount of motion blur, the depth map may only have depth data for the background (not the object). When the depth map is warped and compared with other warped depth maps that have correct depth values ​​for the object, a low confidence value may be given in the corresponding region of the confidence map where the object is expected to be present (i.e., based on the other depth map). If two (or more) warped depth maps have the same (or similar) depth data, a high confidence value may be given in the corresponding region of the confidence map.

[0043] A depth map that is likely to have the correct depth for the foreground object (e.g., by checking for high confidence in the corresponding depth confidence map) may be warped to another viewpoint to determine whether the depth map of the other viewpoint is expected to have an incorrect depth by comparing it with the depth map of the other view. In other words, the depth map of a first viewpoint may be warped to a second viewpoint, and the predicted depth at the second viewpoint is compared to help determine whether the observed depth of the second viewpoint is correct. If not, it may be flagged in the target view and given a low confidence score.

[0044] The method may further include the steps of obtaining at least two depth confidence maps corresponding to the depth map, and warping at least one depth confidence map to at least one depth comparison viewpoint having a corresponding depth map, and the step of comparing the depth maps at each depth comparison viewpoint further includes comparing the corresponding depth confidence maps.

[0045] In some cases, each depth map can be associated with a depth confidence map generated during generation of the corresponding depth map, which can provide an indication of the likelihood that each depth value is accurate. The indication of the likelihood for a particular depth value can be based on available data (e.g., images) used during the estimation of the depth value.

[0046] The present invention also provides a computer program product comprising computer program code which, when executed on a computing device having a processing system, causes the processing system to perform all of the steps of a method for processing multi-view data of a scene.

[0047] The invention also provides a system for processing multi-view data of a scene, the system comprising: The method includes a processor configured to obtain at least two images of a scene from different cameras, determine a sharpness index for each image, and determine a confidence score for each image based on the sharpness index, the confidence scores for use in determining weights when blending the images to synthesize a new virtual image.

[0048] The processor may be further configured to obtain at least one depth map of the scene, warp at least two of the images to a target viewpoint based on the at least one depth map, and blend the images at the target viewpoint to generate a composite image, wherein during blending, each pixel in the images is weighted based on a corresponding confidence score.

[0049] The processor may be further configured to obtain at least one depth map of the scene, warp at least one image to at least one image comparison viewpoint using the at least one depth map such that there are at least two images at each image comparison viewpoint, and compare pixel color values ​​of the images at each comparison viewpoint, wherein determining the confidence score for each image is further based on the comparison of the pixel color values.

[0050] The image comparison viewpoints may include or consist of all viewpoints of the images.

[0051] The image comparison viewpoint may be a target viewpoint, and the processor may be further configured to blend the images at the target viewpoint to generate a composite image, wherein during blending, each pixel in the images is weighted based on a corresponding confidence score.

[0052] The processor may be further configured to acquire at least two depth maps acquired from different sensors or generated from at least one different image of the scene, warp at least one depth map to at least one depth comparison viewpoint such that at each image comparison viewpoint there are at least two depth maps, compare the depth maps at each depth comparison viewpoint, and determine a confidence score for each depth map based on the comparison of the depth maps.

[0053] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief description of the drawings]

[0054] For a better understanding of the present invention and to show more clearly how it may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Figure 1] FIG. 1 illustrates a scene imaged by a multi-camera setup. [Diagram 2] FIG. 2 illustrates a first embodiment for determining a confidence score. [Diagram 3] FIG. 13 illustrates a second embodiment for determining a confidence score. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0055] The present invention will now be described with reference to the drawings.

[0056] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems and methods, are for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems and methods of the present invention will become better understood from the following description, the appended claims, and the accompanying drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to denote the same or similar parts.

[0057] The present invention provides a method for processing multi-view data of a scene, the method comprising obtaining at least two images of the scene from different cameras, determining a sharpness metric for each image, and determining a confidence score for each image based on the sharpness metric, the confidence score for use in determining weights when blending the images to synthesize a new virtual image.

[0058] Figure 1 shows a scene imaged by a multi-camera setup. The illustrated multi-camera setup comprises five capture cameras 104 and one virtual camera 108 for which a virtual view image is synthesized. Figure 1 is used to illustrate potential problems that arise when processing multi-view data for a fast moving object.

[0059] A fast moving circular object 102 is shown by a set of circles that are all captured by all cameras 104 within a certain integration time. The cameras 104 are assumed to be synchronized for simplicity in this example. The magnitude of motion blur caused by the movement of the object 102 is shown by solid lines on each projection plane 106a-e for each of the five cameras 104. The motion blur is relatively small on projection planes 106a, 106b and large on projection planes 106c, 106d. Projection plane 106e shows motion blur that is neither large nor small. A virtual camera 108 is shown for which a virtual image at a target viewpoint will be synthesized.

[0060] Depth estimation is likely to be successful for images corresponding to projection planes 106a, 106b, and 106e. However, for images corresponding to projection planes 106c and 106d, depth estimation is likely to fail due to motion blur. Motion blur may remove texture from the image of the foreground object 102, making it appear semi-transparent. The latter means that depth estimation is likely to make a fast-moving foreground object 102 "see-through," assigning local background depth to pixels in the blurred region. In other words, background texture may be visible through the foreground object 102 in regions with motion blur, and thus a depth map generated from an image with motion blur will show the foreground object 102 as having background depth.

[0061] When an image at the target viewpoint of the virtual camera 108 is synthesized by using the images and depth maps corresponding to the projection planes 106d and 106c (i.e., the projection plane with the largest motion blur), the new virtual image may show a blurred texture of the object 102 that appears to be in the background. Therefore, it may be necessary to use other images that are not blurred. To determine which image to use, it is proposed to determine a sharpness index for each image. The sharpness index is a measure of the blur in each image.

[0062] Conventionally, when synthesizing a new virtual image at the target viewpoint of virtual camera 108, the images corresponding to projection planes 106d and 106e have the higher weight since they are closest, and the images at 106a and 106b have the lowest weight since they are the furthest from virtual camera 108. However, a new confidence score can be given to each image based on the sharpness metric, and the synthesis of the new virtual image can be further weighted based on the confidence scores.

[0063] Thus, when synthesizing a new virtual image, pixels in the sharper images of 106a and 106b used to generate the new virtual image may be weighted higher than pixels in the blurred images of 106d and 106c, and similarly, pixels in the image of 106e may be weighted higher than the image of 106d. Some pixels in the sharper images of 106a and 106b (or any of the images) may not be used because they are not needed for the generation of the new virtual image. For example, the image of 106a contains mostly texture data of the left side of the object 102, and the virtual camera 108 is mostly "viewing" the right side of the object 102. Thus, most of the pixels in the image of 106a may not be used to generate the new virtual image.

[0064] For example, because image 106e is the closest and least blurred (compared to the other images), the overall weighting during compositing (e.g., based on confidence scores and proximity) may be high for image 106e, slightly lower for images 106a and 106b, and lowest for images 106c and 106d.

[0065] This method can be used during rendering or as a pre-processing step before rendering.

[0066] 2 shows a first embodiment for determining confidence scores during rendering. During rendering from a source view image 202 to a target viewpoint, pixels resulting from uncertain source view pixels in the source view image 202 are detected by comparing the sharpness of the incoming warped source view pixels in the target coordinate system. This case is important because it allows for on-the-fly (during rendering) source view analysis and synthesis.

[0067] It should be noted that confidence scores at different viewpoints (e.g., the coordinate system of the source view image 202) can also be determined as a separate pre-processing step in which the source view image 202 is first warped to one or more comparison viewpoints to establish confidence scores for each pixel before using the data to begin rendering.

[0068] Three camera views 208, 210 and 212 from a set of multiple camera views are shown in Figure 2. Source view image 202 of camera view 208 is a sharp image of a circular object. Source view image 202 of camera view 210 is a somewhat sharper image of a circular object showing a sharp region in the center of the image and blurry sections (e.g., due to motion blur) around the sides of the object.

[0069] The source view images 202 of the camera view 212 are blurred images of a circular image. Each source view image 202 is warped to a target viewpoint, and a sharpness index is determined for each warped image 204. The sharpness index is a measure of the sharpness of each warped image. For example, a per-pixel contrast measure can be used to determine the sharpness index. A focus measure can also be used.

[0070] In addition, if the predicted color coming from a source view image differs significantly from the average predicted across all source view images, the contribution of that source view receives lower confidence, resulting in a lower blending weight. This covers the case where a blurry foreground object is incorrectly measured as having background depth, in which case it may be warped to the wrong position (i.e., based on the wrong depth) and end up in a position where other source views predict a different color (e.g., green grass).

[0071] If a source view contribution (e.g., a pixel of the warped image 204) has a much lower image sharpness measure than the average from the other contributions, then that source view contribution also receives a lower confidence, resulting in a lower blending weight. This covers the case where a blurry foreground object has (possibly partially) the correct depth and thus maps to the correct position. Weighting a blurry pixel contribution equally to other sharper contributions may result in image blur.

[0072] The sharpness index can be based on identifying blurry regions of the warped image 204. For example, a pixel-by-pixel contrast value can be determined for each pixel in the warped image 204 to create a contrast map for each warped image 204. The contrast map can then be used as the sharpness index. Alternatively, the sharpness index can be a single value that can be compared to the sharpness indices of other warped images 204.

[0073] The sharpness metrics of the warped images 204 are compared at the comparison viewpoints, and a confidence score 206 is generated for each camera view. The confidence score 206 shown in Figure 2 is a confidence map. The confidence map 206 includes a confidence value for each pixel (or group of pixels) in the corresponding source view image 202.

[0074] The confidence map 206 for camera view 208 indicates a high confidence (shown as white) for the entire source-view image 202 for camera view 208 because the source-view image 202 is sharp and therefore has a high sharpness index. The confidence map 206 for camera view 210 indicates a high confidence for the sharp regions of the object, but a low confidence (shown as black) for the blurry regions of image 202. Similarly, the confidence map 206 for camera view 212 indicates a low confidence for the blurry regions.

[0075] Therefore, the confidence map 206 can be used during synthesis of a new virtual image as part of the weighting: the blurred regions in the camera views 210 and 212 will have a relatively lower weight when synthesizing a circular object compared to the sharp regions in the camera views 208 and 210.

[0076] 3 shows a second embodiment for determining confidence scores in pre-processing (i.e., before rendering). In the second embodiment, uncertain regions of an image are detected by comparing the depth and sharpness of corresponding pixels in multiple source view images 302. Three source view images 302 of a scene are shown along with corresponding depth maps 304 and confidence maps 306. Camera views 308 and 310 show a sharp image 302 of a fast moving circular foreground object. The corresponding depth map 304 shows an estimate of the depth of the circular object.

[0077] A camera view 312 shows the image 302 with motion blur for the object. As a result, the background texture (shown as white) is visible through the semi-transparent foreground object (shown as black). The corresponding estimated depth map 304 shows the background depth where the circular object is expected to be. A dotted line is shown in the depth map 304 corresponding to the camera view 312 to indicate where the depth of the circular object should have appeared in the depth map 304.

[0078] In order to avoid synthesis errors when synthesizing a new virtual image from the three camera views 308, 310 and 312, a confidence map 306 is generated for each camera view, the generation of which is further based on a comparison between the depth maps 304 as well as the source view images 302 in addition to comparing the sharpness metrics.

[0079] Each source view 302 is warped to every other view using an associated depth map 304. A source view pixel is further flagged as having an associated low confidence if the color and / or texture change of the warped source view is significantly different from the color and / or texture change of the target view. For example, when warping source view 308 to source view 310, the colors match as the object's colors and the local texture changes match closely. This is also true when warping source view 310 to source view 308. However, when warping pixels in source view 312 to either 308 or 310, the color and / or texture changes of the object do not match because the depth used to warp the object's pixels was incorrect. Thus, the confidence score is set to a low value for the object's pixels in source view 312.

[0080] Using the depth map 304 as an additional measure for generating the confidence map 306 can increase the accuracy of the confidence map 306. For example, in the case of a complex image with various levels of sharpness, a comparison between the depth maps 304 can further help identify the most blurry regions of the complex image.

[0081] The new image is typically generated from the source view image 302 by warping the source view image 302 to the target viewpoint and blending the warped image. Image regions around the source view 302 for the camera view 312 correspond to blurry objects and therefore have low confidence values ​​on the confidence map 306. The confidence map 306 can be warped to the target viewpoint along with the source view image 302 and used as weights in the blending operation to generate the new virtual image. As a result, fast moving objects appear sharp in the new virtual image.

[0082] In summary, it is proposed to use the sharpness (and potentially color and depth) differences between source views to solve the problems of depth estimation and new view synthesis. During depth estimation, the estimated depth and texture differences (i.e., color and sharpness) between source views can be compared. This information is then used to estimate a confidence score for the image (and corresponding depth map), setting pixels that correspond to motion blur (or any blur in general) to a low confidence. For example, when considering motion blur, the confidence score of the image can provide an indication about where fast moving objects are more likely to be located in 3D space.

[0083] One approach to determining the sharpness index is to compute a local per-pixel contrast measure for each image in the source view image coordinate system, and warp this measure along with color (and optionally depth) using a depth map to the target viewpoint (i.e., the viewpoint from which the new image is generated).

[0084] It may also be possible to simply warp the colors of each image to the target viewpoint. For example, eight source views may provide eight images that can be warped to the target virtual viewpoint, and the results are stored in the GPU memory (i.e., the memory module of the graphics processing unit). In a separate shader, local contrast may be calculated for each warped image. Based on the local contrast, a confidence (and therefore blending weight) is determined. Both approaches may also be used with a comparison viewpoint instead of the target viewpoint, which typically includes the viewpoint of each image.

[0085] In addition, during novel view synthesis, image regions with lower confidence scores receive lower blending weights. The use of confidence scores for novel view synthesis can avoid multiple copies of fast moving objects becoming visible, thus making the newly synthesized image sharper than without the use of confidence scores.

[0086] There are different approaches for when and where to determine the confidence score, and by whom. In one example, the encoder can determine the confidence score when (or after) the image is captured and send / broadcast the confidence score along with the image. The decoder can then receive the image and the confidence score and synthesize a new image using the image and the confidence score. In another example, the decoder can receive the image, determine the confidence score, and synthesize a new image. The confidence score can be determined during rendering (as shown in FIG. 2) or as a pre-processing step before rendering (as shown in FIG. 3).

[0087] As mentioned above, the method utilizes image and depth map warping. Warping can include applying a transformation to a source view image and / or source view depth map, which transformation is based on a viewpoint of the source view image and / or source view depth map and a known target viewpoint. The viewpoint is defined by at least the pose (i.e., position and orientation) of a virtual camera (or sensor) in 3D space. For example, the transformation can be based on the difference between a pose corresponding to the depth map and a known target pose corresponding to the target viewpoint. When referring to warping, it should be understood that forward warping and / or inverse (reverse) warping can be used. In forward warping, source pixels are projected onto the target image using point or triangle primitives constructed in the source view image coordinate system. In reverse warping, target pixels are inversely mapped to positions in the source view image and sampled accordingly.

[0088] Possible warping approaches include using points, using a regular mesh (ie, of a predefined size and topology), and / or using an irregular mesh.

[0089] For example, using the points may include using a depth map (for a given pixel) from a first perspective (view A) to calculate the corresponding location in a second perspective (view B) and fetching the pixel location from view B back to view A (i.e., inverse warp).

[0090] Alternatively, for example, using the points may include using a depth map (for a given pixel) of view A to calculate the corresponding pixel location in view B and mapping the pixel location from view A to view B (i.e., forward warping).

[0091] Using a regular mesh (e.g., two triangles per pixel, two triangles per 2x2 pixels, two triangles per 4x4 pixels, etc.) may involve calculating 3D mesh coordinates from a depth map in view A and texture mapping the data from view A to view B.

[0092] Using an irregular mesh may include generating a mesh topology for view A based on a depth map (and optionally texture and / or transparency data in view A) and texture mapping data from view A to view B.

[0093] An image can be warped by using a corresponding depth map. For example, for an image view A and a depth map in view B, the depth map can be warped to view A, and the image can be warped to a different view C based on warping of the depth pixels corresponding to the image pixels.

[0094] Those skilled in the art can easily develop a processor to perform any of the methods described herein. Thus, each step of the flowchart may represent a respective operation performed by a processor and may be executed by a respective module of the processor.

[0095] As described above, the system utilizes a processor to process data. The processor may be implemented in a variety of ways using software and / or hardware to perform the various functions required. The processor typically uses one or more microprocessors that are programmed using software (e.g., microcode) to perform the functions required. The processor may also be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0096] Examples of circuitry that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).

[0097] In various implementations, the processor may be associated with one or more storage media, e.g., volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on the one or more processors and / or controllers, perform the required functions. The various storage media may be attached within the processor or controller, or may be transportable such that one or more programs stored on the storage media are read into the processor.

[0098] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0099] A single processor or other unit may fulfill the functions of several items recited in the claims.

[0100] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0101] The computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0102] When the term "adapted for" appears in the claims or the specification, it is meant to be synonymous with the term "configured to."

[0103] Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. 1. A method for processing multi-view data of a scene, comprising: acquiring at least two images of the scene from different cameras; determining a sharpness index for each image, the sharpness index being a sharpness map having a plurality of sharpness values ​​each corresponding to one or more pixels of the corresponding image; determining a confidence score for each image based on the sharpness metric, the confidence score for use in determining weights when blending images to synthesize a new virtual image via viewpoint interpolation; A method having the following.

2. obtaining at least one depth map of the scene; warping at least two of the images to a target viewpoint based on the at least one depth map; blending the images at the target viewpoint to generate a composite image; and The method of claim 1 , wherein during blending, each pixel in the image is weighted based on a corresponding confidence score.

3. obtaining at least one depth map of the scene; warping at least one image to at least one image comparison viewpoint using the at least one depth map such that there are at least two images at each image comparison viewpoint; comparing pixel color values ​​of the images at each comparison viewpoint; and The method of claim 1 , wherein determining a confidence score for each image is further based on the comparison of the pixel color values.

4. The method of claim 3 , wherein the image comparison viewpoints include all viewpoints of the images.

5. 4. The method of claim 3, wherein the image comparison viewpoint is a target viewpoint, the method further comprising blending the images at the target viewpoint to generate a composite image, wherein during blending, each pixel in the images is weighted based on a corresponding confidence score.

6. acquiring at least two depth maps, said depth maps being acquired from different sensors or generated from at least one different image of the scene; warping at least one depth map to at least one depth comparison viewpoint such that at each image comparison viewpoint there are at least two depth maps; comparing the depth maps at each depth comparison viewpoint; determining a confidence score for each depth map based on the comparison of the depth maps; 6. The method of claim 1, further comprising:

7. obtaining at least two depth confidence maps corresponding to said depth map; warping at least one depth confidence map together with a corresponding depth map to said at least one depth comparison viewpoint; and The method of claim 5 , wherein comparing the depth maps at each depth comparison viewpoint further comprises comparing corresponding depth confidence maps.

8. A computer program product which, when run on a computing device, causes the computing device to carry out the method of any one of claims 1 to 5 and 7.

9. 1. A system for processing multi-view data of a scene, comprising: acquiring at least two images of the scene from different cameras; determining a sharpness index for each image, the sharpness index being a sharpness map having a plurality of sharpness values ​​each corresponding to one or more pixels of the corresponding image; determining a confidence score for each image based on the sharpness metric, the confidence score for use in determining weights when blending images to synthesize a new virtual image via viewpoint interpolation; 1. A system having a processor configured to execute:

10. The processor further comprises: obtaining at least one depth map of the scene; warping at least two of the images to a target viewpoint based on the at least one depth map; blending the images at the target viewpoint to generate a composite image; further configured to perform The system of claim 9 , wherein during blending, each pixel in the image is weighted based on a corresponding confidence score.

11. the processor: obtaining at least one depth map of the scene; warping at least one image to at least one image comparison viewpoint using the at least one depth map such that there are at least two images at each image comparison viewpoint; comparing pixel color values ​​of the images at each comparison viewpoint; further configured to perform The system of claim 9 , wherein determining a confidence score for each image is further based on the comparison of the pixel color values.

12. The system of claim 11 , wherein the image comparison viewpoints include all viewpoints of the images.

13. 12. The system of claim 11, wherein the image comparison viewpoint is a target viewpoint, and the processor is further configured to perform the step of blending the images at the target viewpoint to generate a composite image, wherein during blending, each pixel in the images is weighted based on a corresponding confidence score.

14. the processor: acquiring at least two depth maps, said depth maps being acquired from different sensors or generated from at least one different image of the scene; warping at least one depth map to at least one depth comparison viewpoint such that at each image comparison viewpoint there are at least two depth maps; comparing the depth maps at each depth comparison viewpoint; determining a confidence score for each depth map based on the comparison of the depth maps; 14. The system of claim 9, further configured to: