Shape-dependent new view synthesis
By estimating the scene surface shape and gradient, and employing a shape-dependent view assignment method, the artifact and sharpness issues in new view compositing are resolved, achieving high-quality view compositing effects.
Patent Information
- Application Number
- CN202480031717.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-11
- Filing Date
- 2024-05-07
- Publication Date
- 2025-12-09
AI Technical Summary
Existing new view synthesis methods are sensitive to calibration and depth errors, resulting in artifacts and reduced sharpness, especially when using a small number of cameras, and adding more cameras is impractical.
By estimating the surface shape within the scene, the view assignment is determined based on the surface gradient. The target view is rendered using a shape-dependent view assignment method, reducing the dependence on camera calibration and generating a new view with accurate textures using a small number of cameras.
It reduces artifacts caused by calibration and depth errors, improves the sharpness of new views, and enables high-quality view synthesis under conditions of limited cameras.
Smart Images

Figure CN121100523A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to rendering images in multi-view video at a target viewpoint. In other words, this invention relates to novel view compositing using a multi-camera capture system. Background Technology
[0002] Synthesizing new views from multiple source view cameras is an active area of research and development. New view synthesizing is particularly challenging when the number of cameras is limited or the goal is to create a large viewing space where virtual cameras can be placed in various locations within the scene. One solution to this problem is to add depth sensors and form a hybrid camera setup that includes both color and depth sensors. The use of depth sensors generally allows for a larger viewing space.
[0003] Existing methods for novel view synthesis are highly sensitive to calibration and depth errors. These errors typically lead to visible artifacts. Furthermore, view fusion methods generally reduce sharpness, while neural network-based methods require good calibration to learn a sound predictive model.
[0004] Artifacts have been found to be a serious problem when using a small number of cameras with a large viewing area. However, adding more cameras is generally not an option. Furthermore, calibration errors are difficult to resolve without using specific calibration modes, and depth measurements obtained via time-of-flight depth sensors may be inaccurate and noisy.
[0005] All these issues typically result in noticeable artifacts during the compositing / rendering of new views, negatively impacting the viewing experience. Therefore, an improved method is needed to compose new views with reduced visual artifacts. Summary of the Invention
[0006] This invention is defined by the claims.
[0007] According to one aspect of the present invention, a method is provided for rendering a target view image of a scene at a target viewpoint, the method comprising:
[0008] Obtain at least two source view images of the scene captured from different locations and at least one depth map of the scene;
[0009] The source view image and / or the depth map are used to estimate the surface shape within the scene, wherein the surface shape indicates the gradient of a surface in the scene relative to the location where the source view image was captured;
[0010] Determined based on the surface shape and the positions of the captured first and second source view images:
[0011] The first part of the scene to be rendered using the first source view image; and
[0012] The second part of the scene to be rendered using the second source view image; and
[0013] A target view image of the scene at the target viewpoint is rendered using the at least one depth map, a first source view image for the first part of the scene, and a second source view image for the second part of the scene.
[0014] This method provides a shape-dependent view assignment for the texture, rather than a mixture of textures from various views. This reduces artifacts caused by calibration and depth errors and increases the sharpness of the new rendered view. This is especially important when capturing the source view image using a small number of cameras.
[0015] The result is a new view rendered with accurate texture values at the target viewpoint, which does not require highly accurate camera calibration. Therefore, a small number of cameras can be used to generate new views with accurate texture values. This is because the gradient of the scene relative to the camera's surface is taken into account when assigning the source view to each part.
[0016] Using surface shapes to determine the first and second portions provides a shape-dependent view assignment. In other words, determining the first and second portions can be viewed as assigning weights to the source view (compared to conventional view blending) that depend on indications of the gradient of surfaces in the scene (i.e., surface shapes). Preferably, at least a portion of the first portion should have 100% weight with respect to the first image, and at least a portion of the second portion should have 100% weight with respect to the second image.
[0017] Estimating surface shape can include assuming surface shape based on the type of objects in the scene (for example, one could assume that a person's head is convex).
[0018] Determining the portion means assigning a source view camera to each portion based on the surface shape relative to the position of each source view camera. The corresponding portion is then rendered using source view images from the assigned source view cameras.
[0019] The method may also include using surface shape to determine which source view camera is most orthogonal to one or more surfaces in the portion. In this case, determining which source view image to use for the portion is based on the source view camera that is most orthogonal to one or more surfaces in the portion.
[0020] Images can be obtained from a multi-camera capture system that includes two or more source view cameras that image the scene.
[0021] In the example, two source view images and two source view depth maps are used. Therefore, the target view can be rendered by warping the first source view image using the first depth map and the second view image using the second depth map. A view assignment map is then used to select which warped view to choose in the target image coordinates.
[0022] In another example, a source view image can be obtained by moving a single camera from one location to another, assuming the scene is static (i.e., there is no motion).
[0023] Estimating the surface shape includes identifying one or more types of objects and / or object features in the scene, and assigning a predetermined surface shape to the one or more objects and / or features based on the identified object types and / or object features;
[0024] It is known that objects of a certain type have a general shape. For example, the surface of a human head can generally be defined as convex. Therefore, by identifying the type of object in the scene, a predetermined surface shape can be assigned to that object.
[0025] Similarly, it is known that various types of object features (e.g., nose, eyebrows, wrinkles, etc.) also have general shapes.
[0026] This simplifies the process of estimating surface shape compared to analyzing depth maps and directly calculating surface gradients.
[0027] The predetermined surface shape can be convex, concave, or flat.
[0028] Generally defining the surface shape as convex, concave, or flat provides a simple format for determining which source view should be used to render which part of the surface.
[0029] For example, for a concave shape, the left side of the surface faces the right side of the multi-camera capture system, and vice versa. Therefore, the rightmost camera can best convey the texture of the left side of the surface, and the leftmost camera can best convey the texture of the right side of the surface.
[0030] Identifying one or more types of objects in a scene may include applying object detection algorithms to images and / or depth maps.
[0031] Identifying one or more types of objects in a scene can include applying a keypoint detection algorithm to an image.
[0032] Keypoint detection is typically performed for a given object category (e.g., humans). Furthermore, keypoint detection can serve as the basis for assumptions about richer surface shape information. For example, using human keypoint detection algorithms, specific body parts (e.g., the left upper arm) can be identified.
[0033] Estimating surface shape can include calculating the surface normal vector in the depth map and determining the angle between the normal vector and the vector between the location of the captured source view image and the surface.
[0034] Determining the first and second portions may include comparing angles determined for different source view images and assigning a source view image to each portion based on the comparison of the angles.
[0035] In some cases, it may be preferable to have a per-pixel indication of the gradient to more accurately represent the shape, rather than assuming, for example, that the surface is convex / concave / flat. Calculating the normal vector can provide a more accurate surface shape, thus providing a more accurate separation between the first and second parts.
[0036] Preferably, determining the first part and the second part involves assigning a camera to each pixel based on the camera whose angle between the normal vector at the pixel location and the vector between the camera and the pixel is the smallest.
[0037] In some examples, determining the first and second parts involves using the weighted sum of image coordinates for each line to determine the seam line separating the first and second parts.
[0038] This method simplifies the determination of the first and second parts because it eliminates the need to determine the surface normals. This method may be particularly advantageous in video call applications, where it is desirable to synthesize new views in real time.
[0039] The sum of the weighted image coordinates for each line can be weighted according to the normalized depth of the target viewpoint.
[0040] The method may also include applying a low-pass filter to the depth map and using the filtered depth map to estimate the surface shape.
[0041] Applying a low-pass filter can reduce noise in the depth map, thereby reducing unnecessary view allocation caused by noise.
[0042] In some cases, it may be preferable to apply separate low-pass filters to pixels belonging to the same depth surface to avoid averaging over sudden depth changes (e.g., between different surfaces at different depths). This would be a non-linear filtering operation.
[0043] Applying a low-pass filter can include applying the low-pass filter to two or more surfaces in a depth map to generate a filtered depth map.
[0044] The method may also include identifying surfaces in the depth map.
[0045] The method may also include identifying one or more objects and / or object features in the scene, and adjusting the first and second parts if objects and / or object features are identified between the first and second parts.
[0046] The area or line between the first and second parts can be referred to as the stitching area / line. If this area / line falls on a specific object or object feature, it can cause unwanted artifacts due to subtle differences in texture variations captured by the camera (e.g., light reflections may differ between cameras due to different camera positions). Therefore, adjusting the sections to move the stitching line / area can reduce artifacts in the target view.
[0047] Object features are the characteristics of an object (e.g., nose, eyebrows, etc.). These features may affect the overall shape of an object (e.g., head). It is understood, of course, that the distinction between an object and object features can be arbitrarily chosen. For example, if the expected scene is of people, the head, arms, body, etc., might be considered objects, while the nose, eyebrows, etc., might be considered object features. However, when the expected scene contains a group of people, the head, arms, body, etc., as well as the nose and eyebrows, might be considered object features. The exact distinction between an object and object features may be based on the scene being imaged and / or personal preference.
[0048] The method may further include calculating the color difference between corresponding colors in the first source view and the second source view in the stitching region, wherein the stitching region is a portion of the target view image that overlaps with both the first and second portions, and adjusting the colors in the stitching region based on the color difference to gradually change the colors from the first portion and the second portion.
[0049] The splicing area may overlap with the first part by less than 20%, 15%, 10%, or 5%, and with the second part by less than 20%, 15%, 10%, or 5%. The splicing area does not overlap with the entire first part or the entire second part.
[0050] The present invention also provides a computer program carrier including computer program code, which, when executed on a computer, causes the computer to perform all the steps described in any of the above methods.
[0051] A computer program carrier may include computer memory (e.g., random access memory) and computer storage devices (e.g., hard disk drives, solid-state drives, etc.). A computer program carrier may be a bitstream carrying computer program code.
[0052] The present invention also provides a system for rendering a target view image of a scene at a target viewpoint, the system comprising a processor configured to:
[0053] Obtain at least two source view images of the scene captured from different locations and at least one depth map of the scene;
[0054] The source view image and / or the depth map are used to estimate the surface shape within the scene, wherein the surface shape indicates the gradient of a surface in the scene relative to the location where the source view image was captured;
[0055] Determined based on the surface shape and the positions of the captured first and second source view images:
[0056] The first part of the scene to be rendered using the first source view image; and
[0057] The second part of the scene to be rendered using the second source view image; and
[0058] Using the at least one depth map and the first source view image for the first part of the scene and the second source view image for the second part of the scene, a target view image of the scene is rendered at the target viewpoint.
[0059] The processor is configured to estimate the surface shape by identifying one or more types of objects and / or object features in the scene, and assigning a predetermined surface shape to the one or more objects and / or features based on the identified types and / or object features;
[0060] The predetermined surface shape can be convex, concave, or flat.
[0061] The processor can be configured to identify one or more types of objects in the scene by applying an object detection algorithm to an image and / or depth map.
[0062] The processor can be configured to identify one or more types of objects in the scene by applying a keypoint detection algorithm to the image.
[0063] The processor can be configured to estimate the surface shape by calculating the normal vector of the surface in the depth map and determining the angle between the normal vector and the vector between the location of the captured source view image and the surface.
[0064] The processor can be configured to determine a first part and a second part by comparing angles determined for different source view images and assigning source view images to each part based on the comparison of the angles.
[0065] The processor can also be configured to apply a low-pass filter to the depth map and use the filtered depth map to estimate the surface shape.
[0066] The processor can be configured to apply a low-pass filter to two or more surfaces in the depth map to generate a filtered depth map.
[0067] The processor can also be configured to identify one or more objects and / or object features in the scene, and to adjust the first and second parts if objects and / or object features are identified between the first and second parts.
[0068] The processor can also be configured to calculate the color difference between corresponding colors in the first source view and the second source view in the stitching region, wherein the stitching region is a portion of the target view image that overlaps with both the first and second portions, and adjusts the colors in the stitching region based on the color difference to gradually change the colors from the first portion and the second portion.
[0069] These and other aspects of the invention will become apparent and will be explained with reference to the embodiments described below. Attached Figure Description
[0070] To better understand the invention and to more clearly illustrate how it can be practiced, reference will now be made to the accompanying drawings by way of example only, wherein,
[0071] Figure 1 The illustration shows two source-view cameras imaging a tilted surface;
[0072] Figure 2 The illustration shows two cameras imaging a second tilted surface;
[0073] Figure 3 The illustration shows the use of two cameras to assign views to a convex surface;
[0074] Figure 4 The illustration shows the use of two cameras to assign views to a concave surface;
[0075] Figure 5 The diagram illustrates the view allocation in a video call scenario that includes a single person;
[0076] Figure 6 The illustration shows how to adjust the stitching lines by recognizing the nose of a person in a new view;
[0077] Figure 7 The diagram illustrates the view assignment for three cameras placed in a triangular configuration;
[0078] Figure 8 The image shows a target view with two parts and a stitched area; and
[0079] Figure 9A method for rendering a target view image of a target viewpoint scene using a multi-camera capture system is shown. Detailed Implementation
[0080] The invention will be described with reference to the accompanying drawings.
[0081] It should be understood that while the detailed description and specific examples indicate exemplary embodiments of the devices, systems, and methods, they are intended for illustrative purposes only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will be better understood from the following description, the appended claims, and the accompanying drawings. It should be understood that the drawings are schematic only and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to denote the same or similar parts.
[0082] This invention provides a method for rendering a target view image of a scene at a target viewpoint. The method includes acquiring at least two source view images of the scene captured from different locations and at least one depth map of the scene. The source view images and / or the depth map are used to estimate surface shapes within the scene, wherein the surface shapes indicate the gradients of surfaces in the scene relative to the locations where the source view images were captured. Based on the surface shapes and locations of the captured first and second source view images, a first portion of the scene to be rendered using the first source view image is determined, and a second portion of the scene to be rendered using the second source view image is also determined. The target view image of the scene at the target viewpoint is rendered using the at least one depth map and the first source view image for the first portion of the scene and the second source view image for the second portion of the scene.
[0083] It has been noted that when synthesizing from a single source view camera in a depth map, the quality of the new view remains higher compared to view blending methods. However, a second or third camera with a depth sensor is often needed to address occlusion issues. Since blending is not a good option, the question becomes which source view should be selected for each pixel or region in the target view. Surface shape and orientation have been found to be good properties for selection. Generally, for better image quality and for a given region in the target view, the source view camera that views that region as orthogonally as possible relative to the 3D surface should be selected.
[0084] Therefore, it is recommended that the surface shape determine which source view contribution is selected for each pixel in the target view. This shape-dependent view assignment borrows the concept of traditional image stitching, but differs in that the image is warped to the target view and assigned using surface shape criteria. The effect of this method is to optimally select pixels from the source views to maintain resolution, and the associated stitching artifacts are less noticeable than blending artifacts.
[0085] Figure 1 The illustration shows two source view cameras 102 and 104 imaging a tilted surface 101. In this example, the entire tilted surface 101 is more oriented toward the source view camera 104 than the source view camera 102. Therefore, the image captured from the source view camera 104 will be used to render a composite view (e.g., at the target viewpoint between cameras 102 and 104).
[0086] The orientation / gradient of the tilted surface 101 can be determined by comparing the angle between the normal vector of the tilted surface 101 and the vector from the cameras 102 and 104 to the tilted surface 101.
[0087] For example, at point 106 on the inclined surface 101, the angle between the normal vector 108 from point 106 and the vector 110 between camera 102 and point 106. The angle between the normal vector 108 from point 106 and the vector 112 between camera 104 and point 106 is greater than the angle between them. Therefore, point 106 on the tilted surface is more oriented toward camera 104 than toward camera 102. The vector between the camera and a point on the surface can be called the source view direction for a given pixel / region.
[0088] Similarly, at point 114 on the inclined surface 101, the angle between the normal vector 116 from point 114 and the vector 120 between camera 102 and point 114 is... The angle between the normal vector 116 from point 114 and the vector 118 between camera 104 and point 114 is greater than the angle between them. Therefore, point 114 on the inclined surface is more oriented toward camera 104 than toward camera 102. Therefore, it can be assumed that all points on surface 101 can be assigned to the image from camera 102.
[0089] It should be understood that for a surface that is assumed to be planar, a single normal vector can be used because the normal vector can be translated across the entire surface.
[0090] Generally speaking, angle and It provides an indication of the gradient of the surface relative to the camera. Here, any index indicating the gradient of the surface relative to the camera is referred to as the surface shape.
[0091] Figure 2 The illustration shows two cameras, 202 and 204, imaging the second tilted surface 201. Figure 2 The diagram illustrates the relationship between... Figure 1 A similar situation. However, in this case, the view assignment changes on surface 201. The angle between the source view direction from camera 202 and the normal vectors from surface 201 (e.g., normal vectors 206, 208, and 210) can be used. And further use the angle between the source view direction from camera 204 and the normal vector from surface 201. To determine the gradient of surface 201 relative to cameras 202 and 204.
[0092] For the first portion 212 of surface 201, where, The pixels are selected from camera 202. However, for the second part 216, where, The pixels are selected from camera 204. There may also be a third part 214, called the stitching part, where the difference between angles is less than a threshold. In the stitching part 214, pixels can be taken from either camera 202 or camera 204, or pixel color, transparency, and depth values from cameras 202 and 204 can be combined to create new color, transparency, or depth values. This can be achieved through filtering operations.
[0093] Figure 1 and Figure 2 The image shows a planar surface. However, many surfaces in a scene will not be planar, or cannot be assumed to be planar. In these cases, a per-pixel surface shape analysis can be performed using the depth map to determine the normal vector for each pixel. However, this analysis can be susceptible to noise in the depth map.
[0094] Therefore, for certain object categories (such as human targets), it can be assumed that the object surface shape is convex or concave to avoid potentially noisy surface shape analysis. For example, for human targets, it can be assumed that most body parts give a convex depth profile. Examples of such objects include the head, neck, shoulders, and arms. However, a counterexample is an open hand facing the depth camera, which would result in a concave depth surface.
[0095] Figure 3The diagram illustrates the view assignment for a convex surface 301 using two cameras 302 and 304. In this case, a normal vector 308 from point 306 can be found such that the angle between the normal vector 308 and the source view directions from cameras 302 and 304 is equal. Thus, the left portion 310 of surface 301 (to the left of point 306) can be assigned to camera 302, and the right portion 312 (to the right of point 306) can be assigned to camera 304.
[0096] In some applications, it can be assumed that point 306 will fall at the geometric center of surface 301, and therefore there is no need to compute normal vector 308. For example, in a video call, it can be assumed that the head of a human subject facing the camera is convex and facing the camera such that the nose points exactly between camera points 302 and 304. Therefore, in this case, it can be assumed that the center of the face is the stitching line (i.e., the line between the first and second parts). Of course, assuming that the noise in the depth map is small, determining normal vectors with equal angles can provide more robust image new view synthesis.
[0097] Figure 4 The view assignment for concave 401 using two cameras 402 and 404 is shown. A normal vector 408 from point 406 can be found such that the angle between normal vector 408 and the source view direction vectors from cameras 402 and 404 is equal. Thus, the right portion 412 of surface 401 (to the right of point 406) can be assigned to camera 402, and the left portion 410 of surface 401 (to the left of point 406) can be assigned to camera 404.
[0098] For example, concave and convex surfaces can be hypothesized by using object detection and classification algorithms on images and / or depth maps and identifying one or more types of objects. Specific types of objects (e.g., heads, shoulders, arms, etc.) can be considered convex or concave, or even planar. Therefore, convex surfaces, concave surfaces, or planes can be matched with surfaces in the depth map.
[0099] Another approach might be to use keypoint detection algorithms on the image to identify the type of object. Therefore, the shape of the object of that type can be assumed. Keypoint detection can be used to identify key points of objects in an image. For example, a human body keypoint detection algorithm can be used to detect the human body in an image. Then, key points (e.g., left knee, right elbow, left eye, etc.) can be used to fit convex or concave surfaces to specific features of the human body (e.g., a convex surface fitting the head).
[0100] Figure 5The diagram illustrates view assignment in a video call scenario involving a single person 506. In this case, two cameras are used to image the person 506, and two source view images 502 and 504 are obtained from the cameras. A new view 508 at the target viewpoint is then synthesized using the two source view images 502 and 504. In this case, the head, neck, and body of the person 506 can all be considered convex.
[0101] To construct the view assignment, pixels in source views 502 and 504 are first warped to the target view using the source view depth map. (The image coordinates in the target view are then used.) In this process, the weighted sum of image coordinates for each line can be calculated to determine the splicing line 510. The splicing line is the line that separates the first region 512 and the second region 514. The texture value of the source view 502 is used for the first region 512, and the texture value of the source view 504 is used for the second region 514.
[0102] The position of splicing line 510 can be determined by calculation: in, It is the position of the splicing line in the coordinates of the target view image, and It is a weighting factor that depends on the normalized depth observed from the target viewpoint. :
[0103] Normalized depth can be obtained by warping one or both of the source view depth maps to the target viewpoint. For example, a depth map can be obtained by determining the disparity map between images 502 and 504 and transforming it into a depth map. A depth sensor can also be used to obtain a depth map.
[0104] Calculating the position of the splice line 510 directly means that it is not necessary to calculate the normal vector for the surfaces involved. Instead, it can be assumed that all regions 512 to the left of the splice line 510 can be assigned to the source view 502, and all regions 514 to the right of the splice line 510 can be assigned to the source view 504, since all surfaces are assumed to be convex.
[0105] It should be understood that if all surfaces can be assumed to be concave, then region 512 will use the texture from source view 504, and region 514 will use the texture from source view 502.
[0106] In some cases, the stitching line 510 is preferably repositioned based on object features in the image. For example, studies have found that depth variations near the nose region can cause artifacts when the stitching line does not cross the center of the nose.
[0107] Figure 6 The splicing line 510 is shown by identifying the nose 602 of a person 506 in a new view 508. Figure 6 Figure a shows the splice line 510 positioned next to the nose 602. It has been found that this positioning of the splice line 510 causes artifacts to be more pronounced near the splice line 510 on the nose 602 (e.g., due to obstruction from one of the source views, or due to the fact that, locally, the surface on either side of the nose is actually concave rather than convex). In this case, the splice line 510 can be adjusted to pass through the center of the nose 602, as shown... Figure 6 As indicated in b. This can be achieved, for example, by detecting the position of the nose 602 in the new view 508 (e.g., via object detection, keypoint detection, etc.) and adjusting the position of the stitching line to pass through the center of the nose 602. Note that it may be more efficient and easier to detect the noise location in one or two source views and warp that location to the target view (using one or more depth maps), and then modify the stitching line in the target view.
[0108] Therefore, knowing that the seam line 510 may cross the center of the face area of the person 506, a positional offset of the seam line 510 can be introduced so that the seam line 510 falls in a more advantageous position than, for example, crossing the cheek of the person 506.
[0109] When the splice line 510 crosses the center of the object (e.g., due to significant depth variations at the center), other objects may cause more noticeable artifacts. Therefore, in other cases, it is preferable to move the splice line 510 away from the center of the object.
[0110] More generally, the movement of the splicing line 510 is an adjustment of portions 512 and 514. In this example, region 512 is expanded near the face to include half of the nose 602, and region 514 is reduced to include only half of the nose 602.
[0111] It should be understood that objects can be identified in the source view, the warped source view (i.e., at the target viewpoint), one or more depth maps, and / or one or more warped depth maps.
[0112] The concepts discussed in this paper can also be extended to a larger number of cameras. Surface shape analysis indicating surface gradients can be generalized to multiple source view cameras.
[0113] Figure 7 The illustration shows the view distribution of three cameras 702, 704, and 706 placed in a triangular configuration. Figure 7a shows the view assignment of the convex surface 701a imaged by cameras 702, 704, and 706. In this case, region 708a will be assigned to camera 702, region 710a will be assigned to camera 704, and region 712a will be assigned to camera 706.
[0114] Figure 7 b illustrates the view assignment of concave 701b captured by cameras 702, 704, and 706. In this case, region 708b will be assigned to camera 702, region 710b will be assigned to camera 704, and region 712b will be assigned to camera 706.
[0115] The concepts described herein enable the creation of an algorithm suitable for real-time view assignment as part of a renderer running on a graphics computing unit (GPU) of a consumer device. For example, the algorithm can warp the depth map of each source view to the target viewpoint, combine the depth maps in the target view, select the closest depth value for each pixel in the target view, filter the combined depth map using a low-pass filter, compute the normal vector for each pixel in the target view, calculate the cosine of the angle between the normal vector and the source view direction, and assign a source view corresponding to the minimum angle for each pixel in the target view.
[0116] The filtering step is included to prevent view assignment (and therefore the location of parts) from becoming sensitive to depth noise. For example, a Gaussian low-pass filter with a large sigma (e.g., >30 pixels) can be used. The normal vector per pixel can be computed by constructing two 3D world vectors with the 3D position of the center pixel. We take the starting point as the vector and the 3D position of each distinct adjacent pixel as the ending point. To avoid parallel vectors, we can select them from a regular image grid at a 90-degree angle. Let these two vectors be represented as... and If we assume that the depth surface can be locally approximated by a plane, then the normal vector n can be approximated by the cross product: For each source view i The angle between the surface normal vector n and the source view direction The cosine can be calculated as: in, It is for the source view. i The normalized source view direction vector is calculated as follows: in, This refers to the source view camera position in world space. This can be determined by calibrating external parameters (e.g., using motion structure algorithms), or it may already be known (e.g., for a camera integrated into a single device). For a given pixel, the view assignment can now be defined as:
[0117] In all the examples provided above, each portion of the target view uses only one source view. However, it will be understood that, depending on the camera setup, source views can be blended by portion. For example, a camera setup might include two sets of cameras, the first set for imaging the right side of the object and the second set for imaging the left side. If a set of cameras is close to each other, it can be assumed that the texture values between the cameras will be similar (e.g., the reflections between a set of cameras will not differ significantly). In this case, the blending of the cameras in the first set can be used to render the right side of the target view, and similarly, the blending of the cameras in the second set can be used to render the left side of the target view.
[0118] Figure 8 A target view image 802 is shown, comprising two portions 804 and 806 and a stitching region 810. The left portion 804 of the target view image 802 is rendered using a first source view image warped to the target viewpoint, and the right portion 806 is rendered using a second source view image warped to the target viewpoint. The line separating the left portion 804 and the right portion 806 is a stitching line 808. The stitching region 810 is defined around the stitching line 810 and overlaps with a portion of the first portion 804 and a portion of the second portion 806.
[0119] When the source view images used for portions 804 and 806 are captured by different cameras and / or from different positions, the different viewing angles of the source view cameras relative to the main light source cause a visible color transition at the stitching line 808. Therefore, it is proposed to modify the color intensity of the warped source view at the stitching region 810 based on the color observed relative to the stitching line 808 position (or more generally, the view assignment map).
[0120] One solution is to blend the contributions from the first and second warped source views in the stitching region 810. However, this could introduce blurring into the stitching region 810. Therefore, it is also proposed to calculate the average color difference of the sub-regions 812 of the stitching region 810 across multiple adjacent pixels and gradually adjust the colors of one or more source views as they move closer to the stitching line 808. The effect is that the original source view colors in the stitching region 810 gradually change in the left portion 804 and the right portion 806, as it takes into account the observed differences between the colors in their corresponding source view images.
[0121] For example, in sub-region 812, corresponding pixels of the warped first source view can be compared with corresponding pixels of the warped second source view. The average color difference between these pixels can be calculated and then used to adjust the color of the center pixel of sub-region 812. This operation can be performed on all pixels in the stitching region 810 so that the colors in the stitching region 810 match at the stitching line.
[0122] It should be understood that colors in the stitching area 810, which are closer to the left region 804 or the right region 806, will have the least adjustment to match adjacent pixels that are not in the stitching area 810. Similarly, colors in the center of the stitching area 810 are likely to have the most adjustment.
[0123] Color adjustments can be made separately between the source views. In other words, for the first source view image, half of the color difference measured at the splicing line 808 can be added (or subtracted), and for the second source view image, half of the color difference measured at the splicing line 808 can be subtracted (or added). The effect of these adjustments is that the color difference between the source view images becomes less noticeable at the splicing line 808.
[0124] For each source view image, the adjustment can be linearly increased from zero compensation at the boundary of the stitching region 810 to the stitching line 808. This can be achieved by calculating the shortest distance from each pixel to the stitching line 808 and dividing it by half the width of the stitching region 810.
[0125] For each color channel (red, green, blue), color adaptation can use the same correction factor, or it can be applied to a transformed color space (such as YUV or others).
[0126] Figure 9A method for rendering a target view image of a scene at a target viewpoint using a multi-camera capture system is illustrated. In step 902, at least two source view images of the scene are obtained from different cameras of the multi-camera capture system, and at least one depth map of the scene is also obtained. In step 904, the source view images and / or the depth map are used to estimate surface shapes within the scene, wherein the surface shapes indicate the gradient of a surface in the scene relative to a camera. The surface shapes are estimated by identifying one or more types of objects and / or object features in the scene, and assigning predetermined surface shapes to the one or more objects and / or object features based on the identified types and / or object features. In step 906, a first portion of the scene to be rendered using the first source view image and a second portion of the scene to be rendered using the second source view image are determined based on the surface shapes corresponding to the first and second source view images and the camera positions. Therefore, in step 908, the target view image of the scene at the target viewpoint can be rendered using the depth map and the first portion of the scene for the first source view image and the second portion of the scene for the second source view image.
[0127] Technicians will be able to easily develop computers for performing any of the methods described herein. Therefore, each step of the flowchart can represent a different action performed by a processor and can be executed by the corresponding module of the processor.
[0128] One or more steps of any method described herein can be performed by one or more processors. A processor includes electronic circuitry suitable for processing data. Any method described herein can be computer-implemented, wherein computer-implemented means that the steps of the method are performed by one or more computers, and wherein a computer is defined as a device suitable for processing data. A computer may be adapted to process data according to prescribed instructions.
[0129] As described above, the system utilizes a processor to perform data processing. A processor can be implemented in various ways, using software and / or hardware, to perform a variety of required functions. A processor typically employs one or more microprocessors, which can be programmed using software (e.g., microcode) to perform the desired functions. A processor can be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuitry for performing other functions.
[0130] Examples of circuits that may be used in various embodiments of this disclosure include, but are not limited to, conventional microprocessors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).
[0131] In various implementations, the processor may be associated with one or more storage media, such as volatile and non-volatile computer memories, such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when run on one or more processors and / or controllers, perform the required functions. The various storage media may be fixed within the processor or controller, or they may be portable, allowing one or more programs stored thereon to be loaded into the processor.
[0132] Those skilled in the art, through studying the accompanying drawings, the disclosure, and the claims, will be able to understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude a plurality.
[0133] A single processor or other unit can perform the functions of several items described in the claims.
[0134] Although specific measures are described in different dependent claims, this does not imply that combinations of these measures cannot be used advantageously.
[0135] Computer programs can be stored / distributed on suitable media such as optical storage media or solid-state media that are provided together with or as part of other hardware, but they can also be distributed in other forms such as via the Internet or other wired or wireless telecommunications systems.
[0136] If the term “suitable” is used in the claims or description, it should be noted that the term “suitable” is intended to be equivalent to the term “configured as”.
[0137] Any reference numerals in the claims should not be construed as limiting the scope.
[0138] Any method described in this document excludes methods for performing such mental acts.
Claims
1. A method for rendering a target view image of a scene at a target viewpoint, the method comprising: Obtain (902) at least two source view images of the scene captured from different locations and at least one depth map of the scene; Using the source view image and / or the depth map to estimate (904) the surface shape within the scene, wherein the surface shape indicates the gradient of a surface in the scene relative to the location where the source view image was captured, wherein estimating the surface shape includes: identifying one or more types of objects and / or object features in the scene, and assigning a predetermined surface shape to one or more objects and / or object features based on the identified types and / or object features; Determined based on the surface shape and the positions of the captured first and second source view images (906): The first part of the scene to be rendered using the first source view image; and The second part of the scene to be rendered using the second source view image; and The target view image of the scene at the target viewpoint is rendered (908) using the at least one depth map, the first source view image for the first part of the scene, and the second source view image for the second part of the scene.
2. The method according to claim 1, wherein, The predetermined surface shape is convex, concave, or planar.
3. The method according to claim 1 or 2, wherein, Identifying one or more types of objects in the scene includes applying an object detection algorithm to the image and / or the depth map.
4. The method according to any one of claims 1 to 3, wherein, Identifying one or more types of objects in the scene includes applying a keypoint detection algorithm to the image.
5. The method according to any one of claims 1 to 4, wherein: Estimating the surface shape includes: calculating the normal vector of the surface in the depth map, and determining the angle between the normal vector and the vector between the location where the source view image was captured and the surface. Determining the first portion and the second portion includes: comparing the angles determined for different source view images, and assigning a source view image to each portion based on the comparison of the angles.
6. The method according to any one of claims 1 to 4, wherein, Determining the first and second portions involves using the weighted sum of image coordinates per line to determine the seam line separating the first and second portions.
7. The method according to claim 6, wherein, The sum of the weighted image coordinates for each line is weighted according to the normalized depth from the target viewpoint.
8. The method according to any one of claims 1 to 7, further comprising: A low-pass filter is applied to the depth map, and the filtered depth map is used to estimate the surface shape.
9. The method according to claim 8, wherein, Applying the low-pass filter includes applying the low-pass filter to two or more surfaces in the depth map to generate a filtered depth map.
10. The method according to any one of claims 1 to 9, further comprising: Identify one or more objects and / or object features in the scene; And if an object and / or object features are identified between the first part and the second part, the first part and the second part are adapted.
11. The method according to any one of claims 1 to 10, further comprising: Calculate the color difference between corresponding colors in the first source view and the second source view within the stitched region, wherein the stitched region is the portion of the target view image that overlaps with both the first and second portions; and The colors in the splicing area are adjusted based on the color differences to gradually change the colors from the first part and the second part.
12. A computer program carrier comprising computer program code, which, when run on a computer, causes the computer to perform all the steps according to any one of claims 1 to 11.
13. A system for rendering a target view image of a scene at a target viewpoint, the system comprising a processor configured to: Obtain (902) at least two source view images of the scene captured from different locations and at least one depth map of the scene; The source view image and / or the depth map are used to estimate (904) the surface shape within the scene, wherein, The surface shape indicates the gradient of a surface in the scene relative to the location where the source view image was captured, wherein the processor is configured to estimate the surface shape by: identifying one or more types of objects and / or object features in the scene, and assigning a predetermined surface shape to one or more objects and / or object features based on the identified types and / or object features; Determining (906) based on the surface shape and the positions where the first source view image and the second source view image are captured: The first part of the scene to be rendered using the first source view image; and The second part of the scene to be rendered using the second source view image; and The target view image of the scene at the target viewpoint is rendered (908) using the at least one depth map, the first source view image for the first part of the scene, and the second source view image for the second part of the scene.
14. The system according to claim 13, wherein: The processor is configured to estimate the surface shape by: calculating the surface normal vector in the depth map, and determining the angle between the normal vector and the vector between the location where the source view image was captured and the surface. The processor is configured to determine the first portion and the second portion by comparing the angles determined for different source view images and assigning a source view image to each portion based on the comparison of the angles.
15. The system according to any one of claims 13 to 14, wherein, The processor is also configured to: Calculate the color difference between corresponding colors in the first source view and the second source view within the stitched region, wherein the stitched region is the portion of the target view image that overlaps with both the first and second portions; and The colors in the splicing area are adjusted based on the color differences to gradually change the colors from the first part and the second part.