Handling reflections in multi-view imaging
By using multiple images and scene geometry to identify light source pixels and analyze texture values, the problem of reflection processing in multi-view imaging is solved, and natural and accurate image synthesis is achieved.
Patent Information
- Application Number
- CN202380067979.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-09-18
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively process reflections in images, especially in multi-view imaging, resulting in reflections that hinder the effect of the ground surface being represented as a single textured plane, resulting in unnatural synthetic artifacts.
By obtaining multiple images of the scene from multiple image sensors, and identifying the light source pixels using the scene's geometry, analyzing the observed pixels, corresponding pixels and light source pixels, the reflections are removed and the images at new viewpoints with appropriate reflections are synthesized.
It realizes effective removal of reflection in multi-view imaging, avoiding unnatural artifacts caused by reflection, and the generated image is more natural and accurate.
Smart Images

Figure CN119948524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of processing reflections in images. In particular, the present invention relates to removing reflections from an image and synthesizing an image at a new viewpoint with appropriate reflections. Background Art
[0002] Indoor scenes, such as sports halls, often present highly reflective ground surfaces. This causes problems for automatic 3D scene reconstruction and re-rendering from cameras. These reflections particularly hinder the representation of the ground surface as a single textured plane. Existing texturing techniques cannot handle reflections and therefore often produce synthetic artifacts that look unnatural.
[0003] A commonly known solution is to use a polarizer in front of the lens of the camera in order to block the reflected light which is usually polarized due to the reflection. However, polarizers are not always used on cameras.
[0004] Another solution could be to use only multi-view depth estimation, but handle reflections separately. However, while reflections can be assigned meaningful depth values (i.e., the depth of the light source from which the light is reflected), the appearance changes rapidly as a function of viewing angle, making it difficult to estimate interpolated and extrapolated view images.
[0005] Thus, there is a need for an improved method of handling reflections, especially for multi-view imaging. Summary of the invention
[0006] The invention is defined by the claims.
[0007] According to an example according to one aspect of the present invention, there is provided a method for removing reflections from an image, the method comprising:
[0008] obtaining two or more images of a scene from one or more image sensors;
[0009] obtaining a geometry of the scene, the geometry providing locations in the scene corresponding to pixels in the image;
[0010] For one or more observed pixels in the image, adjusting the value of each observed pixel by:
[0011] identifying corresponding pixels in other images that correspond to the same location in the scene as the observed pixel;
[0012] identifying a light source pixel corresponding to a location in the scene where a light source is delivering light to a location in the scene corresponding to the observed pixel and the corresponding pixel by using the geometry of the scene and the position of the image sensor relative to the locations of the observed pixel and the corresponding pixel; and
[0013] The observed pixel and the corresponding pixel and the texture value of the light source pixel are analyzed to obtain an adjusted texture value for the observed pixel, wherein reflections are removed.
[0014] The two or more images obtained may be obtained from image sensors at different locations in the scene. Alternatively, the images may be obtained from the same location at different times, or different color bands may be used from the images.
[0015] Certain objects (e.g., a polished floor) often reflect a large amount of light. This causes the object to appear a different color than it actually is when it is being imaged. In multi-view imaging, these reflections cause problems during rendering (e.g., they cause artifacts on rendered reflective surfaces). In particular, when rendering a multi-view frame at a target viewpoint, the reflections from the different images used will all correspond to reflections from different viewpoints, and the resulting image will contain an unnatural mixture of reflections.
[0016] It has been found that these reflections can be removed when using two or more images of the same scene taken from different viewpoints. In particular, this is possible when the image sensors used to obtain the images are placed at different positions relative to each other. This is common in multi-view imaging and therefore conventional workflows for multi-view imaging can be used.
[0017] The idea is that each captured image (from a different viewpoint) will contain different reflections. These reflections are caused by background light sources, which can be luminous objects or surfaces that reflect light towards the objects, and the light is then reflected (or further reflected) towards the image sensor. The light sources causing these reflections can be identified using geometry (where "light source" is used to refer to active or passive light generation as described above), because the position of the camera is known, and the geometry of the scene is also known (e.g., from a point cloud and / or depth map in multi-view imaging). In particular, the reflection angle will be the same from the observed pixel to the camera as from the observed pixel to the light source.
[0018] Therefore, the color of the light source that caused the reflection can be found. It can be assumed that the color of a pixel with a reflection is a combination of the original color of the object in the scene and the color of the light source that has been reflected. How these colors (i.e., texture values) are combined depends on the reflectivity of the object, which is usually unknown.
[0019] However, different images of the same object from different viewpoints provide different levels of reflectance for the same part of the object. In this way, other observed pixels (from other images) corresponding to the same position as the observed pixel can be identified. The original color of the object and its reflectance will be the same for all other observed pixels. In this way, various simultaneous equations containing two unknowns can be constructed from the different images and used to solve for the original color of the object. Therefore, the texture values of the observed pixel and the other observed pixels can be adapted to correspond to the original color of the object (i.e., with the reflection removed).
[0020] Adjusting each observed pixel may also include tracking light from the position of the image sensor to the corresponding observed pixel or corresponding pixel, and simulating reflection of the light from the observed pixel or corresponding pixel. Thus, identifying the light source pixel for the observed pixel and the corresponding pixel includes identifying one or more pixels intersected by the reflected light.
[0021] For example, a ray may be traced from the position of the image sensor to a simulated ground surface. The ground surface may, for example, be a plane estimated for the ground surface of the scene. The observed pixel may be any pixel on the ground surface. In other words, the observed pixel may be a pixel that is intersected by a ray from the image sensor.
[0022] Once the light has intersected the ground surface, the reflection of the light from the observed pixel can be simulated. It should be understood that simulating the reflection of the light usually means reflecting the light with the same angle of incidence. Therefore, the pixels of the object intersected by the reflected light can be marked as light source pixels. For example, the color of the background pixel can be reflected on the ground surface, and thus the background pixel will be regarded as a light source.
[0023] Adjusting each observed pixel may also include analyzing texture values of the corresponding light source pixel and one or more pixels adjacent to the observed pixel along with the observed pixel of the same image under the assumption that the adjacent pixels include similar visual properties as the observed pixel.
[0024] In some cases, more than one light source may be found to cause reflections on an observed pixel. This may increase the number of unknowns, meaning that more information may be needed to solve the original texture of the object. In this way, neighboring pixels can be used. The neighboring pixels are close enough to the observed pixel that it can be assumed that the reflectivity of the neighboring pixel is the same as the reflectivity of the observed pixel. Therefore, the information of such pixels and the corresponding light source pixels identified in the same way as the observed pixels can be used to help solve the original texture.
[0025] Identifying light source pixels may include: identifying specular light source pixels corresponding to locations in the scene where the light source is delivering light to locations in the scene corresponding to the observed pixel and the corresponding pixel, and the light is being reflected to the location of the corresponding image sensor; and identifying non-specular light source pixels adjacent to or proximate to the light source pixels, wherein the non-specular light source pixels correspond to locations in the scene where the light source is delivering light to locations in the scene corresponding to the observed pixel, and the light is expected to be diffused toward the location of the corresponding image sensor, wherein analyzing texture values of the light source pixels includes analyzing texture values of both the specular light source pixels and the non-specular light source pixels.
[0026] Specular reflections present the same reflection angle after the light reflects. However, if the object is not a perfect mirror, non-specular reflections can also occur. Non-specular reflections occur when light is diffused as it is reflected, so a "cone" of light is reflected instead of a straight ray.
[0027] Non-specular reflections can be handled by assuming that the observed pixel reflects light from a set of pixels near the light source (ie, non-specular light source pixels).
[0028] Additionally, the amount of light reflected from each of the non-specular light sources may be weighted (eg, using a Gaussian function) such that the farther the non-specular light source is from the light source pixel, the lower the contribution of the non-specular light source pixel to the texture value of the observed pixel.
[0029] Obtaining the geometry of the scene and the position of the image sensor may include using a structure from motion algorithm on two or more images.
[0030] Obtaining the geometry of the scene may include iteratively fitting one or more surfaces to one or more objects present in the two or more images.
[0031] Obtaining the geometry of the scene may include obtaining depth measurements of the scene from one or more depth sensors. For example, a time-of-flight sensor, a structured light sensor, etc. may be used as a depth sensor.
[0032] The method may also include generating an object texture for an object in the scene using the adjusted texture values of the observed pixels corresponding to the object.
[0033] The method may further include generating a reflectivity map for the object based on the observed texture values of the pixels compared to corresponding adjusted texture values.
[0034] The method may further include sending the object texture and the reflectivity map for virtual view synthesis.
[0035] The object texture and the reflectivity map can be used to render the scene with new accurate reflections. This is because the object texture will not contain any view-dependent reflections, and therefore warping it to the target viewpoint will not generate artifacts due to view-dependent reflections. Thereafter, appropriate reflections can be added (i.e., from the target viewpoint) since the reflectivity value of the object will be known.
[0036] Virtual view synthesis involves the synthesis of new images at different virtual viewpoints.
[0037] The bitstream may include one or more object textures corresponding to one or more objects in the scene, a geometry of the scene, and reflectivity values of at least one of the object textures, the object textures and geometry being suitable for synthesizing a new image at a different virtual viewpoint.
[0038] The present invention also provides a method for synthesizing an image at a target viewpoint, the method comprising:
[0039] receiving a first object texture, a reflectivity value for the first object texture, and a second object texture, wherein the first object texture and the second object texture correspond to objects in a scene;
[0040] receiving a geometry of the scene including the first object texture and the second object texture;
[0041] synthesizing an image having the first object texture and the second object texture at the target viewpoint using the geometry of the scene;
[0042] simulating a reflection on the first object texture based on the reflectivity value and a texture value of the second object texture; and
[0043] A third object texture is received and added to the first object texture with simulated reflection to synthesize the image at the target viewpoint.
[0044] The method may also include receiving a reflectivity map including a plurality of reflectivity values for a plurality of pixels of the first object texture.
[0045] The method may further include simulating reflection on the first object texture based on the reflectivity value and a texture value of the third object texture.
[0046] Simulating reflection may include: tracing a ray from the position of the target viewpoint to a pixel on a first object texture; simulating reflection of the ray from the first object texture; identifying a pixel on a second (or third) object texture that is intersected by the reflected ray; and adjusting a texture value of the first object texture based on the texture value of the intersected pixel. The simulated reflection of the ray may be based on specular reflection and / or non-specular reflection.
[0047] The present invention also provides a computer program product comprising a computer program code, which, when executed on a processor, causes the processor to execute all the steps of the aforementioned method.
[0048] The present invention also provides a processor configured to execute the above-mentioned computer program code.
[0049] The present invention also provides a computer-readable data carrier carrying the above-mentioned computer program code.
[0050] For example, the computer readable data carrier may be a storage medium carrying (ie storing) the computer program code or a bit stream / signal carrying the computer program code.
[0051] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] For a better understanding of the invention and to show more clearly how it may be put into practice, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0053] Figure 1 A cross section of the scene is shown;
[0054] Figure 2 Shows Figure 1 An image of a scene;
[0055] Figure 3 illustrates how two cameras image a given point on the Earth's surface;
[0056] Figure 4 A cross section of a scene with three candidate background walls is shown;
[0057] Figure 5 A cone showing possible trajectories of light diffused at the ground surface;
[0058] Figure 6 An image having a specular light source pixel and a plurality of non-specular light source pixels on a background wall is shown; and
[0059] Figure 7 A method for removing reflections from an image is shown. DETAILED DESCRIPTION
[0060] The present invention will be described with reference to the accompanying drawings.
[0061] It should be understood that the detailed description and specific examples are intended to be for illustration purposes only and are not intended to limit the scope of the invention when indicating exemplary embodiments of the device, system and method. These and other features, aspects and advantages of the device, system and method of the present invention will be better understood from the following description, claims and drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar parts.
[0062] The present invention provides a method for removing reflections from an image. Images of a scene at different viewpoints and the geometric structure of the scene are obtained. For observed pixels in the images, the observed pixels are adjusted by the following operations: identifying other pixels in other images, the other pixels corresponding to the same positions in the scene as the observed pixels, and identifying light source pixels, the light source pixels corresponding to the following positions in the scene, at which the light source is delivering light to the positions in the scene corresponding to the observed pixels and the other pixels. The texture values of the observed pixels, the other pixels, and the light source pixels are analyzed to obtain adjusted texture values with the reflections removed for the observed pixels.
[0063] Figure 1 1 shows a cross section of a scene. Light from background wall 106, for example, reflects toward ground surface 104. Some of the light reflected from ground surface 104 will be reflected toward image sensor 102 (i.e., camera). As a result, the color (I o ) will appear at the camera's location as the true color of the ground surface 104 at location 108 (I g ) and the color of the background surface 106 at position 110 (I s ) is a specific combination of .
[0064] Therefore, it is proposed to remove reflections from a textured 3D surface (ie the imaging surface) by identifying pixels in the image from the camera 102 that cause reflections in the image.
[0065] Figure 2 Shows Figure 1 The texture / color of the background wall 204 and the ground surface 202 can be seen in the image 200. Prior knowledge of the geometry of the scene is used to identify pixels in the source view image 200 that are likely to be light emitters that cause a reflection of a given pixel that images a point on the reflective ground surface 202. For example, knowledge of the geometry of the ground surface 202 relative to the background wall 204 (e.g., Figure 1As shown in FIG. 1 , the light source pixel can be used to identify the light source pixel. In this case, for the observed pixel I o , the light source pixel I has been identified s .
[0066] The color value of a given light source pixel I s and the observed color value I of the reflecting surface pixel o , a reflectance model can be formulated. The reflectance model aims to predict the color of the observed pixel by asserting that o is the original color of the ground surface I g and the color of the light source pixel weighted by the reflectivity α of the ground surface s The color of the ground surface 202 without reflection is estimated by combining:
[0067] I o =αI g +(1-α)I s (1)
[0068] Therefore, there are two unknown quantities: α and I g However, in multi-view imaging, there are usually multiple cameras imaging the same scene from different viewpoints. This enables the geometry of the scene to be obtained from the images of the scene. However, it has been recognized that having multiple images of the scene from different viewpoints also enables the removal of reflections from the images.
[0069] Since multiple cameras view the same surface point, the model equations can be combined across cameras. In fact, when two cameras are used, a closed-form solution for the surface color is produced. More cameras can help compute a robust solution that can also handle uncertainty in the location of light source pixels. The inferred reflection properties can be sent to the client device so that during rendering, the reflections can (optionally) be re-rendered.
[0070] Figure 3 The diagram illustrates how two cameras image a given point on the ground surface 302. The two images 300a and 300b have the same scene from different viewpoints. Due to the different viewpoints, each camera sees a reflection of a different light source from the background wall 304 at the same point on the ground surface 302. The first image 300a shows the reflection of the light source pixel. The second image 300b shows the reflection of the light source Therefore, the reflectance model shown in equation (1) can be written for the observed pixel color of each camera image as:
[0071]
[0072] Where the parameter α is a surface parameter that models the reflectivity of the ground surface plane. If α = 1, there is no reflection at all, while if α = 0, the ground surface is a perfect mirror. Equations (2) and (3) can be rewritten as the color of the ground surface (I g ) solution:
[0073]
[0074] Without loss of generality, the following substitutions can be used:
[0075]
[0076] Equations (6), (7) and (8) can be substituted into equations (4) and (5) to obtain:
[0077]
[0078] Note that E (1) and E (2) are based on color values that can be measured from images 300a and 300b. Thus, the only unknown quantity in this pair of equations is the color of the ground surface. g and the value of β. Solving equation (10) for β and plugging it into equation (9) gives:
[0079]
[0080] Therefore, for both views, the surface brightness / color can be solved in closed form. The resulting color I g is the ground surface color with reflections removed. It can therefore be used to generate a single (view-invariant) texture for ground surface geometry. Note that there may still be non-specular contribution in the result. Non-specular is handled below.
[0081] For multi-view imaging, there may be more than two cameras imaging the scene. Therefore, for k = [1, ..., N] input views, the reflection model can be written as:
[0082]
[0083] Equation (12) is a standard linear regression problem of the form y = ax + b. Thus, for example, β and I can be solved directly (i.e., non-iteratively) under the Gaussian error assumption. g The maximum likelihood estimate of .
[0084] So far, it has been assumed that the scene geometry is known a priori and has no (large) errors. This is usually a reasonable assumption for the ground surface, and so finding the observed pixel color in view k is After extrinsic calibration of the N input views (e.g. using existing Structure from Motion (SfM) techniques), the resulting point cloud can be used to fit the ground surface. To this end, typically, a priori knowledge of the camera's height above the ground surface can be used to define a pivot point for plane fitting. Alternatively, a direct image-based multi-view error metric can be used in conjunction with a starting position and an iterative search procedure to fit the ground surface plane.
[0085] Given the knowledge of the positions of 3D ground points and 3D background points in the scene, we can find the source light in the defined view k pixels. However, the background geometry may be less well known or less accurate than the ground surface. Therefore, describing the background scene geometry with a single plane may be too simple and therefore relatively inaccurate. A solution may be to consider the uncertainty of the background depth in the equations to be solved.
[0086] Figure 4 A cross section of a scene with three candidate background walls 106a, 106b, and 106c is shown. By tracking light rays from each camera through the ground floor 104 at point 108 (which reflect at the same angle toward the background), a plurality of candidate light source positions 110a, 110b, and 110c can be identified, each corresponding to a different background wall 106a, 106b, and 106c having a different depth value (relative to the camera 102).
[0087] In reality, the background wall may not be flat. However, the different background walls 106a, 106b and 106c at different depth values provide a solution to the depth variation in the background.
[0088] Therefore, the reflection model from equation (12) can generally be expanded to:
[0089]
[0090] Among them, for the background depth z m Where , there are m fixed depth levels.
[0091] Therefore, the background depth can deviate slightly from the nominal model. This is the well-known multiple linear regression problem with m explanatory variables. Similar to the linear regression problem, it can be solved using existing non-iterative methods.
[0092] Due to the fact that the number of depth values to be tested can easily exceed the number of views k, equation (13) quickly becomes underdetermined. However, this problem can be solved as follows.
[0093] So far, equation (13) applies to a single pixel of an image. However, it is reasonable to assume that the reflectance properties of the ground surface are fairly constant over space due to constant material properties. This means that two adjacent pixels in the same view should have the same solution for βm, since the adjacent pixels must also be related to a very similar background depth z. m This then allows a local system of equations to be constructed for a set of pixels around a given pixel in the source view, where the equations share a value βm. For a window of 3×3 pixels around the center pixel in the source view, nine equations can be constructed. Combined with 8 different views, this gives 72 equations. One can then, for example, select 10 depth values around the background model depth and still be able to solve the overdetermined system of equations provided in equation (13).
[0094] Points on the ground surface model are typically not visible in all views because objects (such as sports players) that exist on the same ground surface may occlude the points in each view. To handle this, multi-view depth estimation is typically preceded by removing reflections. Depth pixels of objects that extend above the ground surface can be mapped onto the ground surface to determine for each view whether the ground surface point is visible. If the ground surface point is not visible, the corresponding equation can be discarded from the system of equations.
[0095] In many cases, the ground surface will not be a perfect mirror. Figure 5 A cone 502 is shown of possible trajectories of light that is diffused at the ground surface 104. This is often referred to as diffuse or non-specular reflection. This means that multiple background light sources may contribute to the color of the observed ground surface point. To address this issue, the reflectivity model for k views can now be described as:
[0096]
[0097] Therein, a 2D kernel of Δ pixels is used around a background point 110 that would correspond to an ideal specular ray, and off-axis rays are weighted less via a Gaussian function f, depending on the angular deviation from the specular ray. The kernel size Δ[pixel] is typically adjusted to the Gaussian spread parameter σ. Large values of σ are used when the ground surface deviates more from a pure specular surface. This parameter can be fixed a priori or fitted with a least squares procedure.
[0098] Figure 6An image 600 is shown with a specular light source pixel 608 and a plurality of non-specular light source pixels 610 on a background wall 604. Light from the specular light source pixel 608 is reflected directly from a point 606 on the ground surface 602 toward the camera. At the same time, light from the non-specular light source pixels is expected to diffuse from the point 606 toward the camera. The larger the diffusion angle, the lower the expected intensity of the light reflected via diffuse reflection. This is why equation (15) approximates the effect of the non-specular light source pixel 610 on the observed pixel using a Gaussian function that depends on the angular deviation from the specular light. The reason for the contribution.
[0099] For a particular view, ground surface pixels 606 may reflect light from a light source that is not observable in the image itself (e.g., covered objects). In this case, knowledge of the location, brightness, and color of the light source (e.g., from other views) may be injected into the image.
[0100] Once the observed pixel has been Determined ground color I g , the color of the observed pixel can be changed to the ground color, thereby removing the reflection on the observed pixel. This can be done for each pixel on the ground surface to obtain an image without any reflection on the entire ground surface. Of course, reflections can occur in other objects besides the ground surface, and the same method described above can also be used to remove reflections from other objects.
[0101] The removal of reflections can be applied to generate textures for ground surface planes, where the texture is sent as an image or video. Other objects (such as athletes, balls and / or baskets) can be sent as multi-view images with depth. Upon receiving the information, the rendering algorithm usually first draws the background and ground surface, and then performs tile-based rendering using multi-view blending and synthesis in a back-to-back order. It may be desirable to remove the ground surface from which it is reflected. However, the reflection effect is still a desired effect due to its increased realism.
[0102] Since the background texture and scene geometry will also be present on the client side, reflections can be reapplied to any composite view using the reflection model:
[0103]
[0104] Note that the sprite background texture I b The view related background items have now been replaced This is a reasonable assumption when the background is relatively far from the camera. band its associated depth are available on the client side, as they are needed during rendering of a new image at the target viewpoint. However, the value of the reflectivity parameter α will now have to be sent. If α varies spatially, a reflectivity map can be used to reflect the α value at different locations on the scene. However, a single scalar value for α may also work.
[0105] Of course, it should be understood that rendering the ground surface non-reflective may enable other computer graphics methods for adding reflections. The particular method used may depend on the available processing resources for synthesizing the image.
[0106] In practice, when synthesizing an image at a target / new viewpoint, pixels from a received object texture (e.g., image, texture patch, ground / background model, etc.) are warped to the new viewpoint. This is achieved using the geometry of the scene. For example, a depth map can be used to warp pixels by deprojecting them to a common coordinate system and reprojecting them to the target viewpoint. A grid or known depth layer for the patches can also be used.
[0107] The background texture can be identified by tracing rays from the new viewpoint toward the ground. b Reflections of light can be simulated, where the light reflects off the ground at the same angle as the incident light from the new viewpoint to the ground. When a reflected ray intersects a (non-transparent) pixel, this indicates that light emitted and / or reflected from the corresponding object will be reflected by the ground and seen at the target viewpoint. In this way, reflections can be added to the ground (or other objects) that are appropriate when seen from any novel viewpoint.
[0108] Reflections of other objects on the ground surface (e.g., players) can be processed in a similar way to the background. The geometry of these objects relative to the scene (e.g., using a depth map or depth layer) is also sent to render these objects. Thus, reflections from these objects, for example, on the ground, can be added after rendering them.
[0109] Advantageously, reflections from background objects can be added first (e.g., before rendering foreground objects). This means that reflections only need to be simulated for the background object(s), which typically cause more reflections than foreground objects (e.g., consider the reflection of the sky on a body of water). Therefore, the foreground objects can be added afterwards. This essentially separates the synthesis of the image into the rendering of the ground / background (with reflections) and the rendering of the foreground objects (e.g., without reflections).
[0110] Of course, in some cases, it may be preferable to also add reflections of some (or all) foreground objects. This will require more processing resources. This can be done after rendering the foreground objects. In some cases, the reflections of the foreground objects can replace the reflections of the background.
[0111] The rendering algorithm used to synthesize the image at the new viewpoint may perform the following steps to add reflections:
[0112] - draw background sprite texture using its mesh or depth map;
[0113] - Draw the ground surface texture using its mesh or depth map;
[0114] - Add reflections caused by background textures to ground surface pixels;
[0115] - drawing objects (players, balls, etc.) that extend above the ground surface; and
[0116] - Adds reflections caused by these objects to ground surface pixels.
[0117] Figure 7 A method for removing reflections from an image is shown. Step 702 includes obtaining multiple images of a scene from different viewpoints. At least two images from different viewpoints are used. Step 704 includes obtaining the geometric structure of the scene. This can be obtained, for example, using a structure from motion algorithm on the images and / or by using a depth map obtained from a depth sensor. The depth map can also be obtained via parallax estimation between the images.
[0118] Steps 706, 708 and 710 are performed on a per-pixel basis. In other words, the following steps may be performed for various different pixels of each image.
[0119] Step 706 involves identifying corresponding pixels that correspond to the same location in the scene. Specifically, observed pixels in an image may be selected and then corresponding pixels may be identified in other images. In the example provided above, the observed pixels and corresponding pixels are described as I for k views. o (k) The knowledge of the scene geometry and the camera position can be used to identify corresponding pixels for all images. A structure-from-motion algorithm can be used to obtain both the camera position and the scene geometry.
[0120] Step 708 involves identifying light source pixels. Light source pixels are pixels that deliver light toward the observed pixel (and corresponding pixels) in such a way that the delivered light is reflected toward the camera. Delivering light may involve emitting light directly toward the ground (e.g., a light bulb) or already reflected light that is further reflected by the ground surface. For specular reflections, the angle of the light between the light source pixel and the observed pixel is the same as the angle of the light between the observed pixel and the camera (e.g., Figure 1 As mentioned previously, both specular and non-specular reflections can be modeled.
[0121] Step 710 involves analyzing the color of the observed pixel (and corresponding pixels) relative to the color of the light source pixel for a different view to obtain an adjusted texture value (referred to as I in the above example) with the reflection removed. g ). Equations (12), (13) and (15) provide how to obtain the adjusted texture value I g The adjusted texture values for various pixels in the image can be used to generate an object texture with reflections removed. Thus, the object texture can be sent in the image instead of patches of the object with reflections.
[0122] When the adjusted texture value is obtained, the reflectivity of the ground surface is usually also obtained. In this way, the reflectivity of the ground surface can be sent together with the object texture and depth (e.g., as metadata).
[0123] Object texture, depth and reflectivity can be used to synthesize an image at the target viewpoint. This can be achieved by using conventional image synthesis (using texture and depth) and adding reflections to the new synthesized image using conventional computer graphics methods. However, this will add significant latency to the generation of the image.
[0124] For immersive video applications and broadcasts, frame rate is an important factor in the viewing experience. As such, in these cases, the additional processing time required to add reflections can significantly impact the viewing experience by reducing the frame rate. This is a particularly significant problem when using, for example, mobile devices with limited processing resources.
[0125] Therefore, it is proposed to synthesize a new image at the target viewpoint in two separate steps. In the first step, two object textures are rendered at the target viewpoint. This is achieved by adding texture values (e.g., color and transparency pixel values) from the target viewpoint to the new image.
[0126] Methods for rendering an object at a target viewpoint will be known. For example, this may involve warping the texture of an object from its source viewpoint (e.g., the viewpoint of a camera that obtains the object's color values) to the target viewpoint using the scene's geometry (e.g., a depth map, planes at various depths, etc.). More generally, the geometry provides the 3D position for the object's texture values. In other words, the geometry enables finding the depth of an object's pixels relative to the target viewpoint.
[0127] Continuing with the first step, the reflection of the second object texture on the first object texture is simulated. The amount of reflection on the first object texture is based on the reflectivity value of the first object texture. Of course, a spatially varying reflectivity map (e.g., with per-pixel reflectivity values) may be used.
[0128] In most cases, the first and second object textures will be the ground surface and the background surface, respectively. This is because it is expected that the ground surface in most scenes will have the most obvious reflections, and these reflections are usually caused by light reflected or emitted from the background of the scene.
[0129] For example, consider the floor of a relatively reflective basketball court (i.e., ground surface), where various light sources illuminate the court (i.e., background). In this case, the basketball court will act as a ground surface that reflects light from the various light sources surrounding the court. The various light sources will act as the background.
[0130] However, the first object texture more generally includes any object or objects to which it is desired to add reflections, and the second object texture more generally includes any object or objects that cause reflections on the object(s) included in the first object texture.
[0131] It should be appreciated that simulating reflections from a second object texture on a first object texture is less computationally intensive than simulating reflections for an entire scene (e.g., with multiple foreground objects). It should also be appreciated that when simulating reflections on a first object texture, the texture values (including color) of the first object texture will be adapted to some combination of the original texture values and the texture values of the second object texture. The particular combination depends on the reflectivity value(s) of the first object texture.
[0132] In a second step, a third object texture is added to the image over the first object texture. This can be achieved by rendering the third object texture (now with reflections) over the first object texture. Typically, the third object texture will likely include mostly foreground objects. Thus, any "object texture" may include multiple objects. For example, in the previous basketball court example, the third object texture may include a basketball player, a ball, and / or a ring. Of course, more than one foreground object may be included in the third object texture.
[0133] It should be understood that the third object texture is not limited to foreground objects, and may include elements of the background / ground surface where reflection is not desired.
[0134] In so doing, this enables synthesis of a new image at a target viewpoint having reflections that accurately reflect the target viewpoint itself, and without requiring as many computational resources as conventional methods for simulating reflections.
[0135] Of course, for a more natural and accurate image, reflections from foreground objects (i.e., included in the third object texture) may also be added to the ground surface (i.e., the first object texture). This may depend on the available computing resources and, for example, the target frame rate. In this way, the foreground objects may be divided into: a third object texture for foreground objects, for which reflections are desired to be added; and a fourth object for foreground objects, for which reflections are not desired to be added. Similarly, each foreground object may have its own separate object texture.
[0136] A skilled person will be easily able to develop a processor for executing any of the methods described herein. Therefore, each step of the flowchart may represent a different action performed by the processor and may be performed by a corresponding module of the processor.
[0137] As described above, the system utilizes a processor to perform data processing. The processor can be implemented in many ways using software and / or hardware to perform the various functions required. The processor typically employs one or more microprocessors that can be programmed using software (e.g., microcode) to perform the required functions. The processor can be implemented as a combination of dedicated hardware that performs some functions and one or more programmed microprocessors and associated circuits that perform other functions.
[0138] Examples of circuits that may be employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).
[0139] In various embodiments, the processor may be associated with one or more storage media, such as volatile and non-volatile computer memory, such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that perform the desired functions when run on one or more processors and / or controllers. The various storage media may be fixed within the processor or controller or may be transportable so that one or more programs stored thereon may be loaded into the processor.
[0140] By studying the drawings, the disclosure and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the word "one" or "an" does not exclude a plurality.
[0141] The functions implemented by a processor may be implemented by a single processor or by multiple individual processing units which may be considered together to constitute a “processor.” In some cases, such processing units may be remote from each other and communicate with each other in a wired or wireless manner.
[0142] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0143] The computer program may be stored / distributed on suitable media, such as optical storage media or solid-state media provided together with or as part of other hardware, but the computer program may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.
[0144] If the term "suitable for" is used in the claims or the description, it should be noted that the term "suitable for" is intended to be equivalent to the term "configured to". If the term "arranged" is used in the claims or the description, it should be noted that the term "arranged" is intended to be equivalent to the term "system", and vice versa.
[0145] Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method for removing reflections from an image for subsequent use in new view synthesis, the method comprising: obtaining (702) two or more images of a scene from one or more image sensors; obtaining (704) a geometry of the scene, the geometry providing locations in the scene corresponding to pixels in the image; obtaining a position of the one or more image sensors; and For each of the one or more observed pixels in the image corresponding to a location in the scene, determining an original color value for each observed pixel by: identifying (706) corresponding pixels in other images that correspond to the same location in the scene as the observed pixel; identifying (708) a light source pixel corresponding to a location in the scene at which a light source is delivering light to a location in the scene corresponding to the observed pixel and the corresponding pixel using the geometry of the scene and the position of the image sensor relative to the locations of the observed pixel and the corresponding pixel; formulating simultaneous equations for the location of the observed pixel, wherein one of the simultaneous equations corresponds to the color value for the observed pixel and the remaining of the simultaneous equations each corresponds to the color value for one of the corresponding pixels, the simultaneous equations being a combination of an unknown original color value and a color value for the light source pixel, the combination depending on unknown surface parameters that model a reflectivity of a ground surface plane at the location in the scene corresponding to the observed pixel; The simultaneous equations are solved to obtain the original color value with reflections removed for the observed pixel.
2. The method according to claim 1, wherein: Determining the original color value for each observed pixel comprises: tracing light rays from a location on the image sensor to a corresponding observed pixel or the corresponding pixel; and simulating said reflection of said light from said observed pixel or said corresponding pixel, Wherein, identifying the light source pixel for the observed pixel and the corresponding pixel comprises: identifying one or more pixels intersected by the reflected light.
3. The method according to claim 1 or 2, wherein: Determining the original color value of each observed pixel includes analyzing the color values of a corresponding light source pixel and one or more pixels adjacent to the observed pixel together with the observed pixel of the same image.
4. The method according to any one of claims 1 to 3, wherein: Identifying light source pixels includes: identifying a specular light source pixel corresponding to a location in the scene where a light source is delivering light to a location in the scene corresponding to the observed pixel and the corresponding pixel, and the light is reflected to a location of a corresponding image sensor; and identifying a non-specular light source pixel adjacent to the light source pixel, wherein the non-specular light source pixel corresponds to a location in the scene at which a light source is delivering light to a location in the scene corresponding to the observed pixel, The simultaneous equations include color values of both the specular light source pixels and the non-specular light source pixels.
5. The method according to any one of claims 1 to 4, wherein: Obtaining the geometry of the scene and the position of the image sensor includes using a structure from motion algorithm on the two or more images.
6. The method according to any one of claims 1 to 5, wherein: Obtaining the geometry of the scene includes iteratively fitting one or more surfaces to one or more objects present in the two or more images.
7. The method according to any one of claims 1 to 6, further comprising: An object texture is generated for an object in the scene using the raw color values of the observed pixels corresponding to the object.
8. The method according to claim 7, further comprising: A reflectance map is generated for the object based on the color values of the observed pixels compared to corresponding original color values.
9. The method according to claim 8, further comprising: The object texture and the reflectivity map are sent for virtual view synthesis.
10. A method for synthesizing an image at a target viewpoint, the method comprising: receiving a first object texture, a reflectivity value for the first object texture, and a second object texture, wherein the first object texture and the second object texture correspond to objects in a scene; receiving a geometry of the scene corresponding to the first object texture and the second object texture; synthesizing an image having the first object texture and the second object texture at the target viewpoint using the geometry of the scene; simulating a reflection on the first object texture based on the reflectivity value and a texture value of the second object texture; and A third object texture is received and added to the first object texture with simulated reflection to synthesize the image at the target viewpoint.
11. The method according to claim 10, further comprising: A reflectivity map is received, the reflectivity map comprising a plurality of reflectivity values for a plurality of pixels of the first object texture.
12. The method according to claim 10 or 11, further comprising: Reflection on the first object texture is simulated based on the reflectivity value and a texture value of the third object texture.
13. A computer program product comprising computer program code which, when executed on a processor, causes the processor to perform all the steps of the method according to any one of claims 1 to 9 and / or all the steps of the method according to any one of claims 10 to 12.
14. A processor configured to execute computer program code according to claim 13.
15. A computer readable data carrier carrying computer program code according to claim 13.