Handling reflections in multiview images

JP2025531019A5Pending Publication Date: 2026-09-04KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025508691
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-09-18
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

Existing methods struggle to effectively handle reflections in multi-view images, particularly in indoor scenes with highly reflective surfaces, leading to unnatural-looking compositing artifacts and challenges in 3D scene reconstruction and re-rendering.

Method used

A method involving multiple images from different viewpoints is used to identify and remove reflections by analyzing texture values and scene geometry, adapting pixel colors to remove reflections, and synthesizing images at new viewpoints with accurate reflections.

Benefits of technology

This approach allows for the generation of reflection-free textured surfaces and accurate virtual view synthesis, reducing computational resources required for rendering and enhancing image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for removing reflections from an image is provided. An image of a scene and a geometry of the scene are obtained. For an observed pixel in the image, the observed pixel is adapted by identifying a corresponding pixel in another image that corresponds to the same location in the scene as the observed pixel, identifying a light source pixel that corresponds to a location in the scene where the light source provides light to the location in the scene corresponding to the observed pixel and the corresponding pixel, and analyzing texture values ​​of the observed pixel and the corresponding pixel and the light source pixel to obtain an adapted texture value for the observed pixel that has reflections removed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of processing reflections in images, and in particular to removing reflections from images and synthesizing images at new viewpoints with appropriate reflections. [Background technology]

[0002] Indoor scenes, such as sports halls, often exhibit highly reflective ground surfaces. This poses a challenge for automatic 3D scene reconstruction and re-rendering from the camera. Reflections in particular hinder the representation of the ground surface as a single textured surface. Existing texture techniques are unable to cope with reflections and often produce unnatural-looking compositing artifacts as a result.

[0003] A commonly known solution is to use a polarizing sheet in front of the camera lens to block reflected light, which is typically polarized due to reflection, but this does not necessarily mean that a polarizer is used on the camera itself.

[0004] Another solution is to simply use multi-view depth estimation and treat reflections separately. However, while reflections can be assigned significant depth values ​​(i.e., the depth of the light source from which light is reflected), their appearance changes rapidly as a function of viewing angle, making it difficult to estimate interpolated and extrapolated view images. Summary of the Invention [Problem to be solved by the invention]

[0005] Therefore, there is a need for improved methods for dealing with reflections, especially for multi-view images. [Means for solving the problem]

[0006] The invention is defined by the claims.

[0007] According to an embodiment of the present invention, there is provided a method for removing reflections from an image, the method comprising: acquiring two or more images of a scene from one or more image sensors; obtaining a geometry of the scene that provides locations in the scene corresponding to pixels in the image; For one or more observed pixels in the image, the value of each observed pixel is calculated by the following steps: identifying corresponding pixels in other images that correspond to the same location in the scene as the observed pixel; using the geometry of the scene and the position of the image sensor relative to the positions of the observed pixel and the corresponding pixel, to identify a source pixel corresponding to a position in the scene from which the light source is shining light into the position in the scene corresponding to the observed pixel and the corresponding pixel; analyzing texture values ​​of the observed pixel, the corresponding pixel, and the light source pixel to obtain an adapted texture value that removes reflections of the observed pixel; and adapting the It has.

[0008] The two or more images acquired can be from image sensors at different locations within the scene, or alternatively, the images can be acquired from the same location at different times, or different color bands can be used from the images.

[0009] Certain objects (e.g., a polished floor) often reflect a large amount of light, causing the object to appear a different color than it actually is when it is imaged. In multi-view images, these reflections cause problems during rendering (e.g., causing artifacts in the rendered reflective surfaces). In particular, when rendering a multi-view frame at a target viewpoint, the reflections from the different images used all correspond to reflections from different viewpoints, and the resulting image contains an unnatural blend of reflections.

[0010] It has been found that these reflections can be removed when two or more images of the same scene taken from different viewpoints are used. In particular, this is possible when the image sensors used to acquire the images are positioned at different positions relative to each other. This is a common occurrence in multi-view imaging, and therefore conventional workflows for multi-view imaging can be used.

[0011] The idea is that each captured image (from a different viewpoint) will contain different reflections. These reflections are caused by background light sources, which may be emitting objects or surfaces that reflect light toward objects, which then reflect (or are further reflected) toward the image sensor. The light sources causing these reflections (here, "light source" is used to indicate active or passive light generation, as explained above) can be identified using geometry when the camera position is known and the scene geometry is also known (e.g., from a depth map and / or point cloud in multi-view imaging). In particular, the reflection angle is the same from the observed pixel to the camera and from the observed pixel to the light source.

[0012] This allows the color of the light source causing the reflection to be determined. The color of a pixel with a reflection can be assumed to be a combination of the original color of the object in the scene and the color of the reflected light source. How these colors (i.e., texture values) are combined depends on the reflectivity of the object, which is generally unknown.

[0013] However, having different images of the same object from different viewpoints provides different levels of reflectivity for the same portion of the object. Therefore, other observed pixels (from other images) corresponding to the same location as the observed pixel can be identified. The original color of the object and its reflectivity are the same for all other observed pixels. Therefore, various simultaneous equations involving two unknowns can be constructed from different images and used to solve for the original color of the object. Therefore, the texture values ​​of the observed pixel and other observed pixels can be adapted to correspond to the original color of the object (i.e., with reflectivity removed).

[0014] Adapting each observed pixel can further include tracing a ray from a location on the image sensor to a corresponding observed pixel or corresponding pixel and simulating a reflection of the ray from the observed pixel or corresponding pixel, and identifying source pixels for the observed pixel and corresponding pixel includes identifying one or more pixels intersected by the reflected ray.

[0015] For example, a ray can be traced from the position of the image sensor to a simulated ground plane. The ground plane can be, for example, a plane estimated as the ground plane of the scene. The observed pixel can be any pixel on the ground plane. In other words, the observed pixel can be a pixel that intersects with a ray from the image sensor.

[0016] When a ray intersects with the ground surface, a reflection of the ray from the observed pixel can be simulated. It will be understood that simulating the reflection of a ray typically means reflecting the ray at the same angle of incidence. Therefore, a pixel of the object that is intersected by the reflected ray can be labeled as a light source pixel. For example, the color of a background pixel can be reflected on the ground surface, and therefore the background pixel is treated as a light source.

[0017] Adapting each observed pixel may further include analyzing texture values ​​of one or more pixels adjacent to the observed pixel and the corresponding light source pixel, along with the observed pixel of the same image, assuming that the adjacent pixels contain similar visual characteristics as the observed pixel.

[0018] In some cases, more than one light source may be found to cause reflections at an observed pixel. This increases the number of unknowns, which means more information is needed to resolve the object's original texture. In such cases, neighboring pixels can be used. Neighboring pixels are close enough to the observed pixel that the reflectance of the neighboring pixels can be assumed to be the same as the reflectance of the observed pixel. Therefore, the information of such pixels, along with the corresponding light source pixels identified in the same way as the observed pixel, can be used to help determine the original texture.

[0019] Identifying light source pixels includes the steps of identifying specular light source pixels corresponding to positions in the scene where a light source is providing light at a position in the scene corresponding to a pixel corresponding to an observed pixel, the light being reflected toward a corresponding image sensor position; and identifying non-specular light source pixels adjacent or nearby the light source pixels, the non-specular light source pixels corresponding to positions in the scene where a light source is providing light at a position in the scene corresponding to the observed pixel, the light being expected to be diffused toward the corresponding image sensor position; and analyzing texture values ​​of the light source pixels includes analyzing texture values ​​of both the specular light source pixels and the non-specular light source pixels.

[0020] Specular reflection occurs when light reflects at the same angle. However, if the object is not a perfect mirror, non-specular reflection can also occur. Non-specular reflection occurs when light is diffused as it is reflected, so that a "cone" of light is reflected instead of a straight ray of light.

[0021] Non-specular reflections can be addressed by assuming that the observed pixel reflects light from a group of pixels close to the light source (i.e., non-specular light source pixels).

[0022] Additionally, the amount of light reflected from each non-specular light source can be weighted (e.g., using a Gaussian function) so that the further the non-specular light source is from the light source pixel, the lower its contribution to the texture value of the pixel observed from the non-specular light source pixel.

[0023] Obtaining the scene geometry and image sensor position can include using structure from motion algorithms on two or more images.

[0024] Obtaining the geometry of the scene can include iteratively fitting one or more surfaces to one or more objects present in two or more images.

[0025] Obtaining the geometry of the scene can include obtaining depth measurements of the scene from one or more depth sensors, such as time-of-flight sensors, structured light sensors, and the like.

[0026] The method may further include generating an object texture for the object in the scene using the adapted texture values ​​of the observed pixels corresponding to the object.

[0027] The method may further include generating a reflectance map of the object based on the texture values ​​of the observed pixels compared to the corresponding adapted texture values.

[0028] The method may further include transmitting object texture and reflectance maps for virtual view synthesis.

[0029] The object texture and reflectance map can be used to render the scene with new accurate reflections. This is because the object texture does not contain any view-dependent reflectance, and therefore warping the object texture to the target viewpoint does not produce artifacts due to view-dependent reflectance. Then, since the object's reflectance value is known, the appropriate (i.e., from the target viewpoint) reflectance can be added.

[0030] Virtual view synthesis involves the synthesis of novel images at distinct virtual viewpoints.

[0031] The bitstream can include one or more object textures corresponding to one or more objects in the scene, reflectance values ​​for at least one of the object textures, and geometry of the scene, the object textures and geometry suitable for synthesizing novel images at distinct virtual viewpoints.

[0032] The present invention also provides a method for synthesizing an image at a target viewpoint, the method comprising: receiving a first object texture, a reflectance value for the first object texture, and a second object texture, the first and second object textures corresponding to an object in the scene; receiving a geometry of a scene including first and second object textures; synthesizing an image with the first and second object textures at the target viewpoint using the geometry of the scene; simulating a reflection on the first object texture based on the reflectance values ​​and texture values ​​of the second object texture; receiving a third object texture; and adding a third object texture onto the first object texture with simulated reflection to synthesize an image at the target viewpoint.

[0033] The method may further include receiving a reflectance map including a plurality of reflectance values ​​for a plurality of pixels of the first object texture.

[0034] The method may further include simulating a reflection on the first object texture based on the reflectance value and a texture value of the third object texture.

[0035] Simulating the reflection can include tracing a ray from a target viewpoint position to a pixel on a first object texture, simulating a reflection of the ray from the first object texture, identifying a pixel on a second (or third) object texture intersected by the reflected ray, and adapting a texture value of the first object texture based on the texture value of the intersected pixel. Simulating the reflection of the ray can be based on specular and / or non-specular reflection.

[0036] The present invention also provides a computer program product comprising computer program code which, when executed on a processor, causes the processor to perform all of the steps of the method described above.

[0037] The present invention also provides a processor configured to execute the above-described computer program code.

[0038] The invention also provides a computer readable data carrier carrying the above-mentioned computer program code.

[0039] For example, the computer readable data carrier may be a storage medium carrying (ie storing) the computer program code or a bitstream / signal carrying the computer program code.

[0040] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0041] For a better understanding of the present invention and to show more clearly how the same may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Figure 1] A diagram showing a cross section of a scene. [Figure 2] FIG. 2 shows an image of the scene in FIG. 1. [Figure 3] A diagram showing how a given point on the Earth's surface is imaged by two cameras. [Figure 4] FIG. 1 shows a cross section of a scene with three candidate background walls. [Figure 5] A diagram showing the cone of possible trajectories of light scattered by the Earth's surface. [Figure 6] FIG. 1 illustrates an image with a specular source pixel and multiple non-specular source pixels on a background wall. [Figure 7] 1 illustrates a method for removing reflections from an image. DETAILED DESCRIPTION OF THE INVENTION

[0042] The present invention will now be described with reference to the drawings.

[0043] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems, and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar parts.

[0044] The present invention provides a method for removing reflections from images. Images of a scene at different viewpoints and the scene geometry are obtained. For an observed pixel in an image, the observed pixel is adapted by identifying other pixels in other images that correspond to the same location in the scene as the observed pixel, and identifying light source pixels that correspond to locations in the scene where a light source is providing light to the locations in the scene corresponding to the observed pixel and the other pixels. Texture values ​​of the observed pixel, the other pixels, and the light source pixels are analyzed to obtain an adapted, reflection-removed texture value for the observed pixel.

[0045] 1 shows a cross section of a scene. Light from a background wall 106 is reflected, for example, towards a ground surface 104. A portion of the light reflected from the ground surface 104 is reflected towards the image sensor 102, i.e., the camera. Thus, the color (I o ) is the actual color (I g ) and the color of the background surface 106 at position 110 (I s ) appears to be some combination between

[0046] It is therefore proposed to remove reflections from a textured 3D surface (i.e., the imaged surface) through identification of pixels in an image from the camera 102 that caused the reflection in the image.

[0047] Figure 2 shows an image 200 of the scene of Figure 1. The texture / color of a background wall 204 and a ground surface 202 can be seen in the image 200. Prior knowledge of the scene geometry is used to identify pixels in the source view image 200 that are likely to be illuminants causing a reflection for a given pixel imaging a point on the reflective ground surface 202. For example, knowledge of the geometry of the ground surface 202 relative to the background wall 204 (as shown in Figure 1) can be used to identify the light source pixel. In this case, the observed pixel I o About the light source pixel Is has been identified.

[0048] Color value of the light source pixel I s and the observed color value of the reflective surface pixel I o Given, we can formulate a reflectance model, which is the observed pixel color I o However, the original color of the ground surface is g and the color of the light source pixel I weighted based on the ground surface reflectance α s The goal is to estimate the color of the reflection-free ground surface 202 by assuming that the combination of: I o = αI g + (1 - α)I s (1)

[0049] Therefore, there are two unknowns, α and I g However, in multi-view imaging, there are typically multiple cameras capturing the same scene from different viewpoints. This allows the scene geometry to be obtained from the images of the scene. However, it has been recognized that having multiple images of the scene from different viewpoints also allows for the removal of reflections from the images.

[0050] Because multiple cameras observe the same surface points, the model equations can be combined across cameras. In fact, using two cameras provides a closed-form solution for the ground color. More cameras help compute a robust solution that can also deal with uncertainty in the location of the light source pixels. The inferred reflectance properties can be sent to the client device so that the reflections can (optionally) be re-rendered during rendering.

[0051] Figure 3 shows how a given point on the ground plane 302 is imaged by two cameras. The two images 300a and 300b are of the same scene from different viewpoints. Due to the different viewpoints, each camera sees a different light source reflected from a background wall 304 on the same point on the ground plane 302. The first image 300a captures the source pixel I s (1) and the second image 300b shows the reflection of the light source I s (2) Therefore, the reflectance model shown in equation (1) can be written in terms of the observed pixel color of each camera image as I o (1) = α I g + (1-α) I s (1) (2) I o (2) = α I g + (1-α) I s (2) (3)

[0052] Here, the parameter α is a surface parameter that models the reflectance of the ground plane. If α=1, there is no reflection at all, and if α=0, the ground plane is a perfect mirror. Equations (2) and (3) are used to calculate the color of the ground plane (I g ) can be rewritten as follows: TIFF2025531019000002.tif38151

[0053] Without loss of generality, the following substitutions can be used: TIFF2025531019000003.tif46146

[0054] By substituting equations (6), (7) and (8) into equations (4) and (5), the following equation can be obtained: I g = I s (1) + β E (1) (9) I g = I s(2) + β E (2) (10)

[0055] E (1) and E (2) Note that both σ and σ are based on measurable color values ​​from images 300a and 300b. Therefore, the unknown values ​​in this set of equations are the surface color I g and the value of β. Solving equation (10) for β and inserting it into equation (9) gives us TIFF2025531019000004.tif30152

[0056] Therefore, for the two views, the brightness / color of the ground plane can be solved in closed form. The resulting color I g is the ground color with reflections removed. It can therefore be used to generate a single (view-invariant) texture of the ground geometry. Note that non-specular reflection contributions may still be present in the result. Non-specular reflections are dealt with below.

[0057] In multi-view imaging, it is likely that more than two cameras image the scene. Thus, for k = [1, .., N] input views, the reflection model can be written as I s (k) = - β E (k) + I g (12)

[0058] Equation (12) is a standard linear regression problem of the form y = ax + b. Thus, for example, under the Gaussian error assumption, β and I g The maximum likelihood estimate of can be solved directly (i.e., non-iteratively).

[0059] So far, it has been assumed that the scene geometry is known a priori and free from (large) errors. This usually means that the ground plane, and hence the observed pixel color I o (k)is a reasonable assumption about the location where to find the ground plane. For example, after external calibration of N input views using existing Structure from Motion (SfM) techniques, the resulting point cloud can be used to fit a ground plane. To do this, a priori knowledge of the camera's height above the ground plane can typically be used to define a pivot point for plane fitting. Alternatively, the ground plane can be fitted using a direct image-based multiview error metric in combination with a starting position and an iterative search procedure.

[0060] Given knowledge of the locations of 3D ground and background points in the scene, we find the source light I in view k. s (k) However, the background geometry may be less well known or may be available with less accuracy than the ground plane. Therefore, describing the background scene geometry using a single plane would be too simplistic and therefore relatively inaccurate. A solution may be to account for the uncertainty in the background depth in the equation to be solved.

[0061] 4 shows a cross section of a scene with three candidate background walls 106a, 106b, and 106c. By tracing light rays from each camera through the ground plane 104 at point 108 and reflecting at the same angle towards the background, multiple candidate light source positions 110a, 110b, and 110c can be identified, each corresponding to a different background wall 106a, 106b, and 106c with a different depth value (relative to camera 102).

[0062] In reality, the background walls may not be flat, however, each background wall 106a, 106b and 106c at a different depth value provides a solution to the depth variations in the background.

[0063] Therefore, the reflection model from equation (12) can be generally extended to: TIFF2025531019000005.tif33151

[0064] Here, the background depth z m There are m fixed depth levels that can exist as

[0065] Therefore, the background depth may deviate somewhat from the nominal model. This is a well-known multilinear regression problem with m explanatory variables. As with linear regression problems, it can be solved using existing non-iterative methods.

[0066] Equation (13) quickly becomes underdetermined due to the fact that the number of depth values ​​to test can easily exceed the number of views k. However, this problem can be solved as described below.

[0067] So far, equation (13) applies to a single pixel in an image. However, it is reasonable to assume that the reflectance properties of the ground surface are fairly constant across space due to constant material properties. This means that two adjacent pixels in the same view will have very similar background depths z. m Therefore, β m This allows the construction of a local set of equations for a group of pixels around a given pixel in the source view, where the equations have the value β m For a 3x3 pixel window around the central pixel in the source view, it is possible to construct 9 equations. In combination with 8 different views, this gives 72 equations. Then, for example, we can choose 10 depth values ​​around the background model depth and solve the overdetermined system of equations provided in equation (13).

[0068] A point on the ground plane model is usually not visible in all views because an object, such as a sports athlete, that exists on the same ground plane may occlude the point in each view. To address this, multi-view depth estimation is usually performed before reflection removal. Depth pixels of objects that extend on the ground plane can be mapped onto the ground plane to determine whether the ground plane point is visible for each view. If the ground plane point is not visible, the corresponding equation can be dropped from the equation system.

[0069] In many situations, the Earth's surface is not a perfect mirror. Figure 5 shows a cone 502 of possible trajectories of light scattered by the Earth's surface 104. This is commonly known as diffuse or non-specular reflection. This is because multiple background light sources can cause the observed color I of a surface point to be different from the observed color I of a surface point. o (k) To address this, the reflectance model for k views is It can be written as TIFF2025531019000006.tif31155.

[0070] A 2D kernel of Δ pixels around the background point 110 corresponding to the ideal specular ray is used, and off-axis rays are weighted less via a Gaussian function f that depends on the angular deviation from the specular ray. The kernel size Δ [pixels] is typically adjusted to the Gaussian spread parameter σ. Larger values ​​of σ are used when the surface deviates more from a pure specular surface. This parameter can be fixed a priori or fitted by least squares.

[0071] Figure 6 shows an image 600 with a specular light source pixel 608 and multiple non-specular light source pixels 610 on a background wall 604. Light from the specular light source pixel 608 is expected to reflect directly from point 606 on the ground surface 602 towards the camera, while light from the non-specular light source pixels is expected to diffuse from point 606 towards the camera. The larger the diffusion angle, the lower the expected intensity of the diffusely reflected light. This is why observed pixel I o (k)This is why equation (15) uses a Gaussian function that depends on the angular deviation from the specular ray to approximate the contribution of the non-specular light source pixel 610 to .

[0072] For a particular view, it is possible that ground pixels 606 reflect light from a light source that is not observable in the image itself (e.g., obscured by an object), in which case knowledge of the light source's location, brightness, and color (e.g., from other views) can be injected into the image.

[0073] Observed pixel I o (k) Against ground color I g Once determined, the color of the observed pixel can be changed to the color of the ground, thereby removing reflections on the observed pixel. This is done for each pixel on the ground, resulting in an image of the entire ground surface that is completely free of reflections. Of course, reflections may occur on other objects besides the ground, and the same method described above can be used to remove reflections from other objects.

[0074] The reflection removal can be applied to generate a ground plane texture, which is transmitted as an image or video. Other objects, such as athletes, balls, and / or baskets, can be transmitted as multi-view images with depth. Upon receiving the information, the rendering algorithm typically first draws the background and ground plane, then performs patch-based rendering using multi-view blending and compositing in a back-to-front order. A ground plane with reflections removed may be desirable; however, a reflection effect may still be a desirable effect due to its added realism.

[0075] Since the background texture and scene geometry also reside on the client side, the following reflection model can be used to reapply reflections to any composited view. I o (k) = α I g+ (1-α) I b (16)

[0076] Sprite Background Texture I b is the view-dependent background term I s (k) Note that the color I has been replaced by , which is a reasonable assumption to make when the background is relatively far away from the camera. b and its associated depth are available on the client side as they are needed during rendering of the new image at the target viewpoint. However, the value of the reflectance parameter α must be transmitted. If α varies spatially, a reflectance map can be used to reflect α values ​​for different positions on the scene. However, a single scalar value for α can also work.

[0077] Of course, it will be appreciated that other computer graphics methods for adding reflections are possible if a non-reflective ground surface is available, and the particular method used will depend on the processing resources available for compositing the image.

[0078] In practice, when synthesizing an image at a target / novel viewpoint, pixels from the received object texture (e.g., image, texture patch, ground / background model, etc.) are warped to the novel viewpoint. This is achieved using the scene geometry. For example, a depth map can be used to warp pixels by back-projecting them to a common coordinate system and then re-projecting them to the target viewpoint. Meshes or known depth layers for patches can also be used.

[0079] Background Texture I bcan be determined by tracing a ray from the new viewpoint towards the ground. The reflection of the ray can be simulated so that the ray reflects off the ground at the same angle as the incident ray from the new viewpoint to the ground. If the reflected ray intersects a (non-transparent) pixel, this indicates that light emitted and / or reflected from the corresponding object is reflected by the ground and observed at the target viewpoint. Thus, reflections can be added to the ground (or other objects) that are appropriate when viewed from any novel viewpoint.

[0080] Reflections of other objects (e.g., athletes) on the ground surface can be treated in a similar manner to the background. The geometry of these objects relative to the scene (e.g., using a depth map or depth layer) is also transmitted to render these objects. Thus, for example, reflections from these objects on the ground can be added after rendering these objects.

[0081] Advantageously, reflections from background objects can be added first (e.g., before rendering foreground objects). This means that reflections only need to be simulated for background objects, which typically cause more reflections than foreground objects (e.g., consider the reflection of sky over a body of water). Foreground objects can therefore be added later. This essentially separates image compositing into rendering the ground / background (with reflections) and rendering the foreground objects (e.g., without reflections).

[0082] Of course, in some cases it may be preferable to also add reflections of some (or all) foreground objects, which requires more processing resources. This can be done after the foreground objects have been rendered. In some cases, the reflections of the foreground objects can replace the reflections of the background.

[0083] A rendering algorithm for compositing an image at a new viewpoint can perform the following steps to add reflections: - Draw background sprite textures using meshes or depth maps; - Draw ground textures using meshes or depth maps; - Adding background texture induced reflections to ground pixels; - Draw objects (athletes, balls, etc.) that extend over the ground; - Adding the reflections induced by these objects to the ground pixels.

[0084] 7 illustrates a method for removing reflections from an image. Step 702 involves acquiring multiple images of a scene from different viewpoints. At least two images from different viewpoints are used. Step 704 involves acquiring the geometry of the scene. This can be obtained, for example, using a structure from motion algorithm on the image and / or by using a depth map acquired from a depth sensor. The depth map can also be obtained via disparity estimation between the images.

[0085] Steps 706, 708 and 710 are performed for each pixel. In other words, the following steps can be performed for various different pixels per image.

[0086] Step 706 involves identifying corresponding pixels that correspond to the same location in the scene. In particular, a certain observed pixel in an image is selected, and a corresponding pixel is identified in another image. In the example provided above, the observed pixel and corresponding pixel are found to be I for k views. o (k) Using knowledge of the scene geometry and camera position, corresponding pixels can be identified for all of the images. Using the structure from the motion algorithm, both the camera position and the scene geometry can be obtained.

[0087] Step 708 involves identifying a source pixel. A source pixel is a pixel that sources light toward the observed pixel (and corresponding pixel) so that the source light is reflected toward the camera. Source light can include emitting light directly toward the ground (e.g., a light bulb) or previously reflected light (which is further reflected by the ground surface). In the case of specular reflection, the angle of the ray between the source pixel and the observed pixel is the same as the angle of the ray between the observed pixel and the camera (as shown in FIG. 1). As previously mentioned, specular and non-specular reflections can be modeled.

[0088] Step 710 analyzes the color of the observed pixel (and corresponding pixels) relative to the color of the light source pixel for different views to generate an adapted texture value (I in the above example) with the reflections removed. g Equations (12), (13) and (15) involve obtaining the adapted texture value I g The adapted texture values ​​for various pixels in the image can be used to generate an object texture with reflections removed. Thus, the object texture can be transmitted instead of patches of the object with reflections in the image.

[0089] When obtaining the adapted texture values, the ground reflectance is also typically obtained, and thus the ground reflectance can be transmitted along with the object texture and depth (e.g., as metadata).

[0090] Object texture, depth, and reflectance can be used to synthesize an image at the target viewpoint. This can be achieved by using traditional image synthesis (using texture and depth) and adding reflections to the newly synthesized image with traditional computer graphics methods. However, this adds significant delay to the generation of the image.

[0091] For immersive video applications and broadcasts, frame rate is a critical factor in the viewing experience. Therefore, in these cases, the additional processing time required to add reflections can significantly affect the viewing experience by reducing the frame rate. This is a particularly significant issue when using, for example, mobile devices with limited processing resources.

[0092] Thus, we propose to synthesize a new image at a target viewpoint in two separate steps. In the first step, two object textures are rendered at the target viewpoint. This is achieved by adding texture values ​​(e.g., color and transparency pixel values) to the new image from the target viewpoint.

[0093] Methods for rendering objects at a target viewpoint are known. For example, this can involve using the scene's geometry (e.g., a depth map, planes at various depths, etc.) to warp the object's textures from their source viewpoint (e.g., the viewpoint of the camera from which the object's color values ​​were obtained) to the target viewpoint. More generally, the geometry provides the 3D location of the object's texture values. In other words, the geometry makes it possible to find the depth of the object's pixels relative to the target viewpoint.

[0094] Continuing with the first step, the reflection of the second object texture on the first object texture is simulated. The amount of reflection on the first object texture is based on the reflectance values ​​of the first object texture. Of course, a spatially varying reflectance map (e.g., with a reflectance value per pixel) can be used.

[0095] In most cases, the first and second object textures are the ground and background surfaces, respectively, because the ground surface in most scenes is expected to have the most prominent reflections, and these reflections are typically caused by light reflected from or emitted by the scene background.

[0096] For example, consider the floor (i.e., ground surface) of a basketball court, which is relatively reflective and has various light sources (i.e., background) illuminating the court. In this case, the basketball court acts as a ground surface that reflects light from the various light sources around the court. The various light sources act as background.

[0097] However, the first object texture more generally includes any one or more objects to which you want to add reflection, and the second object texture more generally includes any one or more objects that will cause reflections on the objects included in the first object texture.

[0098] It will be appreciated that simulating reflections from a second object texture on a first object texture is much less computationally intensive than simulating reflections of an entire scene (e.g., involving multiple foreground objects). It will also be appreciated that when reflections are simulated on a first object texture, the texture values ​​(including color) of the first object texture are adapted to some combination of the original texture values ​​and the texture values ​​of the second object texture. The particular combination depends on the reflectance values ​​of the first object texture.

[0099] In the second step, a third object texture is added to the image over the first object texture. This can be achieved by rendering the third object texture over the first object texture (here with reflections). In general, the third object texture is likely to contain mostly foreground objects. Any "object texture" can contain multiple objects. For example, in the basketball court example mentioned above, the third object texture could contain basketball players, a ball, and / or a hoop. Of course, more than one foreground object can be included in the third object texture.

[0100] It will be appreciated that the third object texture is not limited to foreground objects only, but may also include background / ground elements where reflection is not required.

[0101] This therefore makes it possible to synthesize novel images at a target viewpoint with reflections that accurately reflect the target viewpoint itself and do not require as many computational resources as traditional methods for simulating reflections.

[0102] Of course, for a more natural and accurate image, reflections from foreground objects (i.e., included in a third object texture) can also be added to the ground plane (i.e., the first object texture). This depends on the available computational resources and, for example, the target frame rate. In this way, the foreground objects can be divided into a third object texture for foreground objects to which reflections are desired and a fourth object texture for foreground objects to which reflections are not desired. Similarly, each foreground object can have its own separate object texture.

[0103] Those skilled in the art can easily develop a processor to perform any of the methods described herein. Accordingly, each step in the flowchart may represent a respective operation performed by a processor, and may be performed by a respective module of the processor.

[0104] As described above, the system utilizes a processor to process data. The processor may be implemented in a variety of ways using software and / or hardware to perform the various functions required. The processor typically uses one or more microprocessors that are programmed using software (e.g., microcode) to perform the required functions. The processor may also be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0105] Examples of circuitry that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).

[0106] In various implementations, the processor may be associated with one or more storage media, e.g., volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. The various storage media may be mounted within the processor or controller, or may be transportable such that one or more programs stored on the storage media can be read by the processor.

[0107] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0108] The functions implemented by a processor may be implemented by a single processor or by multiple separate processing units, which may be considered to constitute a "processor". Such processing units may be remote from each other and may communicate with each other via wired or wireless means.

[0109] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0110] The computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0111] When the term "adapted for" is used in the claims or description, it is meant to be equivalent to the term "configured to." When the term "apparatus" is used in the claims or description, it is intended to be equivalent to the term "system," and vice versa.

[0112] Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A method for removing reflections from an image for later use in a new view composite, The steps include acquiring two or more images of a scene from one or more image sensors, The steps include obtaining the geometry of the scene, which provides the position in the scene corresponding to a pixel in the image, The steps include: acquiring the position of one or more image sensors, For each of the one or more observed pixels in the image corresponding to the position in the scene, the following steps are taken: - A step of identifying a corresponding pixel in another image that corresponds to the same position as the observed pixel in the scene, - Using the geometry of the scene and the position of the image sensor relative to the observed pixels and the corresponding pixels, a step of identifying a light source pixel corresponding to a position in the scene from which a light source is supplying light to the position in the scene corresponding to the observed pixels and the corresponding pixels, - A step of formulating a system of equations for the position of the observed pixel, wherein one of the equations corresponds to the color value of the observed pixel, the remaining equations each correspond to one of the color values ​​of the corresponding pixel, the system of equations is a combination of an unknown original color value and the color value of the light source pixel, and the combination depends on an unknown surface parameter that models the reflectance of the ground surface in the scene at the position corresponding to the observed pixel. For the observed pixels, the step of solving the simultaneous equations to obtain the original color value with the reflection removed, This involves the step of determining the original color value of each observed pixel, A method of having.

2. The step of determining the original color value of each observed pixel is: Tracing a light ray from the position of the image sensor to the corresponding observed pixel and the corresponding pixel, This includes simulating the reflection of the light ray from the observed pixel or the corresponding pixel, The method according to claim 1, wherein the step of identifying light source pixels for the observed pixels and the corresponding pixels includes identifying one or more pixels intersected by reflected light rays.

3. The method according to claim 1, wherein the step of determining the original color value of each observed pixel includes analyzing the color values ​​of one or more pixels adjacent to the observed pixel and corresponding light source pixels in the same image as the observed pixel.

4. The step of identifying the light source pixel is The steps include identifying a specular light source pixel corresponding to the position in the scene where a light source supplies light to the observed pixel and the corresponding pixel, and the light is reflected to the position of the corresponding image sensor, and A step of identifying a non-specular light source pixel adjacent to the light source pixel, wherein the non-specular light source pixel corresponds to a position in the scene where the light source is supplying light to the position in the scene corresponding to the observed pixel. The method according to claim 1, wherein the system of equations includes the color values ​​of both the specular light source pixels and the non-specular light source pixels.

5. The method according to claim 1, wherein the step of obtaining the geometry of the scene and the position of the image sensor includes using a structure from a motion algorithm for the two or more images.

6. The method according to claim 1, wherein the step of obtaining the geometry of the scene includes iteratively fitting one or more surfaces to one or more objects present in the two or more images.

7. The method according to claim 1, further comprising the step of generating an object texture for an object in the scene using the original color values ​​of the observed pixels corresponding to the object.

8. The method according to claim 7, further comprising the step of generating a reflectance map for the object based on the color values ​​of the observed pixels compared to the corresponding original color values.

9. The method according to claim 8, further comprising the step of transmitting the object texture and the reflectance map for virtual view compositing.

10. A method for synthesizing images from a target viewpoint, A step of receiving a first object texture, a reflectance value for the first object texture, and a second object texture, wherein the first object texture and the second object texture correspond to objects in the scene. The steps include receiving the geometry of the scene corresponding to the first object texture and the second object texture, The steps include: using the geometry of the scene to composite the image using the first object texture and the second object texture at the target viewpoint; A step of simulating reflection on the first object texture based on the reflectance value and the texture value of the second object texture, The steps include receiving a third object texture and adding the third object texture to the first object texture with simulated reflections to composite the image at the target viewpoint, A method of having.

11. The method according to claim 10, further comprising the step of receiving a reflectance map having a plurality of reflectance values ​​for a plurality of pixels of the first object texture.

12. The method according to claim 10, further comprising the step of simulating reflection on the first object texture based on the reflectance value and the texture value of the third object texture.

13. A computer program executed by a processor, which causes the processor to perform the method described in any one of claims 1 to 12.

14. A processor configured to execute the computer program described in claim 13.

15. A computer-readable data carrier for carrying a computer program as described in claim 13.