Image processing method for setting transparency and color values of pixels in a virtual image

By classifying virtual image pixels and setting transparency values, the ghosting problem caused by linear combination was solved, achieving high-quality virtual image display.

CN114631120BActive Publication Date: 2026-03-31KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies, when creating virtual images, lead to false transparency areas by linearly combining transparency and color values, introducing ghosting areas and reducing the quality of virtual images and VR/AR.

Method used

By receiving multiple reference view images, transparency information is determined, including components of foreground information, background information, and uncertain information. Based on these components, virtual image pixels are classified, and binary transparency and color values ​​are selected to avoid the occurrence of ghosting areas.

Benefits of technology

It effectively eliminates ghosting areas in virtual images, ensuring clear display of objects on different backgrounds and improving the quality of virtual images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631120B_ABST
    Figure CN114631120B_ABST
Patent Text Reader

Abstract

A method for selecting transparency settings and color values for pixels in a virtual image is provided. A virtual image can be formed by combining reference images taken at different angles to produce a virtual image of viewing an object at a new, un-captured angle. The method includes determining for each pixel of the virtual image what information it carries from the reference view images. The information of the pixel is used to define a pixel class, and the class is used to select what information will be displayed by the pixel and set the color of the pixel based on logical conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more specifically to the field of creating virtual images from a combination of reference view images. In particular, this invention relates to a method for setting the transparency and color values ​​of pixels in a virtual image. Background Technology

[0002] Virtual images of an object can be created by combining reference view images of the object that have been captured from multiple different viewpoints. Virtual images can be created based on combinations of reference view images, allowing the object to be viewed from any angle (including angles where no reference view image was captured).

[0003] Reference view images are often captured against a chroma key background so that the resulting virtual image of the target can be used on different graphical backgrounds (e.g., in virtual reality or augmented reality (VR / AR)). The virtual image formed by combining reference view images can include overlapping areas of any combination of background areas and / or object areas from each of the reference view images.

[0004] Existing image processing techniques process layered images by calculating the color of surface pixels from a linear combination of the color and transparency values ​​of layers in the image. US2009 / 102857 is an example of this technique, which calculates the color of surface pixels based on the transparency values ​​of surface pixels, the transparency of the pixels below, and the color of the pixels below.

[0005] Linear combination techniques, such as those in US2009 / 102857, introduce false transparency areas when used on virtual images. These false transparency areas appear near the edges of objects from each reference view image, meaning that a virtual image formed by overlapping and combining reference view images can contain many false transparency areas.

[0006] False transparency areas are often referred to as ghosting or shadow areas. These ghosting areas cause the graphic background to shine through the object, degrading the quality of the virtual image and consequently reducing the quality of VR / AR in which it can be used. VR / AR users may lose expressiveness in their situation and may be unable to perform their required tasks due to ghosting areas hindering their ability to see objects or read.

[0007] Therefore, there is a need for an image processing method that can set the transparency and color values ​​of pixels in a virtual image without introducing ghosting regions. Summary of the Invention

[0008] This invention is defined by the claims.

[0009] According to an example of an aspect of the present invention, a method is provided for setting corresponding color values ​​and transparency values ​​for a plurality of virtual image pixels in a virtual image, the method comprising:

[0010] Receive multiple reference view images of an opaque object, wherein each of the multiple reference view images is captured from a different viewpoint; and

[0011] Combining the plurality of reference view images to create a virtual image of the object, wherein the virtual image is located at a different viewpoint than any one of the plurality of reference view images, and wherein each pixel of the reference view image corresponds to a pixel of the virtual image through mapping, wherein creating the virtual image includes:

[0012] Determine transparency information, which includes at least foreground and background information components for each of the plurality of virtual image pixels, the components being derived from corresponding pixels of the plurality of reference view images;

[0013] Each virtual image pixel among the plurality of virtual image pixels is classified based on a combination of the components of foreground information and background information;

[0014] For regions far from the edges of the opaque object, a binary transparency value is selected for each virtual image pixel based on the classification of the virtual image pixels to display foreground or background information; and

[0015] The color value of each of the plurality of virtual image pixels is set based on the foreground information or background information selected to be displayed.

[0016] Each of the plurality of reference view images may include depth information. The mapping between reference view image pixels and virtual image pixels may be based on the depth information.

[0017] This method classifies each pixel in a virtual image formed from multiple reference images taken from different angles or viewpoints. Creating the virtual image by combining reference images means using information components from multiple reference images to construct each pixel of the virtual image. Each pixel of a reference view image is mapped to at least one virtual image pixel. Transparency information components mapped to each of the multiple virtual image pixels (e.g., foreground information from reference view image 1 and background information from reference view image 2) can be determined, thereby creating a classification for each pixel based on the combination of its foreground / background information components. The virtual image pixel classification is used to determine which component should be displayed, and subsequently, the color of the virtual image pixel should be set to a specific corresponding value. By selecting one of the transparency information components to be displayed, the invention ensures that the virtual image of an object does not contain background artifacts covering the object, or, if the object was captured on a chroma key (such as a green screen), does not cause ghosting when projected onto a background.

[0018] The transparency value indicates whether a pixel is in the foreground or the background.

[0019] Each pixel of the constructed virtual image includes, for example, an RGBA value, where the alpha channel (α) contains the desired background / foreground information. Reference images can also be encoded in the same format as RGBA pixels. They are warped to create the virtual image. Reference images typically also have an associated depth or disparity map (D) containing 3D information, such that camera shifts result in a warped image manipulated by local 3D information. Therefore, a reference image can contain five channels: RGBA, RGBA, and RGBA.

[0020] Each of the reference view pixels includes transparency information in the form of foreground information, background information, or uncertain information, and determining the transparency information further includes determining an uncertain component for each of the plurality of virtual image pixels.

[0021] Therefore, pixels in the reference images can be encoded as foreground, background, or indeterminate. A combination of these three states from multiple reference images determines the classification for each pixel, and subsequently, the transparency information.

[0022] The edge region of an object can contain a combination of background and foreground components that can contribute to the ghosting effect. This indeterminate component allows the ghosting region to be identified.

[0023] The classification of virtual image pixels (referred to as pixels in this document) may also include the classification of ghosting regions of pixels that include a combination of components including uncertain regions and foreground information.

[0024] This classification allows ghosting areas to be located and targeted for correction.

[0025] The method may include: for each of the virtual image pixels having transparency information that includes only foreground information, selecting a transparency value corresponding to the foreground.

[0026] Selecting foreground information only for display can contribute to removing ghosting effects because it helps remove the presence of background shading on the object surface introduced by ghosting areas. It also ensures that the background information component of the image does not overlay the foreground information component in pixels that include both foreground and background components.

[0027] A new outer boundary of the object in the virtual image can be created by selecting a transparency value corresponding to the foreground for all virtual image pixels that have transparency information that includes only the component and the combination of uncertain components that include only the component that includes background information or only the combination of multiple uncertain components.

[0028] This new outer boundary can connect overlapping boundaries from a reference view image to a single continuous outer edge.

[0029] This pixel classification selection selects pixels at or near the outer boundary of the overlapping reference view image of the object, and reduces any reduction in object size that may occur when defining a new outer boundary.

[0030] The color value setting in this method can be achieved by setting the corresponding color value of each pixel having transparency information that includes only foreground information components to the average color value of the corresponding foreground information components from the reference view image pixels corresponding to each pixel. A weighted combination can be used, where the weight is higher when the target virtual viewpoint is closer to a particular reference view.

[0031] This allows objects in a virtual image to have smooth color transitions across their surfaces, as the reference view image can show objects in different colors due to lighting, reflection, and / or objects with different colors on different surfaces.

[0032] This method can include selecting a binary transparency value for all virtual images. Therefore, the pixels of the entire virtual image can be categorized as foreground or background.

[0033] One limit for transparency is used for the background, and another limit is used for the foreground. For areas around the edges of an object, transparency values ​​between these limits can be set.

[0034] In this way, the color transition of the background at the edges of the object can be made less sharp.

[0035] The method includes, for example, setting the transparency value of each of a plurality of virtual image pixels having transparency information including an indeterminate component but excluding a foreground component, using the color difference between the pixel and at least one adjacent pixel.

[0036] The virtual image pixels that include the indeterminate component but not the foreground component are those that define the edges of the object. They are set to intermediate transparency values ​​instead of a binary foreground / background setting. The transparency is set based on the colors of adjacent virtual image pixels that do include foreground information.

[0037] This allows the edges of an object to be blended with the new background, thus enabling smooth transitions that do not appear sharp and unrealistic. Transparency can be calculated using the color difference between adjacent pixels, the color difference with a known (e.g., green) background color, a weighted combination of Euclidean distances to adjacent pixels, or other averaging methods known in the art, and any combination thereof.

[0038] The transparency information component of each of the multiple virtual image pixels can be derived from the color of the reference view image pixel corresponding to each of the multiple virtual image pixels. Each reference view image pixel, combined to form the virtual image pixel, contributes some color to the virtual image pixel. This color can be used to determine whether the pixel is part of the image background or part of the target object.

[0039] The background of the reference view image can be a chroma key background.

[0040] This simplifies identifying background and foreground components by color selection, since the background is a known and consistent color.

[0041] The method can be a computer program that includes computer program code modules, which, when run on a computer, are adapted to implement any example of the method of the present invention.

[0042] An image processing device is also provided, comprising:

[0043] An input unit is used to receive multiple reference view images of an object, wherein each of the multiple reference view images includes foreground information and background information captured from different viewpoints;

[0044] A processor for processing the plurality of reference view images to generate a virtual image; and

[0045] The output unit is used to output the virtual image.

[0046] The processor is adapted to implement the methods defined above.

[0047] These and other aspects of the invention will become apparent from the embodiments described below and will be set forth with reference to the embodiments described below. Attached Figure Description

[0048] To better understand the invention and to more clearly illustrate how it can be implemented, reference will now be made to the accompanying drawings by way of example only, wherein:

[0049] Figure 1 A shows an example setup for capturing a reference view image of a semi-transparent window through which the background can be seen;

[0050] Figure 1 B shows an example setup for capturing a reference view image of an object in front of a chroma key background;

[0051] Figure 2 An example of a top view from two reference cameras is shown for a reference view image of an object captured against a chroma key background.

[0052] Figure 3 An example of a virtual image formed by two reference view images having a classification used in the method of the present invention is shown;

[0053] Figure 4 A flowchart illustrating the steps of an example of the present invention is shown;

[0054] Figure 5 The illustration shows an example according to the invention after color and transparency values ​​have been set. Figure 3 Virtual images;

[0055] Figure 6 An image processing device is shown. Detailed Implementation

[0056] The invention will be described with reference to the accompanying drawings.

[0057] It should be understood that the detailed descriptions and specific examples are intended for illustrative purposes only when indicating exemplary embodiments of the apparatuses, systems, and methods, and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatuses, systems, and methods of the present invention will become better understood from the following description, the appended claims, and the accompanying drawings. It should be understood that the drawings are merely schematic and not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar parts.

[0058] This invention provides a method for setting the transparency and color values ​​of pixels in a virtual image. A virtual image can be formed by combining reference images captured at different angles to produce a virtual image of an object viewed from a new, uncaptured angle. The method includes determining, for each pixel of the virtual image, what information it carries from the reference view image. The pixel-specific information is used to define pixel categories, and these categories are used to select, based on logical conditions, what information will be displayed by the pixel and to set the pixel's color and transparency values. This invention particularly relates to the blending of a virtual image of an object with a new background, which is assumed to always be farther away than the object.

[0059] Before describing the invention in detail, issues related to image transparency and depth will be described first, as well as conventional methods for creating new images (referred to as "virtual images") from different viewpoints.

[0060] Figure 1 A illustrates an example of transparency in multi-camera capture. The captured scene includes objects with foreground 100 and background 102, which have a semi-transparent window 104 through which the background 102 is visible. The entire scene is captured by a set of cameras 106, each of which generates a reference image. The images are processed to generate new images from a virtual camera 108 at a different viewpoint than each of the cameras in the set of cameras 106. This new image is called a virtual image. This means that it is an image captured from a virtual camera location (i.e., a location where a real camera image is not available).

[0061] For the semi-transparent window 104, each camera 106 will see a different blend of foreground 100 and background 102. A standard method for synthesizing a new virtual image from the viewpoint of the virtual camera 108 is to weight distorted versions of two or more reference views based on pose proximity, deocclusion (stretching), and possible depth. This naive approach can yield sufficient image quality when the pose of the view to be synthesized is close to that of the reference views.

[0062] Figure 1 B illustrates an example of the correlation between transparency and chroma keying. In this case, the background is set to a screen of a known color, such as green. This allows for easy extraction of the foreground from the background in image processing.

[0063] This is a special case of transparency, where the object to be captured is opaque but is considered completely transparent outside its outer boundary (to show the background screen). Of course, the object does not need to be a closed solid shape, and it can have an opening through which the screen is visible.

[0064] The problem here is more serious because the green screen background 102 needs to be replaced by a new background to create a virtual image. For example, the foreground is overlaid on a separate background to create the desired overall image. If transparency is ignored, the blended pixels that still contain the green component will remain at the object boundaries. This is very noticeable when using the Naive View blending method that ignores transparency, or when using the Linear Combination method.

[0065] The standard method for chroma keying is to calculate a transparency map. Using the standard alpha-matting equation, the image color of a pixel can be described as a linear combination of the foreground and background colors:

[0066] i = αf + (1-α)b

[0067] In this equation, i represents the determined pixel color, f represents the foreground color, b represents the background color, and α is the transparency value. This equation is applied to the entire image on a per-pixel basis, thus creating a transparency map.

[0068] Considering the color characteristics of a specific background used for chroma keying (e.g., green), one of the many algorithms available in the literature can be used to estimate the foreground color f and transparency value α for a pixel. The color data can be used to determine the transparency of each reference image pixel by comparing it to a known background color. If the pixel color equals the chroma key color, then the pixel is transparent. If the pixel contains only a fraction of the chroma key color, then it is semi-transparent. If it contains no fraction of the chroma key color, then it is not transparent.

[0069] Different algorithms exist for this purpose. They typically combine processing of the local neighborhood around the pixel with known color characteristics of the background material. Even without a green screen, the transparency around the depth step detected in the depth map can still be estimated. In such cases, blended pixels often exist around the edges of the defocused foreground.

[0070] The value of α varies between 0 and 1. In this document, the selection symbol indicates that 1 represents zero transparency, so the foreground color is visible, and 0 represents full transparency, so the background color is visible.

[0071] Therefore, the transparent area of ​​an image (α=0) is the area where the background is visible, while the opaque area of ​​an image (α=1) is the area where the foreground is visible.

[0072] Typically, these variables are estimated using small neighborhoods of pixels. When applied to reference view images of k = 1…N and taking into account the presence of perspective information generated by a depth sensor or depth estimation process, the following data is obtained:

[0073] f1,d1,α1,b1

[0074]

[0075] f N ,d N ,α N ,b N

[0076] This data is used to reference pixels in each of the reference images 1…N, which correspond to pixels in the virtual image. Multiple sets of this data are used across the virtual image as a whole to determine the values ​​of all its pixels.

[0077] For all raw pixels on the green screen, α = 0 because only the chroma key color is seen. For pixels on the foreground object, α = 1 because no chroma key color is seen. For pixels on the boundary or in transparent areas, 0 ≤ α ≤ 1 because the pixel color fraction is the chroma key color.

[0078] A transparency map (i.e., a set of α values ​​for the entire image) can be used to blend a single reference view with a new background. However, known rendering methods that use multiple reference views cannot handle transparency.

[0079] Currently, multiple reference views are distorted to the synthetic viewpoint, after which multiple predictions are blended to predict a single new synthetic view (virtual image). The following equation calculates the color of a particular virtual image pixel based on the reference view image pixels from each reference view image mapped to a single virtual image pixel:

[0080]

[0081] The tilde on i This refers to the distortion of the reference view images before they are combined to form a weighted prediction. In this equation, the w value is the weighting factor assigned to each reference view image or each pixel within a reference view image. (The text then abruptly shifts to a different topic: "with subscripts...") The value refers to the color value of a pixel in a specific reference view image, which is given as an index number. (No index) It is the calculated color value of the virtual image pixel corresponding to all the reference view image pixels.

[0082] The reference view image is distorted before combination to determine which reference view image pixels correspond to virtual image pixels, as the mapping between the two will vary between the reference image and the desired viewpoint. This equation must be applied to each virtual image pixel to determine its color based on its corresponding reference image pixel. The above equation only ignores the presence of α.

[0083] A straightforward approach would be to distort and weight α in a manner similar to how colors are averaged. Again, this process calculates the value for each individual pixel based on the properties of the reference image pixels mapped to that particular virtual image pixel, and must be performed for each virtual image pixel. Again, the w value is a weighting factor assigned to each reference view image or each pixel within a reference view image. An α value with a subscript refers to the transparency value of a pixel in a specific reference view image, where the reference view image is identified by its subscript number. An α without a subscript is the calculated transparency value for each corresponding virtual image pixel in the reference image pixel:

[0084]

[0085] The resulting virtual image pixel color and transparency It can now be used to composite captured multi-view foreground objects onto a new (graphic) background.

[0086] However, this method of handling transparency results in artifacts around the boundaries of the foreground object.

[0087] Figure 2 Two reference cameras 106, C1, and C2 are shown for capturing the foreground object 202 in front of the green screen 102.

[0088] A naive weighted combination of the transparency values ​​of the reference view results in ghosting regions 200, where the background "illuminates" through the foreground object. These semi-transparent ghosting regions 200 appear near the boundaries of the foreground object 202 generated by this method.

[0089] Figure 2 Different regions along the green screen 102 are shown. Two reference cameras 106 will view different parts of the scene when pointed at those regions. The first regions are designated B1 and B2. This means that when imaging in this direction, the first camera C1 observes background information B1, and the second camera C2 observes background information B2. The last regions are designated F1 and F2. This means that when imaging in this direction, the first camera C1 observes foreground information F1, and the second camera C2 observes foreground information F2. There are conflicts between the reference images in other regions. U refers to an undefined region near the boundary.

[0090] The 3D shape of an object can cause ghosting problems. However, depth estimation and / or filtering errors can also contribute to this issue.

[0091] Figure 3Two elliptical images 300 and 302 of a foreground object, distorted from two reference cameras positioned differently for imaging the scene, are shown.

[0092] like Figure 2 In the text, the two-letter combination indicates whether, for each distorted reference view, a pixel is a foreground pixel (F), a background pixel (B), or an indeterminate pixel (U) characterized by shading from both the foreground and background. The subscript indicates which reference view image the pixel originates from, and thus which camera the pixel originated from.

[0093] The reference view image of object 202 captured by camera C1 is ellipse 300. This has been combined with the reference view image of object 202 (ellipse 302) captured by camera C2. The area where the edge of one of the reference view image ellipses in the virtual image overlaps with the interior of the other reference view ellipse produces a ghosting region 200.

[0094] The method of the present invention will now be described.

[0095] Figure 4 A method 400 for setting corresponding color and transparency values ​​for multiple virtual image pixels is shown.

[0096] In step 402, multiple reference view images of the object are received. Each of the multiple reference view images includes foreground and background information captured from different viewpoints.

[0097] Then, in step 403, the reference view images are mapped such that there is a mapping between pixels of each reference image and a single virtual image. This mapping depends on the viewpoint of each reference image and involves, for example, distorting and combining images. A correspondence then exists between pixels in the reference images and pixels in the virtual images. Specifically, the content of each pixel in the virtual image is determined based on the content of its “corresponding” (i.e., linked via mapping) pixel from the reference images. Typically, this is a many-to-one mapping, but this is not substantial. For example, due to occlusion, more than one input pixel will be mapped to the same output.

[0098] Then, step set 404 is executed to combine multiple reference view images, thereby creating a virtual image of the object. The virtual image is located at a different viewpoint than any of the multiple reference view images.

[0099] In step 406, transparency information is determined, which includes at least foreground and background information components for each of the plurality of virtual image pixels. These components are derived from corresponding pixels in the plurality of reference view images.

[0100] For example, Figure 3A region of the image is designated as B1 and F2. B1 is the component of background information already obtained from the reference image camera C1, and F2 is the component of foreground information already obtained from the reference image camera C2.

[0101] For a pixel in this region, sets B1 and F2 constitute the "foreground information and background information components".

[0102] In step 408, each of the multiple virtual image pixels is classified based on a combination of foreground and background information components.

[0103] Classification is used to determine which component should be displayed. Specifically, for each pixel in the virtual image, a component of foreground / background information derived from the corresponding pixel in multiple reference view images (i.e., based on the mapping used to derive the virtual image from the reference images) is determined. Pixels are classified according to the components of the foreground / background information pixels. Transparency values ​​can then be derived. In the simplest implementation, transparency values ​​may have only a binary component corresponding to either "foreground" or "background." As is apparent from the description above, transparency data can be derived from color data. Transparency values ​​determine whether foreground or background scene data has been recorded. Essentially, the green screen color is measured, and the transparency value depends on how close the observed color is to the green screen color. More complex methods exist.

[0104] This invention utilizes the non-linear dependence of the foreground / background / uncertainty information (F, B, and U) components on transparency. This is, for example, by classifying each pixel of a reference image through one of the three distinct components:

[0105]

[0106] Here, F explicitly refers to the foreground, B explicitly refers to the background, and U refers to "indeterminate" pixels. Δ is the threshold transparency value, to which the pixel transparency α is compared. The threshold transparency can be set according to the object being imaged and can be 0.5 or less.

[0107] Alternative locations can be used to classify pixels based on irrelevant ranges, for example:

[0108]

[0109] In chroma keying examples, these indeterminate pixels are pixels with blended color values, where the background color is still visible in the output color. In some examples, more classifications can be used, especially for objects with different levels of transparency. In images with multiple foreground objects, or for objects with different transparency, multiple foreground classifications can exist corresponding to each object or a region of each object.

[0110] The invention can also be implemented using only the F and B components, since the threshold of α can be adjusted to ensure that all possible values ​​are covered only by F and B.

[0111] In step 410, foreground or background information is selected to be displayed by each of the plurality of virtual image pixels based on its category. This involves selecting a transparency value from one of two binary values. The virtual image pixels are thus selected to display only foreground or background information in a binary manner to avoid ghosting issues. This will be discussed further below.

[0112] In step 412, the color value of each of the multiple virtual image pixels is set based on the information selected to be displayed.

[0113] Return to reference Figure 3 The ghosting region comes from the following combination of components:

[0114] C1,C2={(F1,U2),(U1,F2)}.

[0115] This set of values ​​represents a classification of the properties of that part of the virtual image. This classification for a specific pixel is used to select which component will be displayed by the pixel by encoding appropriate transparency values, and is used to set the corresponding color.

[0116] Below is an example of an algorithm that selects the component to be displayed and sets the color based on the transparency information of a reference view image:

[0117] If (C1 == F)

[0118] If (C2 == F)

[0119] Create a weight for the reference color and set the output opacity to 1.

[0120] (i.e., transparency setting = foreground)

[0121] Otherwise, select F with transparency of 1 (i.e., transparency setting = foreground).

[0122] otherwise

[0123] If (C2 == F)

[0124] Select F with transparency of 1 (i.e., transparency setting = foreground).

[0125] otherwise

[0126] Choose any color with 0% transparency (i.e., transparency setting = background).

[0127] In the following cases, this example algorithm will always return an opacity of 0:

[0128] C1,C2=U1,U2.

[0129] In this scenario, accurate transparency estimation is unnecessary as long as the classification is performed correctly. The algorithm above simply returns transparency values ​​as 0 or 1 (and thus classifies pixels as background or foreground), thus generating no intermediate values.

[0130] Figure 5 This demonstrates how to apply the above algorithm. Figure 3 This is an example of a virtual image created from a virtual image. The combined projection of the foreground of all reference views results in a new outer boundary, which is a combination of elliptical images 300 and 302. This new outer boundary can be identified by the following category combinations:

[0131] C1,C2={(B1,U2),(U1,B2),(U1,U2)}.

[0132] These are virtual pixels that have transparency information including indeterminate components but excluding foreground components.

[0133] The example algorithm above, for instance, selects any desired color and makes the new outer boundary pixels completely transparent to blend with the new background. However, the effect of doing so is that the foreground object is slightly reduced in size because pixels with only uncertain information (which appear at points around the outer edge of the object) will become transparent.

[0134] Another issue is that objects may exhibit sharp color transitions to the background. This requires knowing precisely the foreground color f1…f2 for each virtual image pixel 1…N. N and the corresponding transparency values ​​α1…α N In certain situations, these parameters can be used directly to blend edges with a new background. However, in practice, it is very difficult to estimate these parameters accurately.

[0135] Therefore, examples of the present invention can use intermediate transparency values ​​for this boundary region. Thus, for the region at the edge of an object, transparency values ​​can be set between extreme values. The transparency value of each pixel in this edge region can be set using the color of at least one adjacent pixel.

[0136] In this way, a new method for defining the outer boundaries of blended objects can be implemented. For each pixel of the virtual image, the method uses pure foreground pixels from the spatial neighborhood of the pixel in question—next to, around, or nearby—to form the foreground color. The estimation. These foreground pixels are defined based on the set of pixels from the foreground in any of the distorted reference views:

[0137] C1,C2={(F1,B2),(B1,F2),(F1,F2)}.

[0138] When both foreground elements are available for one or more pixels used in the averaging, the method can select a color from one of the reference view images, or use the average color value of both. The resulting transparency can be calculated based on the color difference between the current (uncertain) pixel and its neighboring pixels and / or as a weighted combination of Euclidean distances to neighboring pixels and / or using other averaging or weighting methods known in the art.

[0139] The closer to the object's boundary, the more transparent the output pixels should be, because as you approach the object's edge, the increase in pixel color fraction will be the chroma key color. This will smoothly blend the object's new outer boundary with the background.

[0140] This invention is generalized to use more than two reference views as used in the examples. Again, combinations of pixel classes derived from multiple reference views can be used to determine the settings for output color (selection or blending of multiple colors) and output transparency.

[0141] The above method, which uses binary selection of transparency based on pixel classification, is a non-linear approach to processing reference image pixels. It has already been described above in conjunction with opaque objects (although optionally with fully transparent areas).

[0142] If the object has semi-transparent areas, the above method can be modified. For example, the least opaque reference image can be dominant, and its transparency value is used as the output transparency value for the virtual image pixels. This, for example, makes it possible to process images with hazy, foggy, or dirty windows, or images of hair. Instead of selecting the minimum transparency, a non-linear selection can be made.

[0143] The following presents a GPU implementation of the process of combining reference images.

[0144]

[0145] The code receives two vector components (the x and y positions of the pixels) from the reference view images and outputs four vectors (three color components and one transparency component). For the two reference view image pixels t1 and t2, their opacity "a" is compared with the opacity standard "opaque = 0.9f". If both reference image pixels meet or exceed this opacity threshold, the color of the virtual image pixel t3 is set to a blend of t1 and t2, and the opacity of the virtual image pixel is set to 1.0.

[0146] If neither reference image pixel is opaque according to the opacity criterion, their opacities are compared. The virtual image pixel t3 is set to match the color and opacity of the least opaque reference image pixel. In this case, the virtual image pixel can have an opacity value that is neither 0 nor 1, but rather an intermediate value.

[0147] Figure 6 An image processing device 600 is shown. The input unit 602 receives multiple reference view images of an object, such as those captured by the set of cameras 106.

[0148] Processor 604 processes multiple reference view images to generate a virtual image. This is provided at output unit 606 to, for example, display 608. The processor implements the above method.

[0149] In alternative examples, the algorithm may include one or more intermediate transparency values ​​that pixels can be set to. Furthermore, the algorithm may set the transparency of pixels to values ​​other than 1 when C1,C2 = U1,U2, for example, when creating a virtual image of an object with varying degrees of transparency from each viewpoint, or for objects requiring artificial transparency.

[0150] By studying the accompanying drawings, disclosure, and claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude multiple. A single processor or other unit can implement the functions of several items recited in the claims. Although specific measures are recited in dissimilar dependent claims, this does not indicate that combinations of these measures cannot be advantageously used. If a computer program has been discussed above, it may be stored / distributed on a suitable medium, such as an optical storage medium or solid-state medium provided together with or as part of other hardware, but it may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. If the term "suitable" is used in the claims or description, it should be noted that the term "suitable" is intended to be equivalent to the term "configured as." No reference numerals in the claims should be construed as limiting the scope.

[0151] Typically, examples of image processing devices, image processing methods, and computer programs implementing the methods are indicated by the following embodiments.

[0152] Example:

[0153] 1. A method (400) for setting corresponding color values ​​and transparency values ​​for multiple virtual image pixels in a virtual image, the method comprising:

[0154] (402) Receive multiple reference view images of an opaque object, wherein each of the multiple reference view images includes depth information captured from different viewpoints; and

[0155] (404) Combining the plurality of reference view images to create a virtual image of the object, wherein the virtual image is located at a different viewpoint from any one of the plurality of reference view images, and wherein each reference view image pixel corresponds to a virtual image pixel by mapping, wherein creating the virtual image includes:

[0156] (406) Determine transparency information, the transparency information including at least foreground and background information components for each of the plurality of virtual image pixels, the components being derived from the corresponding pixels of the plurality of reference view images;

[0157] (408) Classify each virtual image pixel among the plurality of virtual image pixels based on the combination of the components of foreground information and background information;

[0158] (410) For regions far from the edges of the opaque object, a binary transparency value is selected for each virtual image pixel based on the classification of the virtual image pixels; and

[0159] (412) Set the color value of each of the plurality of virtual image pixels based on the information selected to be displayed.

[0160] 2. The method according to embodiment 1, wherein each of the reference view pixels includes depth information in the form of foreground information, background information, or uncertain information, and wherein determining the transparency information further includes determining an uncertain component for each of the plurality of virtual image pixels.

[0161] 3. The method according to embodiment 2 includes: classifying each of the plurality of virtual image pixels having transparency information as a combination of uncertain components and foreground information components into a ghosting region responsible for the ghosting effect.

[0162] 4. The method according to embodiment 2 or 3 includes:

[0163] For each of the plurality of virtual image pixels having transparency information that includes only foreground information, a first binary transparency value corresponding to non-transparency is selected; and

[0164] For each of the plurality of virtual image pixels having transparency information that includes only background information, a second binary transparency value corresponding to transparency is selected.

[0165] 5. The method according to any one of embodiments 2 to 4, comprising creating a new outer boundary of the object in the virtual image by selecting a first binary transparency value corresponding to non-transparency for a virtual image pixel having transparency information that includes only a combination of components and uncertain components that provide background information or only a combination of multiple uncertain components.

[0166] 6. The method according to any one of embodiments 1 to 5, comprising: setting the color of each of the plurality of virtual image pixels having transparency information including only foreground information as the average color of the foreground component of a reference view image.

[0167] 7. The method according to any one of embodiments 1 to 6, comprising selecting a binary transparency value for all pixels of the virtual image.

[0168] 8. The method according to any one of embodiments 1 to 6, comprising: setting a transparency value between the binary values ​​for a region at the edge of the object.

[0169] 9. The method according to embodiment 8, comprising: setting the transparency value of each of the plurality of virtual image pixels having transparency information including an indeterminate component but excluding a foreground component, using the color of at least one adjacent pixel.

[0170] 10. The method according to any one of embodiments 1 to 9, comprising: for each of the plurality of virtual image pixels, using the color of the corresponding reference view image pixel to determine transparency information derived from the reference view image.

[0171] 11. The method according to any of the foregoing embodiments, wherein the background is a chroma key.

[0172] 12. A computer program including a computer program code module, wherein when the program is run on a computer, the computer program code module is adapted to implement the method according to any one of embodiments 1 to 11.

[0173] 13. An image processing apparatus (600), comprising:

[0174] The input unit (602) is used to receive multiple reference view images of an object, wherein each of the multiple reference view images includes foreground information and background information captured from different viewpoints;

[0175] Processor (604) for processing the plurality of reference view images to generate a virtual image; and

[0176] Output unit (606), which is used to output the virtual image,

[0177] The processor is adapted to implement the method according to any one of embodiments 1 to 12.

[0178] More specifically, the invention is defined by the appended claims.

Claims

1. A method (400) for setting respective color values and transparency values for a plurality of virtual image pixels in a virtual image, the method comprising: (402) receiving a plurality of reference view images of an opaque object, wherein each of the plurality of reference view images is captured from a different viewpoint; and (404) combining the plurality of reference view images to create a virtual image of the object, wherein the virtual image is at a different viewpoint than any of the plurality of reference view images, and wherein each reference view image pixel is mapped to correspond to a virtual image pixel, wherein creating the virtual image comprises: (406) determining transparency information for each of a plurality of reference view image pixels, wherein each of the reference view pixels includes transparency information in the form of foreground information (F), background information (B), or uncertain information (U); (408) classifying each of the plurality of virtual image pixels based on the transparency information of the corresponding reference view pixel, wherein a ghost region is given virtual image pixels having corresponding reference view pixels that include transparency information in the form of both uncertain information (U) and foreground information (F); (410) deriving transparency information for each of the virtual image pixels from corresponding pixels of the plurality of reference view images, including selecting a binary transparency value for each of the virtual image pixels for a region away from an edge of the opaque object based on the classification of the virtual image pixel to display foreground information or background information; and (412) setting the color value of each of the plurality of virtual image pixels based on the foreground information or the background information selected to be displayed.

2. The method of claim 1, including selecting a binary transparency value for each of the virtual image pixels in the ghost region.

3. The method of claim 1 or 2, including: selecting a first binary transparency value corresponding to non-transparent for each of the plurality of virtual image pixels having corresponding transparency information including only foreground information; and selecting a second binary transparency value corresponding to transparent for each of the plurality of virtual image pixels having corresponding transparency information including only background information.

4. The method of claim 1 or 2, including creating a new outer boundary of the object in the virtual image by selecting a first binary transparency value corresponding to transparent for virtual image pixels having corresponding transparency information including both background information and uncertain information or only uncertain information.

5. The method of claim 1 or 2, comprising: using a color of at least one neighboring pixel to set the transparency value of each of the plurality of virtual image pixels having corresponding transparency information including uncertain information and not foreground information.

6. The method of claim 1 or 2, comprising: for each of the plurality of virtual image pixels having corresponding transparency information comprising only foreground information, setting the color of the virtual image pixel to the average color of the reference view image pixels.

7. The method of claim 1 or 2, comprising selecting a binary transparency value for all pixels of the virtual image.

8. The method of claim 1 or 2, comprising: for regions at the edge of the object, setting transparency values between the binary transparency values.

9. The method of claim 1 or 2, comprising using the color of the reference view image pixels to determine the transparency information for each of the reference view image pixels.

10. The method of claim 1 or 2, wherein, The background is chroma key.

11. The method of claim 1 or 2, wherein, In the step of determining transparency information for a reference view pixel: determining foreground information (F) by comparing a transparency value of the reference view pixel with a first threshold; and / or determining background information (B) by comparing the transparency value of the reference view pixel with a second threshold; and determining indeterminate information when the transparency value of the reference view pixel of the respective reference view neither meets the first threshold nor the second threshold.

12. A computer program comprising computer program code adapted to perform the method of any of claims 1 to 11 when the program is run on a computer.

13. An image processing device (600) comprising: an input (602) for receiving a plurality of reference view images of an object, wherein each of the plurality of reference view images comprises transparency information and is captured from a different viewpoint; a processor (604) for processing the plurality of reference view images to generate a virtual image; and an output (606) for outputting the virtual image, wherein the processor is adapted to perform the method of any of claims 1 to 11.

Citation Information

Patent Citations

  • Antialiasing of two-dimensional vector images

    US20090102857A1

  • Method and apparatus for generating a three dimensional image

    CN106664397A

  • Image generation device and image generation method

    CN110291564A