Improvement of Texture and Alpha Channel in Multi-Plane Images
A multi-step image processing technique for multi-plane images addresses artifacts by normalizing weights, applying local averaging, and scaling texture values, enhancing image quality and reducing rendering errors.
Patent Information
- Application Number
- JP2024575433
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-26
- Publication Date
- 2025-07-23
AI Technical Summary
Existing multi-plane image rendering technologies suffer from artifacts such as 'dark pixels', 'black holes', and object boundary issues, which degrade the quality of rendered images.
A multi-step image processing technique involving alpha-channel normalization, local averaging, and texture-channel scaling is applied to multi-plane images to correct these artifacts, ensuring that weight sums are normalized, alpha and texture values are adjusted to local averages, and texture values match the source image at the reference camera position.
The technique significantly reduces or eliminates artifacts in rendered images, improving image quality and fidelity.
Smart Images

Figure 2025523504000001_ABST
Abstract
Description
Technical Field
[0001] 1. Cross - reference to related applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 357,669, filed on July 1, 2022, and European Patent Application No. 22182507.8, filed on July 1, 2022, each of which is hereby incorporated by reference in its entirety.
[0002] 2. Field of disclosure Various exemplary embodiments generally relate to multiplane imaging (MPI), and more specifically, but not limited to, relate to the editing of multiplane images.
Background Art
[0003] 3. Background Multiplane images embody a relatively new approach for storing volumetric content. MPI can be used to render both still images and videos. For example, it uses 32 texture planes and transparency (alpha) information per camera to represent a three - dimensional (3D) scene within a frustum. Exemplary applications of MPI include computer vision and graphics, image editing, photo - animation, robotics, and virtual reality.
Summary of the Invention
Means for Solving the Problems
[0004] This specification discloses various embodiments of an image processing technique directed to improving the quality of a visible image generated by rendering a multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position. In one exemplary embodiment, the image processing technique includes the following operations: (A) for a first set of pixels, scaling each weight of the layers such that the sum of the scaled weights is normalized to 1; (B) for a second set of pixels, replacing each alpha value and texture value in the layers with corresponding local average values; and (C) for a third set of pixels, scaling the corresponding texture values in the layers such that the texture values of the third set match the respective texture values of a source image captured from the reference camera position for the resulting visible image rendered for the reference camera position, one or more of which are included.
[0005] According to one exemplary embodiment, an apparatus for improving a first multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position is provided. The apparatus includes at least one processor and at least one memory including program code. The at least one memory and the program code, together with the at least one processor, cause the apparatus to at least: scale each weight of the layers such that the sum of the scaled weights equals a predetermined fixed value for each pixel of the first set of pixels; replace each alpha value and texture value in the layers with corresponding local average values for each pixel of the second set of pixels; and scale the corresponding texture values in the layers such that the texture value of each pixel of the third set matches the respective texture value of a reference image captured from the reference camera position for the resulting visible image rendered for the reference camera position for each pixel of the third set of pixels.
[0006] According to another exemplary embodiment, a method for improving a first multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position is provided. The method includes: for each pixel of a first set of pixels, scaling each weight of the layers such that the sum of the scaled weights becomes a predetermined fixed value, wherein the scaling of each weight is performed using at least one processor and at least one memory including program code; for each pixel of a second set of pixels, replacing each alpha value and texture value in the layer with a corresponding local average value, wherein the replacement is performed using the at least one processor and the at least one memory; and for each pixel of a third set of pixels, scaling the corresponding texture value in the layer such that the texture value of each pixel in the resulting visible image rendered for the reference camera position matches the respective texture value of a reference image captured from the reference camera position, wherein the scaling of the corresponding texture value is performed using the at least one processor and the at least one memory.
[0007] According to yet another exemplary embodiment, there is provided a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations, the operations including a method for enhancing a first multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position, the method comprising: for each pixel of a first set of pixels, scaling each weight of the layers such that the sum of the scaled weights becomes a predetermined fixed value, the scaling of each weight being performed using at least one processor and at least one memory including program code; for each pixel of a second set of pixels, replacing each alpha value and texture value in the layer with a corresponding local average value, the replacement being performed using the at least one processor and the at least one memory; and for each pixel of a third set of pixels, scaling the corresponding texture value in the layer such that the texture value of each pixel of the third set matches the respective texture value of a reference image captured from the reference camera position for a resulting visible image rendered for the reference camera position, the scaling of the corresponding texture value being performed using the at least one processor and the at least one memory.
Brief Description of the Drawings
[0008] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent from the following detailed description and the accompanying drawings by way of example.
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3
[0012]
Figure 4
[0013]
Figure 5A
Figure 5B
Figure 5C
Figure 5D
[0014]
Figure 6
[0015]
Figure 7
[0016]
Figure 8
[0017]
Figure 9
[0018]
Figure 10
[0019]
Figure 11
[0020]
Figure 12
[0021] The present disclosure and aspects thereof can be embodied in various forms including hardware, devices or circuits controlled by a computer-implemented method, computer program product, computer system and network, user interface, and application programming interface; and hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The above is only intended to give a general idea of the various aspects of the present disclosure and is in no way intended to limit the scope of the present disclosure.
[0022] In the following description, numerous details of device configuration, timing, operation, etc. are set forth in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to those skilled in the art that these specific details are merely exemplary and are not intended to limit the scope of the present application.
[0023] Furthermore, while the present disclosure focuses primarily on examples where various circuits are used in a digital projection system, it will be understood that these are merely examples. It should be further understood that the disclosed systems and methods can be used in any device that needs to project light, such as in movie theaters, consumer, and other commercial projection systems, head-up displays, virtual reality displays, and the like.
[0024] Video Coding According to Exemplary Embodiments FIG. 1 shows an exemplary process of a video delivery pipeline (100) showing various stages from video / image capture to video / image content display, according to an embodiment. A sequence of video / image frames (102) can be captured or generated using an image generation block (105). The frames (102) can be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and / or image data (107). Alternatively, the frames (102) may be captured on film by a film camera. The film can then be converted to a digital format to provide the video / image data (107).
[0025] In the production phase (110), data (107) can be edited to provide a video / image production stream (112). The data of the video / image production stream (112) may be provided to a post-production block (115) for post-production editing by a processor (or one or more processors, such as a central processing unit (CPU)). Post-production editing in block (115) may include, for example, adjusting or modifying the color or luminance in a specific region of an image according to the creative intent of the video producer to improve image quality or achieve a specific appearance for the image. This part of post-production editing is sometimes referred to as "color timing" or "color grading". Other edits (such as scene selection and sequencing, image cropping, addition of computer-generated visual special effects, removal of artifacts, etc.) can be performed in block (115) to result in a "final" version (117) of the production for distribution. The improvement of texture and alpha channels in the multi-plane images disclosed hereinafter in this specification may be performed in block (115). During post-production editing (115), the video and / or image can be viewed on a reference display (125).
[0026] Following post-production (115), the data of the final version (117) may be delivered to an encoding block (120) for further downstream delivery to a decoding and playback device such as a television set, a set-top box, a movie theater, etc. In some embodiments, the encoding block (120) can include audio and video encoders such as those defined by ATSC, DVB, DVD, Blu-Ray, and other delivery formats to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by a decode unit (130) to generate a corresponding decoded signal (132) that represents a copy or a close approximation of the signal (117). The receiver may be attached to a target display (140) that may have somewhat or completely different characteristics than the reference display (125). In such a case, a display management (DM) block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Depending on the embodiment, the decode unit (130) and the display management block (135) may include individual processors or may be based on a single integrated processing unit.
[0027] Multiplane Imaging A multi-plane image includes a plurality of image planes, and each of the image planes is a "snapshot" of a 3D scene at a certain depth relative to the camera position. The information stored in each plane includes texture information (e.g., represented by R, G, B values) and transparency information (e.g., represented by an alpha (A) value). Here, the acronyms R, G, B represent red, green, and blue, respectively. There are various ways to generate a multi-plane image. For example, two or more input images from two or more cameras located at different known viewpoints can be processed together to generate a corresponding multi-plane image. Alternatively, single-view synthesis of a multi-plane image may be performed using a source image captured by a single camera. For at least some multi-plane image generation algorithms, unfortunately, the corresponding multi-plane image may exhibit one or more types of artifacts when rendered on a reference display (125) or a target display (140). The various embodiments disclosed herein can be beneficially used to reduce the occurrence of such artifacts and / or to substantially completely suppress such artifacts.
[0028] FIG. 2 shows a 3D scene representation using a multi-plane image (200) according to an embodiment. The multi-plane image (200) has D planes or layers (P0, P1, …, P(D−1)). Here, D is an integer greater than 1. The layers are indexed such that the layer farthest from the reference camera position (RCP) is the 0th layer and is at a distance (or depth) d0 from the RCP along the Z dimension of the 3D scene. The index is incremented by 1 for each successive layer closer to the RCP. The plane (layer) closest to the RCP has the index value (D−1) and is at a distance (or depth) d from the RCP along the Z dimension D-1It is located at. The planes (P0, P1, …, P(D-1)) are orthogonal to the base plane (202) parallel to the XZ coordinate plane. The RCP is at a vertical height h above the base plane (202). The XYZ triad shown in FIG. 2 indicates the general orientation of the multi-plane image (200) and the planes (P0, P1, …, P(D-1)) with respect to the X, Y, and Z dimensions of the 3D scene.
[0029] Let the RGB value for the i-th layer be C i and denote it. The horizontal size of the layer is H×W, where H is the height (Y dimension) of the layer and W is the width (X dimension) of the layer. The pixel value (x, y) for color channel c is C i denoted as (x, y, c). The α value for the i-th layer is A i denoted as, and the corresponding pixel value (x, y) for the alpha channel is A i denoted as (x, y). The depth distance from the i-th layer to the reference camera position (RCP) is d i denoted as. The source image from the original reference view (by the camera fixed at the RCP) is denoted as R. Here, the texture pixel value is denoted as R(x, y, c). In MPI, the effective distance between two adjacent layers typically has a fixed value. For example, the different layers (P0, P1, …, P(D-1)) of the multi-plane image (200) are spaced equidistantly in terms of disparity (the reciprocal of depth).
[0030] As already described above, a multi-plane image such as the multi-plane image (200) can be generated using single-view synthesis from a single source image R or using multi-view synthesis from two or more source images. Such synthesis may be performed, for example, during the production phase (110). The corresponding MPI synthesis algorithm may typically output a multi-plane image (200) containing XYZ-resolved pixel values of the form {(C i , A i ) i = 0, …, D-1}.
[0031] {(Ci , A i ) By processing the multi-plane image (200) represented by i = 0, …, D - 1}, the MPI rendering algorithm can generate a visible image corresponding to the RCP or a new virtual camera position different from the RCP. An exemplary MPI rendering algorithm (often referred to as an "MPI viewer") that can be used for this purpose may include warping and compositing steps. Other suitable MPI viewers may be used. The rendered multi-plane image (200) can be viewed, for example, on the reference display (125).
[0032] During the warping step of the MPI rendering algorithm, each layer (C of the multi-plane image (200) i , A i ) can be warped from the RCP viewpoint position (v s ) to a new viewpoint position (v t ) as follows, for example:
Equation
Equation
[0001] T a represents the distance to the plane that is fronto-parallel to the source camera at the depth σd i .
[0033] During the synthesis step of the MPI rendering algorithm, a new visible image C t can be generated, for example, using processing operations corresponding to the following equations.
Equation
Equation
Equation
Equation
[0034] A to B in FIG. 3 shows pseudo-codes (300, 302) that can be used to implement Equation (8) according to an embodiment. More specifically, the pseudo-code (300) defines a first function that can be called to generate a weight W i s based on the alpha values of the multi-plane image (200). The pseudo-code (302) defines a second function C that can be called to render the image C s In its "STEP-1", the pseudo-code (302) calls the first function. In some cases, the weight {W i} can be known from some other process. In such a case, it is not necessary to call the pseudo-code (300) during the execution of the pseudo-code (302).
[0035] In an exemplary embodiment, the post-production editing (115) approximately minimizes the difference between the image C s and the source image R from which the corresponding multi-plane image (200) was generated, such that the image C sincludes a process aimed at adjusting. Mathematically, the goal of such a process can be formulated as follows. This process is [Number] used for. During such a conversion, the resulting rendered view corresponding to the reference camera position (RCP) is given by Equation (9). [Number] The optimization criterion for finding a suitable "optimal" conversion algorithm may be formulated using Equation (11) as follows: [Number] However, it should be noted that Equation (11) represents an underdefined problem because the field of adjustable parameters for this purpose is too large for a deterministic solution. Thus, according to various disclosed embodiments, the introduction of additional (e.g., implicit) constraints is used to obtain an approximately optimal solution. The validity of such additional constraints has been experimentally verified, and representative results according to various exemplary embodiments will be described later with reference to FIGS. 10-11.
[0036] Postproduction Editing of Multiplane Images FIG. 4 is a flowchart showing a method (400) for editing a multi-plane image (200) according to an embodiment. The method (400) uses, as input, a multi-plane image (200) that can be generated as described above, for example. The editing method (400) is applied to process the input multi-plane image (200), thereby converting the input multi-plane image into a corresponding output multi-plane image (440). When the multi-plane image (440) is rendered using, for example, the MPI rendering algorithm as described above, the appearance of artifacts in the corresponding visible image is advantageously reduced or completely suppressed compared to that in a similar rendering of the input multi-plane image (200).
[0037] The method (400) includes a first processing block (410) in which the alpha channel of the multi-plane image (200) is subjected to a normalization process. The resulting multi-plane image (412) is applied to a second processing block (420) of the method (400) in which the alpha channel and the texture channel of the image (412) are subjected to a refinement process of the alpha channel and the texture channel. The resulting multi-plane image (422) is applied to a third processing block (430) of the method (400) in which the texture channel of the image (422) is subjected to a scaling process. The output of the third processing block (430) is the multi-plane image (440). Exemplary embodiments of the processing blocks (410, 420, 430) are described in more detail below.
[0038] In some embodiments of the method (400), one of the processing blocks (410, 420, 430) may not be present. In some other embodiments of the method (400), two of the processing blocks (410, 420, 430) may not be present. In some embodiments of the method (400), the order in which the processing blocks (410, 420, 430) are executed may be different from the order shown in FIG. 4.
[0039] Alpha-Channel Normalization One exemplary type of artifact that can be observed during the rendering of the multi-plane image (200) is the "dark pixel" artifact. Analysis of the alpha-channel values {A i (x,y)} corresponding to the "dark pixel" artifact reveals that the alpha-channel values for the pixels corresponding to such an artifact are significantly lower than the alpha-channel values for the artifact-free pixels in the same multi-plane image layer, while the texture channel has a similar range of values for both sets of pixels. A similar trend appears in the weight space {W i s (x,y)}.
[0040] Figures 5A to 5D graphically show exemplary analysis results showing certain characteristics of the "dark pixel" artifact according to an embodiment. The analysis performed includes calculating the mean of absolute difference (MAD) between the source image R and the rendered image according to Equation (12) for each pixel position (x, y). [Number] The analysis performed further includes calculating the sum of weights according to Equation (13) for each pixel position (x, y): [Number]
[0041] Figures 5A to 5C respectively show the scatter plots of MAD(x, y, c) calculated for the sub-channels R, G, and B of the texture channel. [Number] Figure 5D shows the corresponding histogram of... As can be observed in Figure 5D, most of the values of the weight... [Number] cluster around 1. When... becomes smaller, the distribution of MAD(x, y, c) becomes wider and the error between the source image and the rendered image becomes larger, which is clear from the scatter plots of Figures 5A to 5C. The analysis results shown graphically in Figures 5A to 5D are... [Number] cluster around 1. When... becomes smaller, the distribution of MAD(x, y, c) becomes wider and the error between the source image and the rendered image becomes larger, which is clear from the scatter plots of Figures 5A to 5C. The analysis results shown graphically in Figures 5A to 5D are... [Number] When... becomes smaller, the distribution of MAD(x, y, c) becomes wider and the error between the source image and the rendered image becomes larger, which is clear from the scatter plots of Figures 5A to 5C. The analysis results shown graphically in Figures 5A to 5D are... [Number] Suggests that correcting the value of
[0042] When rendering at the reference camera position (RFC, Figure 2), when a multi-plane image (440) is constructed with the intention of matching the source image R, the following equation is approximately satisfied for the pixel position (x, y).
Equation
Equation
Equation
Equation
[0043] In an exemplary embodiment of the processing block (410), the normalization of the weight W for the pixel (x, y) i (x, y) can be performed according to Equation (16).
Equation
Equation
[0044] In an exemplary embodiment with a processing block (410), the normalized weights
Number
Number
Number
Number
[0045] Starting from Equation (18a), the modified alpha value for the (D - 1)th layer can be calculated using Equation (19).
Number
Number
Number
Number
Number
Number
[0046] FIG. 6 shows pseudocode (600) that can be used to implement at least a portion of the processing of processing block (410) according to an embodiment. The pseudocode (600) includes three code blocks respectively labeled as STEP-1, STEP-2, and STEP-3. STEP-1 performs the conversion of alpha-channel values to weights. In an exemplary embodiment, STEP-1 of the pseudocode (600) can be implemented using the pseudocode (300) (see A in FIG. 3). STEP-2
Number
Number
[0047] Refinement of Alpha and Texture Channels The process implemented in the first processing block (410) of the method (400) can, by itself, remove a relatively large number of artifacts, but after that process is completed, some types of artifacts still remain. One such artifact is the "black hole" artifact where the RGB pixel values for the corresponding pixels in the visible image generated by rendering the multi-plane image (412) are nulled, i.e., (0, 0, 0), while the source image R does not have in it the feature that can cause null values. Note that the "black hole" artifact typically appears as isolated small "holes" and does not occur in large quantities.
[0048] Since the rendered pixel values are a weighted linear combination of the corresponding pixel values from D layers, the weights in all D layers
Number
Number
[0049] A to B of FIG. 7 show pseudo - code (710, 720) that can be used to perform local averaging in a second processing block (420) of a method (400) according to an embodiment. More specifically, the pseudo - code (710) in FIG. 7A is configured to perform local averaging for the alpha channel. The pseudo - code (720) in FIG. 7B is similarly configured to perform local averaging for the texture channel.
[0050] Both of the pseudo - codes (710, 720) use an average filter, which is a filter that calculates the average value within a sliding window. In some embodiments, the size B of the sliding window used for the pseudo - code (710) a may be different from the size B of the sliding window used for the pseudo - code (720). c For example, the sliding window size may be B a = 19 and B c = 9. In some other embodiments, the sliding window sizes may be the same. The local channel averages calculated using the pseudo - codes (710, 720) can be expressed as follows.
Equation
[0051] FIG. 8 shows pseudo - code (800) that can be used to implement at least a part of the processing of a processing block (420) according to an embodiment. The pseudo - code (800) includes three code blocks labeled STEP - 1, STEP - 2, and STEP - 3 respectively. STEP - 1 of the pseudo - code (800) is configured to perform local averaging corresponding to equations (25), (26), and can be implemented using the pseudo - code (710, 720) (see A, B in FIG. 7). STEP - 2 of the pseudo - code (800) is configured to (i) search for "black hole" artifacts and (ii) use the corresponding local average to replace the "black hole" pixel values with the corresponding local average calculated in STEP - 1 of the pseudo - code (800).
[0052] The "black hole" artifact conditions that can be used to perform the search (i) of STEP - 2 can be expressed as follows.
Number
Number
Number
Number
[0053] STEP - 3 of the pseudo - code (800) is such that the pseudo - code (300, 302) uses the refined values generated in STEP - 2 of the pseudo - code (800)
Number
[0054] Texture-Channel Scaling FIG. 9 shows pseudo-code (900) that can be used to implement at least a portion of the processing of a processing block (430) according to an embodiment. As shown below in FIGS. 11A-11B, generally, the processing of the processing block (430), particularly the processing of the pseudo-code (900), can be beneficially used to reduce object boundary artifacts in a visible image generated by rendering an output multi-plane image (440). The processing of the processing block (430) can also reduce blurring in some portions of the visible image.
[0055] The input for the pseudo-code (900) is a source image R and a refined texture-channel value calculated in STEP-2 of the pseudo-code (800) [Number] may also include. The scaling operation of the pseudo-code (900) is applied selectively to the pixels of the visible image C that satisfy the following conditions (HN) : [Number] As already described above, the visible image C (HN) may be calculated in STEP-3 of the pseudo-code (800).
[0056] The effect of the scaling factor β applied in the pseudo-code (900) can be more clearly shown by Equation (31). [Number] More specifically, Equation (31) states that the scaling factor β is selected such that after scaling, the scaled pixel values of the visible image C (HN) are the same as the corresponding pixel values of the source image R. Rearranging Equation (31) gives the following equation for calculating the scaling factor β. [Number] The process of the pseudo-code (900) is configured to replace the pixel values that satisfy condition (30) with the corresponding scaled pixel values. Pixel values that do not satisfy condition (30) remain unchanged. Equation (33) provides an exemplary mathematical formula for such conditional scaling. [Number] This equation may be used for the pseudo-code (900) shown in FIG. 9.
[0057] Examples of Improvements A - B of FIG. 10 show exemplary visual improvements corresponding to the first processing block (410) of a method (400) according to an embodiment. More specifically, A of FIG. 10 shows a portion of a visible image generated by rendering an input multi - plane image (200) depicting a violinist during performance. The above - mentioned "dark pixel" artifacts can be clearly seen within the region surrounded by the contour around the upper arm of the violinist. B of FIG. 10 shows the same portion of the visible image generated by rendering the multi - plane image (412). A comparison between the region surrounded by the contour of the image shown in A of FIG. 10 and the corresponding region of the image shown in B of FIG. 10 provides a visual indication to the extent that the "dark pixel" artifacts can be corrected by the alpha - channel normalization of the first processing block (410) of the method (400).
[0058] A - B of FIG. 11 show exemplary visual improvements corresponding to the third processing block (430) of a method (400) according to an embodiment. More specifically, A of FIG. 11 shows a portion of a visible image generated by rendering a multi - plane image (422) obtained from the "violinist during performance" input multi - plane image (200) corresponding to A of FIG. 10. B of FIG. 11 shows the same portion of the visible image generated by rendering the output multi - plane image (440). A comparison between the region surrounded by the contour of the image shown in A of FIG. 11 and the corresponding region of the image shown in B of FIG. 11 provides a visual representation to the extent that the "object boundary" and "blur" artifacts can be corrected by the texture - channel scaling of the third processing block (430) of the method (400).
[0059] Exemplary Hardware FIG. 12 is a block diagram showing a computing device (1200) according to an embodiment. The device (1200) can be used, for example, in a post-production block (115). The device (1200) includes an input / output (I / O) device (1210), an image-enhancement engine (IEE, 1220), and a memory (1230). The I / O device (1210) can be used to enable the device (1200) to receive at least a portion of a video / image production stream (112) and output at least a portion of a final video / image stream (117). The I / O device (1210) can also be used to connect the device (1200) to a reference display (125).
[0060] The memory (1230) may include, for example, a buffer for receiving an input multi-plane image (200) by a video / image production stream (112). The input multi-plane image (200) may be in the form of, for example, an image file. Once the input multi-plane image (200) is received, the memory (1230) can provide the image file to the IEE (1220) for processing there. The IEE (1220) includes a processor (1222) and a memory (1224). The memory (1224) can store program code that, when executed by the processor (1222), enables the IEE (1220) to execute a method (400). The program code can include, among other things, program code embodying the various pseudocodes described above. Once the IEE (1220) converts the input multi-plane image (200) to a corresponding output multi-plane image (440) by executing the method (400), the IEE (1220) can perform its rendering process and provide a corresponding visible image for viewing on the reference display (125). The visible image may be in the form of, for example, an appropriate image file output through the I / O device (1210).
[0061] For example, according to the exemplary embodiments disclosed above in the summary section and / or with respect to any one or part or all of FIGS. 1 to 12 or any combination thereof, an apparatus for improving a first multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position is provided. The apparatus includes at least one processor (e.g., 1222 in FIG. 12) and at least one memory (e.g., 1224 in FIG. 12) including program code, and the at least one memory and the program code, together with the at least one processor, cause the apparatus to at least: for each pixel of a first set of pixels, scale the respective weights of the layers (e.g., STEP-2 in FIG. 6) to make the sum of the scaled weights equal to a predetermined fixed value; for each pixel of a second set of pixels, replace the respective alpha values and texture values in the layer with corresponding local average values (e.g., STEP-2 in FIG. 8); and for each pixel of a third set of pixels, scale the corresponding texture values in the layer such that the texture value of each pixel of the third set matches the respective texture value of a reference image captured from the reference camera position for the resulting visible image rendered with respect to the reference camera position (e.g., STEP-1 in FIG. 9). In various embodiments, the predetermined fixed value can be one or other appropriately selected positive fixed value.
[0062] In some embodiments of the above apparatus, the second set is an empty set. The second set of pixels is empty, for example, when there is no "black hole" artifact. The first and third sets of pixels are typically not empty.
[0063] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are configured to further cause the device to generate a second multi-plane image (e.g., 412 in FIG. 4). This is by at least: converting the alpha value of the first multi-plane image into a corresponding weight value (e.g., STEP-1 in FIG. 6); identifying a set of first pixels based on the corresponding weight value; and calculating an alpha value for the second multi-plane image using the scaled weights (e.g., STEP-3 in FIG. 6).
[0064] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are configured to further cause the device to calculate an alpha value for the second multi-plane image by recursive backpropagation of the scaled weights (e.g., equations (19)-(22)).
[0065] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are configured to further cause the device to identify a set of second pixels by at least finding one or more null texture values in a visible image generated based on the second multi-plane image (e.g., equation (23)).
[0066] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are configured to further cause the device to generate a third multi-plane image (e.g., 422 in FIG. 4) based on the second multi-plane image, wherein a second set of pixels of the third multi-plane image has a corresponding local average value as the pixel value therein.
[0067] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are further configured to cause the device to perform alpha-channel normalization for a third multi-plane image (e.g., STEP-3 in FIG. 8).
[0068] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are further configured to cause the device to perform alpha-weight conversion for alpha-channel normalization.
[0069] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are further configured to cause the device to generate a fourth multi-plane image (e.g., 440 in FIG. 4) based on the third multi-plane image, and a third set of pixels of the fourth multi-plane image has alpha values and texture values that cause a match.
[0070] In some embodiments of any of the above devices, the at least one memory and program code, together with the at least one processor, are further configured to cause the device to generate another visible image by rendering a fourth multi-plane image for a virtual camera position different from the reference camera position.
[0071] For example, according to another exemplary embodiment disclosed above with reference to the summary section and / or any one or part or all of FIGS. 1 to 12 or any combination thereof, a method for improving a first multi-plane image represented by a plurality of layers corresponding to different respective distances from a reference camera position is provided. The method includes: for each pixel in a first set of pixels, scaling the respective weights of the layers (e.g., STEP-2 in FIG. 6) to make the sum of the scaled weights equal to a predetermined fixed value, wherein the scaling of each weight is executed by at least one processor (e.g., 1222 in FIG. 12) and at least one memory (e.g., 1224 in FIG. 12) including program code; for each pixel in a second set of pixels, replacing the respective alpha value and texture value in the layer with corresponding local average values (e.g., STEP-2 in FIG. 8), wherein the replacement is executed by the at least one processor and the at least one memory; for each pixel in a third set of pixels, scaling the corresponding texture value in the layer so that the texture value of each pixel in the third set matches the respective texture value of a reference image captured from the reference camera position for the resulting visible image rendered for the reference camera position (e.g., STEP-1 in FIG. 9), wherein the scaling of the corresponding texture value is executed by the at least one processor and the at least one memory.
[0072] In some embodiments of the above method, the method further includes generating a second multi-plane image (e.g., 412 in FIG. 4). This is by at least: converting the alpha value of the first multi-plane image into a corresponding weight value (e.g., STEP-1 in FIG. 6); identifying a first set of pixels based on the corresponding weight values; and calculating an alpha value for the second multi-plane image using the scaled weights (e.g., STEP-3 in FIG. 6).
[0073] In some embodiments of any of the above methods, the method further includes calculating alpha values for a second multi-plane image by recursively backpropagating the scaled weights (e.g., Equations (19)-(22)).
[0074] In some embodiments of any of the above methods, the method further includes identifying a second set of pixels by at least finding one or more null-texture values in a visible image generated based on the second multi-plane image (e.g., Equation (23)).
[0075] In some embodiments of any of the above methods, the method further includes generating a third multi-plane image (e.g., 422 in FIG. 4) based on the second multi-plane image, wherein a second set of pixels of the third multi-plane image has corresponding local average values as pixel values therein.
[0076] In some embodiments of any of the above methods, the method further includes performing alpha-channel normalization for the third multi-plane image (e.g., STEP-3 in FIG. 8).
[0077] In some embodiments of any of the above methods, the method further includes performing an alpha-weight transformation for alpha-channel normalization.
[0078] In some embodiments of any of the above methods, the method further includes generating a fourth multi-plane image (e.g., 440 in FIG. 4) based on the third multi-plane image, wherein a third set of pixels of the fourth multi-plane image has alpha values and texture values that cause the said match.
[0079] In some embodiments of any of the above methods, the method further includes generating another visible image by rendering a fourth multi-plane image for a virtual camera position different from the reference camera position.
[0080] For example, according to yet another exemplary embodiment disclosed above in the summary section and / or with respect to any one or part or all of FIGS. 1-12 or any combination thereof, a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations is provided. The operations include a method for enhancing a first multi-plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position, the method comprising: for each pixel of a first set of pixels, scaling each weight of the layers (e.g., STEP-2 in FIG. 6) such that the sum of the scaled weights equals a predetermined fixed value, wherein the scaling of each weight is performed by at least one processor (e.g., 1222 in FIG. 12) and at least one memory (e.g., 1224 in FIG. 12) including program code; for each pixel of a second set of pixels, replacing each alpha value and texture value in the layer with a corresponding local average value (e.g., STEP-2 in FIG. 8), wherein the replacement is performed by the at least one processor and the at least one memory; and for each pixel of a third set of pixels, scaling the corresponding texture value in the layer such that the texture value of each pixel of the third set matches the respective texture value of a reference image captured from the reference camera position for the resulting visible image rendered with respect to the reference camera position (e.g., STEP-1 in FIG. 9), wherein the scaling of the corresponding texture value is performed by the at least one processor and the at least one memory.
[0081] Regarding the processes, systems, methods, heuristics, etc. described herein, although the steps of such processes, etc. are described as occurring according to an ordered sequence, it should be understood that such processes can be implemented using the described steps in an order other than that described herein. Further, it should be understood that certain steps can be executed simultaneously, other steps can be added, or certain steps described herein can be omitted. In other words, the description of the processes herein is provided for the purpose of exemplifying certain embodiments and should in no way be construed as limiting the scope of the claims.
[0082] Thus, it should be understood that the above description is intended to be illustrative and not limiting. Many embodiments and applications other than the examples provided will be apparent upon reading the above description. The scope should not be determined with reference to the above description, but rather with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. Future developments are expected and intended in the technologies discussed herein, and the disclosed systems and methods to be incorporated into such future embodiments. In short, it should be understood that this application is capable of modification and variation.
[0083] All terms used in the claims are intended to be given their broadest reasonable interpretation and their ordinary meaning as would be understood by one skilled in the art of the technology described herein, unless explicitly indicated to the contrary herein. In particular, the use of singular articles such as "a," "the," "said," etc. should be read as reciting one or more of the indicated elements, unless the claim describes an explicit limitation to the contrary.
[0084] The abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, for the purpose of better presenting the flow of the present disclosure, it can be seen that various features are grouped together in various embodiments. This method of disclosure should not be construed as reflecting an intention that the claimed embodiments incorporate more features than are explicitly recited in each claim. Rather, as reflected by the following claims, the subject matter of the invention lies in less than all of the features of a single disclosed embodiment. Thus, the following claims are incorporated herein by reference, and each claim stands on its own as a separately claimed subject matter.
[0085] The present disclosure includes references to exemplary embodiments, but the present specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the present disclosure, which are apparent to those of ordinary skill in the art to which the present disclosure pertains, are considered to be within the principles and scope of the present disclosure, as represented, for example, by the following claims.
[0086] Some embodiments may be embodied in the form of methods and apparatuses for performing those methods. Some embodiments may also be embodied in the form of program code recorded on a tangible medium such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy (registered trademark) disk, a CD-ROM, a hard drive, or any other non-transitory machine-readable storage medium, and when the program code is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing the patented invention. When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates in a manner similar to a specific logic circuit.
[0087] Unless otherwise stated explicitly, each numerical value and range should be construed as approximate as if the word "about" or "approximately" preceded the value or range.
[0088] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the subject matter recited in the claims to facilitate the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0089] The elements in the claims of the following methods are, if any, described in a specific order using corresponding labeling, but unless the claim recitation otherwise implies a specific order for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that specific order.
[0090] References to "one embodiment" or "an embodiment" in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of the present disclosure. The appearance of the phrase "in an embodiment" in various places in this specification does not necessarily refer to the same embodiment, and distinct or alternative embodiments do not necessarily exclude each other. The same applies to the term "implementation".
[0091] Unless otherwise specified herein, the use of ordinal adjectives such as "first", "second", "third", etc. to refer to one of a plurality of similar objects merely indicates that different instances of such similar objects are being referred to, and does not imply that the similar objects so referred to must be in a corresponding order or sequence in time, space, ranking, or any other way.
[0092] Unless otherwise specified in this specification, in addition to its plain meaning, the conjunction "when... " may also or alternatively be construed to mean "at the time of... " or "upon... " or "in response to determining... " or "in response to detecting... ", and such construction may depend on the particular context in question. For example, the phrases "when... is determined" or "when [stated condition] is detected" may be construed to mean "upon determining... " or "in response to determining... " or "upon detecting [stated condition or event]" or "in response to detecting [stated condition or event]".
[0093] Also, for the purposes of this description, the terms "couple", "coupling", "coupled", "connect", "connection", or "connected" refer to any manner known in the art or later developed in which energy is permitted to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, but not required. Conversely, terms such as "directly coupled", "directly connected", etc. mean that no such additional elements are present.
[0094] The functionality of the various elements shown in the figures, including any functional blocks labeled as and / or referred to as "processor" and / or "controller", can be provided through dedicated hardware as well as through the use of hardware capable of executing software in association with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Further, the explicit use of the term "processor" or "controller" should not be construed to refer only to hardware capable of executing software, and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage. Conventional and / or custom other hardware may also be included. Similarly, any switches shown in the figures are conceptual only. Their functionality may be carried out through the operation of program logic, through dedicated logic, through an interaction of program control and dedicated logic, or even manually, and the particular technique can be selected by the implementer as more specifically understood from the context.
[0095] As used herein, the terms "circuit" and "circuits" refer to (a) a circuit implementation of only hardware (such as an implementation with only analog and / or digital circuit configurations), (b) (where applicable) (i) a combination of analog and / or digital hardware circuits with software / firmware, and (ii) a hardware processor (including a digital signal processor), software, and any portion of memory that together operate to cause a device such as a mobile phone or a server to perform various functions, and (c) a hardware circuit and / or processor such as a microprocessor or a portion of a microprocessor that requires software (such as firmware) to operate, but may not exist when the software is not required for operation, one or more or all of these. This definition of circuit applies to all uses of this term in this application, including in the claims. As a further example, as used herein, the term "circuit" also covers simply a hardware circuit or processor (or processors), or a portion of a hardware circuit or processor, and its (or their) associated software and / or firmware implementation. The term "circuit" covers, for example, a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit within a server, a cellular network device, or other computing or network device, if applicable to the elements of a particular claim.
[0096] It should be understood by those skilled in the art that any block diagram in this specification represents a conceptual diagram of an exemplary circuit embodying the principles of the present disclosure. Similarly, any flowchart, flow diagram, state transition diagram, pseudocode, etc. is substantially represented in a computer-readable medium and thus represents various processes that can be executed by such a computer or processor, whether or not such a computer or processor is explicitly shown.
[0097] The "Summary of the Invention" in this specification is intended to introduce some exemplary embodiments, and additional embodiments are described in the "Detailed Description of the Invention" and / or with reference to one or more drawings. The "Summary of the Invention" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
1. An apparatus for improving a first multi - plane image represented by a plurality of layers corresponding to respective different distances from a reference camera position, the apparatus comprising: at least one processor; at least one memory containing program code, wherein the at least one memory and the program code, together with the at least one processor, cause the apparatus to at least: scale the respective weights of the plurality of layers such that the sum of the scaled weights equals a predetermined fixed value for each pixel of a first set of pixels; replace the respective alpha values and texture values in the plurality of layers with corresponding local average values for each pixel of a second set of pixels; scale the corresponding texture values in the plurality of layers such that the texture value of each pixel of the third set of pixels matches the respective texture value of a reference image captured from the reference camera position for a resulting visible image rendered with respect to the reference camera position, for each pixel of the third set of pixels. An apparatus.
2. The apparatus according to claim 1, wherein the second set is an empty set.
3. The at least one memory and the program code, together with the at least one processor, further cause the apparatus to at least: convert the alpha value of the first multi - plane image into a corresponding weight value; identify the first set of pixels based on the corresponding weight value; generate a second multi - plane image by calculating an alpha value for the second multi - plane image using the scaled weights. The apparatus according to claim 1.
4. The apparatus according to claim 3, wherein the at least one memory and the program code, together with the at least one processor, further cause the apparatus to calculate the alpha value for the second multi - plane image by recursive back - propagation of the scaled weights.
5. The at least one memory and the program code, together with the at least one processor, are further configured to cause the apparatus to identify the second set of pixels by finding one or more null texture values in a visible image generated based at least on the second multi-plane image, for the apparatus according to claim 3 or 4.
6. The at least one memory and the program code, together with the at least one processor, are configured to further cause the apparatus to generate a third multi-plane image based on the second multi-plane image, wherein the second set of pixels of the third multi-plane image has the corresponding local average value as a pixel value therein, for the apparatus according to claim 3.
7. The at least one memory and the program code, together with the at least one processor, are configured to further cause the apparatus to perform alpha-channel normalization on the third multi-plane image, for the apparatus according to claim 6.
8. The at least one memory and the program code, together with the at least one processor, are configured to further cause the apparatus to perform an alpha-weight transformation for the alpha-channel normalization, for the apparatus according to claim 7.
9. The at least one memory and the program code, together with the at least one processor, are configured to further cause the apparatus to generate a fourth multi-plane image (e.g., 440 in FIG. 4) based on the third multi-plane image, wherein the third set of pixels of the fourth multi-plane image has an alpha value and a texture value that cause the match, for the apparatus according to any one of claims 6 to 8.
10. The at least one memory and the program code, together with the at least one processor, are configured to further cause the apparatus to generate another visible image by rendering the fourth multi-plane image for a virtual camera position different from the reference camera position, for the apparatus according to claim 9.
11. A method for improving a first multi-plane image represented by a plurality of layers corresponding to different respective distances from a reference camera position, the method comprising: scaling the respective weights of the plurality of layers such that the sum of the scaled weights for each pixel in a first set of pixels is a predetermined fixed value, the scaling of each weight being performed using at least one processor and at least one memory including program code; replacing, for each pixel in a second set of pixels, the respective alpha values and texture values in the plurality of layers with corresponding local average values, the replacement being performed using the at least one processor and the at least one memory; scaling the corresponding texture values in the plurality of layers such that the texture value of each pixel in the third set of pixels matches the respective texture value of a reference image captured from the reference camera position for a resulting visible image rendered for the reference camera position, the scaling of the corresponding texture values being performed using the at least one processor and the at least one memory, the method. **Claim 12** At least: converting the alpha values of the first multi-plane image to corresponding weight values; identifying the first set of pixels based on the corresponding weight values; generating a second multi-plane image by calculating alpha values for the second multi-plane image using the scaled weights, further comprising the step of The method according to claim 11. **Claim 13** The method according to claim 12, further comprising the step of calculating the alpha values for the second multi-plane image by recursive backpropagation of the scaled weights. **Claim 14** The method according to claim 12, further comprising the step of identifying the second set of pixels by finding one or more null texture values in a visible image generated based at least on the second multi-plane image. **Claim 15** The method according to claim 12, further comprising generating a third multi-plane image based on the second multi-plane image, wherein the second set of pixels of the third multi-plane image has the corresponding local average value as a pixel value therein.
16. The method according to claim 15, further comprising performing alpha-channel normalization on the third multi-plane image.
17. The method according to claim 16, further comprising performing an alpha-weight transformation for the alpha-channel normalization.
18. The method according to claim 15, further comprising generating a fourth multi-plane image based on the third multi-plane image, wherein the third set of pixels of the fourth multi-plane image has alpha values and texture values that cause the match.
19. The method according to claim 18, further comprising generating another visible image by rendering the fourth multi-plane image at a virtual camera position different from the reference camera position.
20. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including the method according to any one of claims 11 to 19.