Encoding Scheme for Video Data Using Downsampling / Upsampling and Nonlinear Filtering of Depth Maps

By performing nonlinear filtering and downsampling of the depth map of immersive video, combined with the analysis of texture maps, the problem of difficulty in effectively reducing the pixel rate in the prior art is solved, and more efficient coding and better user experience is achieved.

CN114868401BActive Publication Date: 2025-05-27KONINKLIJKE PHILIPS NV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080088330.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-12-17
Publication Date
2025-05-27
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reduce pixel rate in immersive video encoding without damaging user quality, and detailed analysis may lead to errors and lead to a decline in end user experience.

Method used

The processed depth map is generated by nonlinear filtering and downsampling the depth map of the source view, and the depth map and texture map are encoded to generate a video code stream. Nonlinear filtering may include enlarging the area of ​​the foreground object to avoid damage to the foreground object by downsampling.

Benefits of technology

It realizes reducing pixel rate without damaging video quality, improving encoding efficiency, and effectively retaining the details of the foreground object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114868401B_ABST
    Figure CN114868401B_ABST
Patent Text Reader

Abstract

Methods for encoding and decoding video data are provided. In one encoding method, source video data including one or more source views is encoded into a video bitstream. Before encoding, depth data of at least one of the source views in the source views is non-linearly filtered and downsampled. After decoding, the decoded depth data is upsampled and non-linearly filtered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video coding. In particular, the present invention relates to methods and apparatuses for encoding and decoding immersive video. Background Art

[0002] Immersive video (also known as six degrees of freedom (6DoF) video) is a video of a three-dimensional (3D) scene that allows for the reconstruction of views of the scene for viewpoints that vary in position and orientation. It represents a further development of three degrees of freedom (3DoF) video, which allows for the reconstruction of views for viewpoints that have arbitrary orientations but are only at fixed spatial points. In 3DoF, the degrees of freedom are angles - namely, pitch angle, roll angle, and yaw angle. 3DoF video supports head rotation - in other words, a user consuming video content can look in any direction within the scene, but cannot move to different positions within the scene. 6DoF video supports head rotation and additionally supports choosing the position within the scene from which to view the scene.

[0003] To generate 6DoF video, multiple cameras are required to record the scene. Each camera generates image data (commonly referred to as texture data in this context) and corresponding depth data. For each pixel, the depth data represents the depth at which the corresponding image pixel data is observed through a given camera. Each of the multiple cameras provides a corresponding view of the scene. In many applications, it may be impractical or inefficient to transmit all the texture data and depth data for all views.

[0004] To reduce redundancy between views, it has been proposed to streamline the views and package the views into "texture atlases" for each frame of the video stream. This approach attempts to reduce or eliminate the overlapping portions between multiple views, thereby improving efficiency. The non-overlapping portions of the different views remaining after streamlining can be referred to as "tiles". An example of this approach is described by Alvaro Collet et al. in "High-quality streamable free-viewpoint video" (ACM Trans.Graphics (SIGGRAPH), 34(4), 2015). Summary of the Invention

[0005] There is a desire to improve the quality and encoding efficiency of immersive video. As described above, the method of using streamlining (i.e., excluding redundant texture tiles) to produce texture atlases can help reduce the pixel rate. However, streamlining the views typically requires detailed analysis, which is not error-free and may result in a reduction in quality for the end user. Therefore, there is a need for robust and simple methods for reducing the pixel rate.

[0006] The present invention is defined by the claims.

[0007] According to an example of one aspect of the present invention, a method for encoding video data including one or more source views, each source view including a texture map and a depth map, is provided. The method includes:

[0008] Receiving the video data;

[0009] Processing the depth map of at least one source view to generate a processed depth map, the processing including:

[0010] Non-linear filtering, and

[0011] Downsampling; and

[0012] Encoding the processed depth map and the texture map of the at least one source view to generate a video bitstream.

[0013] Preferably, at least part of the non-linear filtering is performed before downsampling.

[0014] The inventors have found that non-linear filtering of the depth map before downsampling can help avoid, reduce or mitigate errors introduced by downsampling. In particular, non-linear filtering can help prevent small or thin foreground objects from partially disappearing or completely disappearing from the depth map due to downsampling. It has been found that in this regard, non-linear filtering is superior to linear filtering because linear filtering introduces intermediate depth values at the boundaries between foreground objects and the background. This makes it difficult for the decoder to distinguish object boundaries from large depth gradients.

[0015] The video data may include 6DoF immersive video.

[0016] The non-linear filtering includes expanding the area of at least one foreground object in the depth map.

[0017] Enlarging the foreground object before downsampling can help ensure better avoidance of damage to the foreground object during the downsampling process - in other words, better retention of the foreground object in the processed depth map.

[0018] Foreground objects can be identified as a set of local pixels at relatively small depths. The background can be identified as pixels at relatively large depths. The perimeter pixels of the foreground object can be locally distinguished from the background, for example, by applying a threshold to the depth values in the depth map.

[0019] Nonlinear filtering may include morphological filtering, especially grayscale morphological filtering, e.g., maximum filter, minimum filter, or other order filters. When the depth map contains depth levels with special meanings (e.g., depth level "zero" indicates invalid depth), then although such depth levels have actual values, such depth levels should preferably be considered as foreground. For this reason, these levels are preferably retained after downsampling. Thus, their area will also expand.

[0020] The nonlinear filtering may include applying a filter designed using a machine learning algorithm.

[0021] The machine learning algorithm can be trained to reduce or minimize the reconstruction error of the reconstructed depth map after the processed depth map has been encoded and decoded.

[0022] The trained filter can similarly help to retain foreground objects in the processed (downsampled) depth map.

[0023] The method may also include using a machine learning algorithm to design a filter, where the filter is designed to reduce the reconstruction error of the reconstructed depth map after the processed depth map has been encoded and decoded, and where the nonlinear filtering includes applying the designed filter.

[0024] Nonlinear filtering may include processing through a neural network, and the design of the filter may include training the neural network.

[0025] The nonlinear filtering may be performed by a neural network including multiple layers, and the downsampling may be performed between two of the multiple layers.

[0026] The downsampling may be performed through a max pooling (or min pooling) layer of the neural network.

[0027] The method may include processing the depth map according to multiple sets of processing parameters to generate corresponding multiple processed depth maps, and the method further includes: selecting, after the corresponding processed depth maps have been encoded and decoded, the set of processing parameters that reduces the reconstruction error of the reconstructed depth map; and generating a metadata bitstream identifying the selected set of parameters.

[0028] This can allow optimizing the parameters for a given application or a given video sequence.

[0029] The processing parameters may include the definition of the nonlinear filtering performed and / or the definition of the downsampling performed. Alternatively or additionally, the processing parameters may include the definition of the processing operations to be performed at the decoder when reconstructing the depth map.

[0030] For each set of processing parameters, the method may include: generating a corresponding processed depth map according to the set of processing parameters; encoding the processed depth map to generate an encoded depth map; decoding the encoded depth map; reconstructing the depth map according to the decoded depth map; and comparing the reconstructed depth map with the depth maps of at least one source view to determine a reconstruction error.

[0031] According to another aspect, a method for decoding video data including one or more source views is provided, the method including:

[0032] receiving a video bitstream including encoded depth maps and encoded texture maps for at least one source view;

[0033] decoding the encoded depth maps to produce decoded depth maps;

[0034] decoding the encoded texture maps to produce decoded texture maps; and

[0035] processing the decoded depth maps to generate reconstructed depth maps, wherein the processing includes:

[0036] upsampling, and

[0037] non-linear filtering.

[0038] The method may further include: before the step of processing the decoded depth maps to generate the reconstructed depth maps, detecting that the resolution of the decoded depth maps is lower than the resolution of the decoded texture maps.

[0039] In some coding schemes, depth maps may be downsampled only in certain cases or only for certain views. By comparing the resolution of the decoded depth maps with the resolution of the decoded texture maps, the decoding method can determine whether downsampling was applied at the encoder. This can avoid the need for metadata in the metadata bitstream to signal which depth maps were downsampled and to what extent these depth maps were downsampled. (In this example, it is assumed that the texture maps are encoded at full resolution.)

[0040] To generate the reconstructed depth maps, the decoded depth maps may be upsampled to the same resolution as the decoded texture maps.

[0041] Preferably, the non-linear filtering in the decoding method is adapted to compensate for the effects of the non-linear filtering applied in the encoding method.

[0042] Non-linear filtering may include reducing the area of at least one foreground object in the depth map. This may be appropriate when non-linear filtering during encoding includes increasing the area of at least one foreground object.

[0043] Non-linear filtering may include morphological filtering, in particular grayscale morphological filtering, for example, a maximum filter, a minimum filter or other order filters.

[0044] Non-linear filtering during decoding preferably compensates for or reverses the effect of non-linear filtering during encoding. For example, if non-linear filtering during encoding includes a maximum filter (grayscale enhancement), then non-linear filtering during decoding may include a minimum filter (grayscale attenuation), and vice versa. When the depth map contains depth levels with special meanings (for example, the depth level "zero" indicates an invalid depth), then such depth levels should preferably be considered as foregrounds even though such depth levels have actual values.

[0045] Preferably, at least part of the non-linear filtering is performed after the upsampling. Optionally, all non-linear filtering is performed after upsampling.

[0046] The processing of the decoded depth map may be at least partially based on the decoded texture map. The inventors have recognized that: the texture map contains useful information that helps to reconstruct the depth map. In particular, in the case where the boundaries of foreground objects are changed by non-linear filtering during encoding, the analysis of the texture map can help to compensate for or reverse the changes.

[0047] The method may include: upsampling the decoded depth map; identifying the peripheral pixels of at least one foreground object in the upsampled depth map; determining whether the peripheral pixels are more similar to the foreground object or the background based on the decoded texture map; and applying non-linear filtering only to the peripheral pixels determined to be more similar to the background.

[0048] In this way, the texture map is used to help identify the pixels that have been converted from the background to the foreground due to non-linear filtering during encoding. Non-linear filtering during decoding may help to restore these identified pixels to parts of the background.

[0049] The non-linear filtering may include smoothing the edges of at least one foreground object.

[0050] The smoothing process may include: identifying peripheral pixels of at least one foreground object in the upsampled depth map; for each peripheral pixel, analyzing the number and / or arrangement of foreground pixels and background pixels in the neighborhood around the peripheral pixel; identifying outlying peripheral pixels projected from the object into the background based on the result of the analysis; and applying non-linear filtering only to the identified peripheral pixels.

[0051] The analysis may include counting the number of background pixels in the neighborhood, wherein if the number of background pixels in the neighborhood is higher than a predefined threshold, the peripheral pixel is identified as an outlier deviating from the object.

[0052] Alternatively or additionally, the analysis may include identifying the spatial pattern of foreground pixels and background pixels in the neighborhood, wherein if the neighborhood of the peripheral pixel matches one or more predefined spatial patterns, the peripheral pixel is identified as an outlier.

[0053] The method may further include receiving a metadata stream associated with the video stream, the metadata stream identifying a set of parameters, and the method optionally further includes processing the decoded depth map according to the identified set of parameters.

[0054] The processing parameters may include the definition of the non-linear filtering and / or the definition of the upsampling to be performed.

[0055] The non-linear filtering may include applying a filter designed using a machine learning algorithm.

[0056] The machine learning algorithm may be trained to reduce or minimize the reconstruction error of the reconstructed depth map after the processed depth map has been encoded and decoded.

[0057] The filter may be defined in a metadata stream associated with the video stream.

[0058] There is also provided a computer program comprising computer code which, when the program runs on a processing system, causes the processing system to implement the method outlined above.

[0059] The computer program may be stored on a computer-readable storage medium. This may be a non-transitory storage medium.

[0060] According to another aspect, there is provided a video encoder configured to encode video data including one or more source views, each source view including a texture map and a depth map, the video encoder comprising:

[0061] An input unit configured to receive the video data;

[0062] A video processor configured to process the depth map of at least one source view to generate a processed depth map, the processing including:

[0063] Non-linear filtering, and

[0064] Downsampling;

[0065] An encoder configured to encode the texture map and the processed depth map of the at least one source view to generate a video bitstream; and

[0066] An output unit configured to output the video bitstream.

[0067] According to another aspect, there is provided a video decoder configured to decode video data including one or more source views, the video decoder including:

[0068] A bitstream input unit configured to receive a video bitstream, wherein the video bitstream includes an encoded depth map and an encoded texture map for at least one source view;

[0069] A first decoder configured to decode the encoded depth map according to the video bitstream to generate a decoded depth map;

[0070] A second decoder configured to decode the encoded texture map according to the video bitstream to generate a decoded texture map;

[0071] A reconstruction processor configured to process the decoded depth map to generate a reconstructed depth map, wherein the processing includes:

[0072] Upsampling, and

[0073] Non-linear filtering, and

[0074] An output unit configured to output the reconstructed depth map.

[0075] Referring to the embodiments described below, these and other aspects of the present invention will become apparent and be elucidated. Description of the Drawings

[0076] For a better understanding of the present invention and to more clearly show how the present invention can be put into practice, reference will now be made, by way of example only, to the accompanying drawings, in which:

[0077] Figure 1 An example of encoding and decoding immersive video using an existing video codec is illustrated;

[0078] Figure 2is a flowchart showing a method for encoding video data according to an embodiment;

[0079] Figure 3 is a block diagram of a video encoder according to an embodiment;

[0080] Figure 4 is a flowchart illustrating a method for encoding video data according to another embodiment;

[0081] Figure 5 is a flowchart showing a method for decoding video data according to an embodiment;

[0082] Figure 6 is a block diagram of a video decoder according to an embodiment;

[0083] Figure 7 illustrates a method for selectively applying non - linear filtering to specific pixels in a decoding method according to an embodiment;

[0084] Figure 8 is a flowchart illustrating a method for decoding video data according to another embodiment; and

[0085] Figure 9 illustrates a usage process for encoding and decoding video data using neural network processing according to an embodiment. Detailed Description

[0086] The present invention will be described with reference to the accompanying drawings.

[0087] It should be understood that although the detailed description and specific examples indicate exemplary embodiments of the apparatus, system, and method, the detailed description and specific examples are for illustrative purposes only and are not intended to limit the scope of the present invention. These and other features, aspects, and advantages of the apparatus, system, and method of the present invention will be better understood from the following description, claims, and drawings. It should be understood that these drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout all the drawings to indicate the same or similar parts.

[0088] Methods for encoding and decoding immersive video are disclosed. In one encoding method, source video data including one or more source views is encoded into a video bitstream. Before encoding, non - linear filtering and downsampling are performed on the depth data of at least one of the source views in the source video. Downsampling the depth map helps reduce the amount of data to be transmitted and thus helps reduce the bit rate. However, the inventors have found that simple downsampling can cause thin or small foreground objects (e.g., wires) to disappear from the downsampled depth map. Embodiments of the present invention attempt to mitigate this effect and retain small and thin objects in the depth map.

[0089] Embodiments of the present invention may be suitable for implementing parts of technical standards, for example, immersive video of ISO / IEC 23090-12 MPEG-I Part 12. Where possible, the terms used herein are chosen to be consistent with the terms used in MPEG-I Part 12. Nevertheless, it should be understood that the scope of the present invention is not limited to MPEG-I Part 12, nor to any other technical standard.

[0090] It is helpful to set forth the following definitions / interpretations:

[0091] A "3D scene" refers to visual content in a global reference coordinate system.

[0092] An "atlas" is the aggregated content that aggregates patches from one or more view representations after the packaging process into a pair of pictures including a texture component picture and a corresponding depth component picture.

[0093] An "atlas component" is a texture component or a depth component of the atlas.

[0094] "Camera parameters" define the projection used to generate a view representation from a 3D scene.

[0095] "Thinning" is the process of identifying and extracting occluded regions across views to obtain patches.

[0096] "Rendering" is an embodiment of the process of creating a viewport or omnidirectional view corresponding to a viewing position and orientation from a 3D scene representation.

[0097] A "source view" is the source video material before encoding corresponding to the format of the view representation, which can be acquired by capturing a 3D scene with a real camera, or by projecting onto a surface using a virtual camera with source camera parameters.

[0098] A "target view" is defined as a perspective viewport or omnidirectional view at a desired viewing position and orientation.

[0099] A "view representation" includes a 2D sample array of a texture component and a corresponding depth component, which represents the projection of a 3D scene onto a surface using camera parameters.

[0100] A machine learning algorithm is any self-training algorithm that processes input data to generate or predict output data. In some embodiments of the present invention, the input data includes one or more views decoded from a bitstream, and the output data includes the prediction result / reconstruction result of the target view.

[0101] Machine learning algorithms suitable for use in the present invention will be obvious to those skilled in the art. Examples of suitable machine learning algorithms include decision tree algorithms and artificial neural networks. Other machine learning algorithms (e.g., logistic regression, support vector machines, or naive Bayes models) are suitable alternatives.

[0102] The structure of an artificial neural network (or simply neural network) is inspired by the human brain. A neural network consists of multiple layers, each layer including a plurality of neurons. Each neuron includes mathematical operations. In particular, each neuron can include different weighted combinations of a single type of transformation (e.g., the same type of transformation, sigmoid, etc., but weighted differently). During the process of processing input data, the mathematical operations of each neuron are performed on the input data to produce a numerical output, and the output of each layer in the neural network (e.g., sequentially) is fed into one or more other layers. The last layer provides the output.

[0103] Methods for training machine learning algorithms are well known. Generally, such methods include obtaining a training data set, which includes training input data entries and corresponding training output data entries. An initialized machine learning algorithm is applied to each input data entry to generate predicted output data entries. The error between the predicted output data entries and the corresponding training output data entries is used to modify the machine learning algorithm. This process can be repeated until the error converges and the predicted output data entries are sufficiently similar to the training output data entries (e.g., ±1%). This is generally referred to as a supervised learning technique.

[0104] For example, in the case where the machine learning algorithm is formed by a neural network, the mathematical operations (weights) of each neuron can be modified until the error converges. Known methods for modifying neural networks include gradient descent, backpropagation algorithms, etc.

[0105] A convolutional neural network (CNN or ConvNet) is a class of deep neural network that is most commonly used for analyzing visual images. A CNN is a regularized version of a multi-layer perceptron.

[0106] Figure 1A system for encoding and decoding immersive video is illustrated in simplified form. An array of cameras 10 is used to capture multiple views of a scene. Each camera captures a conventional image (referred to herein as a texture map) and a depth map of the view in front of the texture map. A set of views including texture data and depth data is provided to an encoder 300. The encoder encodes both the texture data and the depth data into a conventional video bitstream - in this case, an efficient video coding (HEVC) bitstream. This is accompanied by a metadata bitstream to inform a decoder 400 of the meaning of different parts of the video bitstream. For example, the metadata tells the decoder which parts of the video bitstream correspond to texture maps and which parts correspond to depth maps. Depending on the complexity and flexibility of the encoding scheme, more or less metadata may be required. For example, a very simple scheme may tightly prescribe the structure of the bitstream such that little or no metadata is required for unpacking the bitstream at the decoder side. In cases where the bitstream has more optional possibilities, a larger amount of metadata will be required.

[0107] The decoder 400 decodes the encoded (texture and depth) views. The decoder 400 passes the decoded views to a synthesizer 500. The synthesizer 500 is coupled to a display device, such as a virtual reality headset 550. The headset 550 requests the synthesizer 500 to synthesize and render a particular view of the 3D scene using the decoded views based on the current position and orientation of the headset 550.

[0108] Figure 1 An advantage of the system shown is that it can use a conventional 2D video codec to encode and decode texture data and depth data. However, a disadvantage is that a large amount of data needs to be encoded, transmitted, and decoded. Therefore, it is desirable to reduce the data rate while degrading the quality of the reconstructed views as little as possible.

[0109] Figure 2 An encoding method according to a first embodiment is illustrated. Figure 3 An illustration of a video encoder that can be configured to perform Figure 2 the method. The video encoder includes: an input section 310 configured to receive video data; a video processor 320 coupled to the input section and configured to receive the depth map received by the input section; an encoder 330 arranged to receive the processed depth map from the video processor 320; and an output section 370 arranged to receive and output the video bitstream generated by the encoder 330. The video encoder 300 also includes a depth decoder 340, a reconstruction processor 350, and an optimizer 360. These components will be described in more detail in connection with a second embodiment of the encoding method described below with reference to Figure 4 The encoding method described below with reference to

[0110] ReferenceFigure 2 and Figure 3 For the method of the first embodiment, it starts at step 110, where the input unit 310 receives video data including a texture map and a depth map. In steps 120 and 130, the video processor 320 processes the depth map to generate a processed depth map. Such processing includes non-linear filtering of the depth map in step 120 and downsampling of the filtered depth map in step 130. In step 140, the encoder 330 encodes the processed depth map and the texture map to generate a video bitstream. Then the generated video bitstream is output via the output unit 370.

[0111] The source view received at the input unit 310 may be a view captured by an array of cameras 10. However, this is not necessarily the case, and the source view need not be the same as the view captured by the cameras. Some or all of the source views received at the input unit 310 may be synthesized or otherwise processed source views. The number of source views received at the input unit 310 may be more or less than the number of views captured by the array of cameras 10.

[0112] In Figure 2 the embodiment, non-linear filtering 120 and downsampling 130 are combined in a single step. A "max pooling 2×2" downscaling filter is used. This means that each pixel in the processed depth map takes the maximum pixel value in the 2×2 neighborhood of four pixels in the original input depth map. This choice of non-linear filtering and downsampling stems from the following two insights:

[0113] 1. The result of downscaling should not contain intermediate (i.e., "in-between") depth levels. Such intermediate depth levels are produced when, for example, a linear filter is used. The inventors have recognized that after view synthesis at the decoder side, intermediate depth levels often produce incorrect results.

[0114] 2. Thin foreground objects represented in the depth map should be preserved. Otherwise, for example, relatively thin objects will disappear on the background. Note that it is assumed that the foreground (i.e., nearby objects) is encoded as a high (bright) level and the background (i.e., distant objects) is encoded as a low (dark) level (difference convention). Alternatively, when using the z - coordinate encoding convention (the z - coordinate increases with the distance from the lens), a "min pooling 2×2" downscaler will have the same effect.

[0115] This processing operation effectively increases the size of all local foreground objects and thus preserves small and thin objects. However, the decoder should preferably be aware of which operations have been applied, as it preferably undoes the introduced bias and shrinks all objects to align the depth map with the texture again.

[0116] According to the present embodiment, the memory requirements for the video decoder are reduced. The original pixel rate is: 1Y + 0.5CrCb + 1D, where Y = luminance channel, CrCb = chrominance channel, and D = depth channel. According to this example, by using 4-fold (2×2) downsampling, the pixel rate becomes: 1Y + 0.5CrCb + 0.25D. Thus, a 30% reduction in pixel rate can be achieved. Most practical video decoders are 4:2:0 and do not include a monochrome mode. In this case, a 37.5% pixel reduction is achieved.

[0117] Figure 4 is a flowchart illustrating an encoding method according to a second embodiment. Similar to the Figure 2 method, the method starts with the input section 310 of the video encoder receiving a source view in step 110. In steps 120a and 130a, the video processor 320 processes the depth map according to multiple sets of processing parameters to generate corresponding multiple processed depth maps (each depth map corresponding to one set of processing parameters). In this embodiment, the purpose of the system is to test each of these depth maps to determine which depth map will produce the best quality at the decoder side. Each of the processed depth maps is encoded by the encoder 330 in step 140a. In step 154, the depth decoder 340 decodes each encoded depth map. The decoded depth map is passed to the reconstruction processor 350. In step 156, the reconstruction processor 350 reconstructs the depth map based on the decoded depth map. Then, in step 158, the optimizer 360 compares each reconstructed depth map with the original depth map of the source view to determine the reconstruction error. The reconstruction error quantifies the difference between the original depth map and the reconstructed depth map. Based on the result of the comparison, the optimizer 360 selects the set of parameters that results in the reconstructed image having the minimum reconstruction error. This set of selected parameters is used to generate the video bitstream. The output section 370 outputs the video bitstream corresponding to the selected set of parameters.

[0118] Note that the operations of the depth decoder 340 and the reconstruction processor 350 will be described in more detail below with reference to the decoding method (see Figures 5 - 8 ).

[0119] Effectively, the video encoder 300 iteratively implements the decoder to allow it to predict how the bitstream will be decoded at the remote decoder. The video encoder 300 selects the set of parameters that gives the best performance (in terms of minimizing the reconstruction error for a given target bitrate or pixel rate) at the remote decoder. As Figure 4As shown in the flowchart, this optimization can be iteratively performed, where the parameters of the non-linear filter 120a and / or the downsampling 130a are updated in each iteration after the comparison 158 performed by the optimizer 360. Alternatively, the video decoder can test a fixed number of parameter sets and perform the above operations sequentially or in parallel. For example, in a highly parallel implementation, there can be N encoders (and decoders) in the video encoder 300, each of which is configured to test a set of parameters for encoding the depth map. This can increase the number of parameter sets that can be tested in the available time, but at the cost of an increase in the complexity and / or size of the encoder 300.

[0120] The parameters to be tested can include the parameters of the non-linear filter 120a, the parameters of the downsampling 130a, or both. For example, the system can perform downsampling experiments with various factors in one or two dimensions. Similarly, the system can perform experiments with different non-linear filters. For example, instead of the maximum filter (which assigns the maximum value in the local neighborhood to each pixel), other types of order filters can be used. For example, the non-linear filter can analyze the local neighborhood around a given pixel and assign the second highest value in the neighborhood to the pixel. This can provide a similar effect to the maximum filter while helping to avoid sensitivity to a single outlier. The kernel size of the non-linear filter is another parameter that can be varied.

[0121] Note that the processing parameters at the video decoder can also be included in the parameter set (which will be described in more detail below). In this way, the video encoder can select a set of parameters that contribute to optimizing the quality versus the bitrate / pixel rate for both encoding and decoding. This optimization can be performed for a given scene or a given video sequence, or more generally on a training set of various scenes and video sequences. Therefore, the optimal set of parameters changes for each sequence, each bitrate, and / or each allowed pixel rate.

[0122] The useful or necessary parameters required for the video decoder to correctly decode the video bitstream can be embedded in the metadata bitstream associated with the video bitstream. The metadata bitstream can be sent / transmitted to the video decoder together with the video bitstream, or separately from the video bitstream.

[0123] Figure 5 is a flowchart of a method for decoding video data according to an embodiment. Figure 6is a block diagram of a corresponding video decoder 400. The video decoder 400 includes an input unit 410, a texture decoder 424, a depth decoder 426, a reconstruction processor 450, and an output unit 470. The input unit 410 is coupled to the texture decoder 424 and the depth decoder 426. The reconstruction processor 450 is arranged to receive the decoded texture map from the texture decoder 424 and the decoded depth map from the depth decoder 426. The reconstruction processor 450 is arranged to provide the reconstructed depth map to the output unit 470.

[0124] Figure 5 The method of starts at step 210, where the input unit 410 receives a video bitstream and optionally a metadata bitstream. In step 224, the texture decoder 424 decodes the texture map from the video bitstream. In step 226, the depth decoder 426 decodes the depth map from the video bitstream. In steps 230 and 240, the reconstruction processor 450 processes the decoded depth map to generate a reconstructed depth map. The processing includes upsampling 230 and non-linear filtering 240. This processing (especially non-linear filtering 240) can also depend on the content of the decoded texture map, which will be described in more detail below.

[0125] Now reference will be made to Figure 8 to describe in more detail Figure 5 an example of the method of . In this embodiment, the upsampling 230 includes nearest neighbor upsampling, where each pixel in a 2×2 pixel block in the upsampled depth map is assigned the value of one of the pixels from the decoded depth map. This "nearest neighbor 2×2" upscaler scales the depth map to its original size. Just like the max pooling operation at the encoder, this process at the decoder avoids generating intermediate depth levels. The characteristics of the upscaled depth map can be predicted in advance compared to the original depth map at the encoder: the "max pooling" downscaling filter tends to expand the area of foreground objects. Therefore, some depth pixels in the upsampled depth map are foreground pixels but should be background pixels instead. However, there are usually no background depth pixels that should be foreground pixels. In other words, after upscaling, the objects are sometimes too large, but usually not too small.

[0126] In the present embodiment, in order to undo the bias (foreground objects with increased size), the non-linear filtering 240 of the upscaled depth map includes color adaptation, conditional filtering, attenuation filtering ( Figure 8Steps 242, 244, and 240a) in []. The weakening part (min operator) ensures that the size of the object is reduced, while color adaptation ensures that the depth edges end at the correct spatial positions - that is, where the transitions in the full-scale texture map indicate the edges should be. Due to the non-linear way in which the weakening filtering acts (i.e., whether a pixel is weakened), the resulting object edges will be noisy. Adjacent edge pixels can give different results in the "whether to weaken" classification for different minimized inputs. This noise has an adverse effect on the object edge smoothness. The inventors have recognized that this smoothness is an important requirement for a view synthesis result with sufficient perceptual quality. Therefore, the non-linear filtering 240 also includes contour smoothness filtering (step 250) to smooth the edges in the depth map.

[0127] Non-linear filtering 240 according to this embodiment will now be described in more detail. Figure 7 A small magnified area of the upsampled depth map representing the filter kernel before non-linear filtering 240 is shown. The grey squares indicate foreground pixels; the black squares indicate background pixels. The peripheral pixels of the foreground object are marked as X. These pixels can represent the expanded / enlarged area of the foreground object caused by non-linear filtering at the encoder. In other words, there is uncertainty as to whether the peripheral pixels X are true foreground pixels or background pixels.

[0128] The steps taken to perform adaptive weakening are:

[0129] 1. Find local foreground edges - that is, the peripheral pixels of the foreground object (marked as X in []. This can be done by applying a local threshold to distinguish foreground pixels from background pixels. The peripheral pixels are then identified as those foreground pixels adjacent to the background pixels (in the sense of 4-connected in this example). This is done by the reconstruction processor 450 in step 242. The depth map can (for efficiency) contain wrapped regions from multiple camera views. Edges on the boundaries of such regions are ignored as they do not indicate object edges. Figure 7

[0130] Figure 7 2. For the identified edge pixels (e.g., the central pixel in the 5×5 kernel in []. Determine the average foreground texture color and the average background texture color in the 5×5 kernel. This is done only based on "confident" pixels (marked with a dot ●) - in other words, the calculation of the average foreground texture and the average background texture excludes the uncertain edge pixels X. They also exclude pixels from potentially adjacent patch regions from other camera views, for example.

[0131]

[0132] 3. Determine the similarity to the foreground - that is, foreground confidence:

[0132]

[0133] Wherein: D indicates the (e.g., Euclidean) color distance between the color of the central pixel and the average color of the background pixel or the foreground pixel. If the central pixel is relatively more similar to the average foreground color in the neighborhood, the confidence will be close to 1. If the central pixel is relatively more similar to the average background color in the neighborhood, the confidence will be close to zero. In step 244, the reconstruction processor 450 determines the similarity of the identified peripheral pixels to the foreground.

[0134] 4. Mark all peripheral pixels with C 前景 <threshold (e.g., 0.5) as X.

[0135] 5. Weaken all marked pixels - that is, take the minimum value in the local (e.g., 3×3) neighborhood. In step 240a, the reconstruction processor 450 applies this non-linear filtering to the marked peripheral pixels (whose similarity to the background is greater than its similarity to the foreground).

[0136] As mentioned above, this process will be noisy and will result in jagged edges in the depth map. The steps taken to smooth the object edges represented in the depth map are:

[0137] 1. Find local foreground edges - that is, the peripheral pixels of the foreground object (such as those pixels marked as X in Figure 7 .

[0138] 2. For these edge pixels (e.g., the central pixel in Figure 7 ), count the number of background pixels in the 3×3 kernel around the pixel of interest.

[0139] 3. Mark all edge pixels with count > threshold.

[0140] 4. Weaken all marked pixels - that is, take the minimum value in the local (e.g., 3×3) neighborhood. This step is performed by the reconstruction processor 450 in step 250.

[0141] This smoothing process will tend to convert outlier or prominent foreground pixels into background pixels.

[0142] In the above example, the method uses the number of background pixels in a 3×3 kernel to identify whether a given pixel is an outlier peripheral pixel projected from a foreground object. Other methods can also be used. For example, as an alternative or complement to counting the number of pixels, the positions of foreground and background pixels in the kernel can also be analyzed. If all the background pixels are on one side of the pixel under discussion, then that background pixel is more likely to be a foreground pixel. On the other hand, if the background pixels are all scattered around the pixel under discussion, then that pixel may be an outlier or noise and is more likely to be a true background pixel.

[0143] Pixels in the kernel can be classified as foreground or background in a binary manner. Each pixel is encoded with a binary flag, where a logical "1" indicates background and a logical "0" indicates foreground. Then the neighborhood (i.e., the pixels in the kernel) can be described by an n-bit binary number, where n is the number of pixels in the kernel surrounding the pixel of interest. An exemplary method of constructing the binary number is shown in the following table:

[0144] <![CDATA[b 7 = 1]]> <![CDATA[b 6 = 0]]> <![CDATA[b 5 = 1]]> <![CDATA[b 4 = 0]]> <![CDATA[b 3 = 0]]> <![CDATA[b 2 = 1]]> <![CDATA[b 1 = 0]]> <![CDATA[b 0 = 1]]>

[0145] In this example, b = b 7 b 6 b 5 b 4 b 3 b 2 b 1 b 0 = 10100101 2 = 165. (Note that the algorithm described above with reference to Figure 5 corresponds to counting the number of non-zero bits in b (= 4).

[0146] Training involves counting the frequency of whether the pixel of interest (the central pixel of the kernel) is foreground or background for each value of b. Assuming that the costs for false alarms and misses are equal, then if the pixel (in the training set) is more likely to be a foreground pixel than a background pixel, it is determined to be a foreground pixel, and vice versa.

[0147] The implementation of the decoder will construct b and retrieve the answer (the pixel of interest is foreground or the pixel of interest is background) from a look-up table (LUT).

[0148] The method of non - linear filtering of depth maps (e.g., enhancement and attenuation as described above) at both the encoder and the decoder is counter - intuitive because it is generally expected to remove information from the depth map. However, the inventors have surprisingly found that for a given bitrate, smaller depth maps generated by non - linear downsampling methods can be encoded with higher quality (using a conventional video codec). This quality gain exceeds the loss in reconstruction; thus, the net effect is an improvement in end - to - end quality while reducing the pixel rate.

[0149] As referenced above Figure 3 and Figure 4 described, a decoder can be implemented inside the video encoder to optimize the parameters of non - linear filtering and downsampling, thereby reducing reconstruction errors. In this case, the depth decoder 340 in the video encoder 300 is substantially the same as the depth decoder 426 in the video decoder 400; and the reconstruction processor 350 at the video encoder 300 is substantially the same as the reconstruction processor 450 at the video decoder 400. These corresponding components perform substantially the same processes.

[0150] As described above, when the parameters of non - linear filtering and downsampling at the video encoder have been selected to reduce reconstruction errors, the selected parameters can be signaled in a metadata bitstream that is input to the video decoder. The reconstruction processor 450 can use the parameters signaled in the metadata bitstream to assist in the correct reconstruction of the depth map. The parameters of the reconstruction process can include, but are not limited to, the upsampling factor in one or two dimensions, the kernel size for peripheral pixels identifying foreground objects, the kernel size for attenuation; the type of non - linear filtering to be applied (e.g., whether to use a minimum filter or other types of filters), the kernel size for identifying foreground pixels for smoothing, and the kernel size for smoothing.

[0151] An alternative embodiment will now be described with reference to Figure 9 In this embodiment, instead of using hand - coded non - linear filters for the encoder and decoder, a neural network architecture is used. The neural network is split to model the depth downscaling operation and the depth upscaling operation. The network is trained end - to - end and learns how to downscale optimally and upscale optimally. However, during deployment (i.e., encoding and decoding of real sequences), the first part is before the video encoder, and the second part is after the video decoder. Thus, the first part provides non - linear filtering 120 for the encoding method; and the second part provides non - linear filtering 240 for the decoding method.

[0152] The network parameters (weights) of the second part of the network can be transmitted as metadata using a bitstream. Note that different sets of neural network parameters may be created using different encoding configurations (different downscaling factors, different target bitrates, etc.). This means that for a given bitrate of the texture map, the upscaling filter for the depth map will work optimally. This can improve performance because texture coding artifacts change the luminance and chrominance characteristics, and especially at object boundaries, this change will cause different weights in the depth upscaling neural network.

[0153] Figure 9 An example architecture for this embodiment is shown, where the neural network is a convolutional neural network (CNN). The symbols in the figure have the following meanings:

[0154] I = Input 3-channel full-resolution texture map

[0155] = Decoded full-resolution texture map

[0156] D = Input 1-channel full-resolution depth map

[0157] D down = Downscaled depth map

[0158] = Downscaled decoded depth map

[0159] C k = Convolution with a k×k kernel

[0160] P k = Downscaling by a factor of k

[0161] U k = Upscaling by a factor of k

[0162] Each vertical black bar in the figure represents a tensor of input data or intermediate data - in other words, the tensor of input data to the layers of the neural network. The dimensions of each tensor are described by a triple (p, w, h), where w and h are the width and height of the image respectively, and p is the number of planes or channels of the data. Thus, the input texture map has dimensions (3, w, h) - three planes corresponding to the three color channels. The downsampled depth map has dimensions (1, w / 2, h / 2).

[0163] Downscaling P k may include the average of the downscaling by a factor of k, or a max pooling (or min pooling) operation with a kernel size of k. The downscaling average operation may introduce some intermediate values, but the subsequent layers of the neural network can (e.g., based on texture information) solve this problem.

[0164] Note that during the training phase, the decoded depth map is not used Instead, the uncompressed downscaled depth map D is used down . The reason for this is that the training phase of the neural network requires the calculation of derivatives, which is not possible for the non-linear video encoder function. In fact, this approximation may be effective - especially for higher quality (higher bitrate). During the inference phase (i.e., for processing real video data), the uncompressed downscaled depth map D is clearly not available to the video decoder down . Therefore, the decoded depth map is used Also note that the decoded full-resolution texture map is used during both the training phase and the inference phase Calculating derivatives is not required because this is auxiliary information and not data processed by the neural network

[0165] Due to the possible complexity constraints at the client device, the second part of the network (after video decoding) typically consists of only a small number of convolutional layers

[0166] The availability of training data is crucial for using deep learning methods. In this case, this training data is easily obtained. Before video encoding, uncompressed texture images and full-resolution depth maps are used on the input side. The second part of the network uses the decoded texture map and the downscaled depth map (via the first half of the network as input for training), and the error is evaluated with respect to the full-resolution depth map of the ground truth which is also used as input. Thus, essentially, the patches from the high-resolution source depth map act as both input and output for the neural network. Therefore, the network has some aspects of both the autoencoder architecture and the UNet architecture. However, the proposed architecture is not just a combination of these methods. For example, the decoded texture map enters the second part of the network in the form of auxiliary data to optimally reconstruct the high-resolution depth map

[0167] In Figure 9 the example shown, the inputs to the neural network at the video encoder 300 include the texture map I and the depth map D. The downsampling P 2 is performed between two other layers of the neural network. There are three neural network layers before the downsampling and two layers after the downsampling. The output of the part of the neural network at the video encoder 300 includes the downscaled depth map D down . This is encoded by the encoder 320 in step 140

[0168] The encoded depth map is transmitted in the video bitstream to the video decoder 400. The encoded depth map is decoded by the depth decoder 426 in step 226. This results in the downscaled decoded depth map This is the upsampling (U 2 ) to be used in a part of the neural network at the video decoder 400. Another input to this part of the neural network is the decoded full-resolution texture map generated by the texture decoder 424 This second part of the neural network has three layers. It produces a reconstructed estimated result compared with the original depth map D as an output to generate the resulting error e.

[0169] From the above, it will be apparent that the neural network processing can be implemented by the video processor 320 at the video encoder 300 and can be implemented by the reconstruction processor 450 at the video decoder 400. In the example shown, the non-linear filtering 120 and downsampling 130 are performed in an integrated manner by a part of the neural network at the video encoder 300. At the video decoder 400, the upsampling 230 is performed separately before the non-linear filtering 240 performed by the neural network.

[0170] It should be understood that Figure 9 the arrangement of the neural network layers shown is non-limiting and can be changed in other embodiments. In the example, the network produces a 2×2 downsampled depth map. Of course, different scaling factors can also be used.

[0171] In several of the above embodiments, max filtering, max pooling, enhancement or similar operations are mentioned at the encoder. It should be understood that these embodiments assume that the depth is encoded as 1 / d (or other similar inverse relationship), where d is the distance from the camera. In this assumption, high values in the depth map indicate foreground objects and low values in the depth map indicate background. Therefore, by applying max-type or enhancement-type operations, the method tends to expand the foreground objects. The corresponding reverse process at the decoder can be to apply min-type or weakening-type operations.

[0172] Of course, in other embodiments, the depth can be encoded as d or log d (or another variable having a direct correlation with d). This means that foreground objects are represented by low values of d and background is represented by high values of d. In such embodiments, min filtering, min pooling, weakening or similar operations can be performed at the encoder. Again, this will tend to expand the foreground objects (which is the goal). The corresponding reverse process at the decoder can be to apply max-type or enhancement-type operations.

[0173] Figure 2 , Figure 4 , Figure 5 , Figure 8 and Figure 9 the encoding method and decoding method of Figure 3 andFigure 6 The encoder and decoder can be implemented in hardware, software, or a hybrid of both (e.g., implemented as firmware running on a hardware device). To the extent that an embodiment is implemented partially or fully in software, the functional steps illustrated in the process flow diagrams can be performed by a suitably programmed physical computing device (e.g., one or more central processing units (CPUs), graphics processing units (GPUs), or neural network accelerators (NNAs)). Each process and its individual component steps illustrated in the flowcharts can be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including computer program code that, when the program is run on one or more physical computing devices, is configured to cause the one or more physical computing devices to perform the above-described encoding method or decoding method.

[0174] The storage medium can include volatile and non-volatile computer memories such as RAM, PROM, EPROM, and EEPROM. The various storage media can be fixed within the computing device or removable, such that one or more programs stored on the storage medium are loaded into the processor.

[0175] Metadata according to an embodiment can be stored on the storage medium. A bitstream according to an embodiment can be stored on the same storage medium or a different storage medium. It is not necessary to embed the metadata in the bitstream. Similarly, the metadata and / or the bitstream (with metadata in the bitstream or with the metadata separated from the bitstream) can be transmitted as a signal modulated onto an electromagnetic carrier.

[0176] The signal can be defined according to digital communication standards. The carrier can be an optical carrier, a radio frequency wave, a millimeter wave, or a near-field communication wave. It can be wired or wireless.

[0177] To the extent that an embodiment is implemented partially or fully in hardware, the blocks shown in the Figure 3 and Figure 6 block diagrams can be separate physical components, can be the logical breakdown of a single physical component, or can all be implemented in an integrated manner in one physical component. In an implementation, the function of one block shown in the drawings can be split among multiple components, and in an implementation, the functions of multiple blocks shown in the drawings can be combined in a single component. For example, although Figure 6 the texture decoder 424 and the depth decoder 46 are shown as separate components, their functions can be provided by a single unified decoder component.

[0178] Generally, examples of methods for encoding and decoding data, computer programs implementing these methods, and video encoders and decoders are indicated by the following embodiments.

[0179] Example:

[0180] 1. A method for encoding video data including one or more source views, each source view including a texture map and a depth map, the method comprising:

[0181] Receiving (110) the video data;

[0182] Processing the depth map of at least one source view to generate a processed depth map, the processing including:

[0183] Non-linear filtering (120), and

[0184] Downsampling (130); and

[0185] Encoding (140) the processed depth map and the texture map of the at least one source view to generate a video bitstream.

[0186] 2. The method according to embodiment 1, wherein the non-linear filtering includes enlarging the area of at least one foreground object in the depth map.

[0187] 3. The method according to embodiment 1 or embodiment 2, wherein the non-linear filtering includes applying a filter designed using a machine learning algorithm.

[0188] 4. The method according to one of the foregoing embodiments, wherein the non-linear filtering is performed by a neural network including multiple layers, and the downsampling is performed between two of the multiple layers.

[0189] 5. The method according to one of the foregoing embodiments, wherein the method includes processing (120a, 130a) the depth map according to multiple sets of processing parameters to generate corresponding multiple processed depth maps,

[0190] The method further includes:

[0191] Selecting a set of processing parameters that reduces the reconstruction error of the reconstructed depth map after the corresponding processed depth map has been encoded and decoded; and

[0192] Generating a metadata bitstream identifying the selected set of parameters.

[0193] 6. A method for decoding video data including one or more source views, the method comprising:

[0194] The receiver (210) receives a video bitstream including an encoded depth map and an encoded texture map for at least one source view;

[0195] Decode the encoded depth map (226) to produce a decoded depth map;

[0196] Decode the encoded texture map (224) to produce a decoded texture map; and

[0197] Process the decoded depth map to generate a reconstructed depth map, wherein the processing includes:

[0198] Upsampling (230), and

[0199] Non-linear filtering (240).

[0200] 7. The method according to embodiment 6, further comprising: before the step of processing the decoded depth map to generate the reconstructed depth map, detecting that the resolution of the decoded depth map is lower than the resolution of the decoded texture map.

[0201] 8. The method according to embodiment 6 or embodiment 7, wherein the non-linear filtering includes reducing the area of at least one foreground object in the depth map.

[0202] 9. The method according to any one of embodiments 6-8, wherein the processing of the decoded depth map is at least partially based on the decoded texture map.

[0203] 10. The method according to any one of embodiments 6-9, comprising:

[0204] Upsample the decoded depth map (230);

[0205] Identify (242) the peripheral pixels of at least one foreground object in the upsampled depth map;

[0206] Determine (244) whether the peripheral pixels are more similar to the foreground object or the background based on the decoded texture map; and

[0207] Apply non-linear filtering (240a) only to the peripheral pixels determined to be more similar to the background.

[0208] 11. The method according to any one of the foregoing embodiments, wherein the non-linear filtering includes smoothing the edges of at least one foreground object (250).

[0209] 12. The method according to any one of embodiments 6-11, further comprising receiving a metadata bitstream associated with the video bitstream, the metadata bitstream identifying a set of parameters,

[0210] The method further includes processing the decoded depth map according to a set of identified parameters.

[0211] 13. A computer program comprising computer code which, when the program runs on a processing system, causes the processing system to implement an embodiment according to one of Embodiments 1 to 12.

[0212] 14. A video encoder (300) configured to encode video data including one or more source views, each source view including a texture map and a depth map, the video encoder comprising:

[0213] An input unit (310) configured to receive the video data;

[0214] A video processor (320) configured to process the depth map of at least one source view to generate a processed depth map, the processing including:

[0215] Non-linear filtering (120), and

[0216] Downsampling (130);

[0217] An encoder (330) configured to encode the texture map and the processed depth map of the at least one source view to generate a video bitstream; and

[0218] An output unit (360) configured to output the video bitstream.

[0219] 15. A video decoder (400) configured to decode video data including one or more source views, the video decoder comprising:

[0220] A bitstream input unit (410) configured to receive a video bitstream, wherein the video bitstream includes an encoded depth map and an encoded texture map for at least one source view;

[0221] A first decoder (426) configured to decode the encoded depth map according to the video bitstream to produce a decoded depth map;

[0222] A second decoder (424) configured to decode the encoded texture map according to the video bitstream to produce a decoded texture map;

[0223] A reconstruction processor (450) configured to process the decoded depth map to generate a reconstructed depth map, wherein the processing includes:

[0224] Upsampling (230), and

[0225] nonlinear filtering (240), and

[0226] an output unit (470) configured to output the reconstructed depth map.

[0227] Hardware components suitable for use in embodiments of the present invention include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs). One or more blocks may be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuitry for performing other functions.

[0228] More specifically, the present invention is defined by the claims Define .

[0229] Those skilled in the art will be able to understand and realize variations of the disclosed embodiments when practicing the claimed invention by studying the drawings, the disclosure, and the claims. In the claims, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude a plurality. A single processor or other unit may implement the functions of several items recited in the claims. Although certain measures are recited in mutually different dependent claims, this does not indicate that a combination of these measures cannot be used to advantage. If a computer program is discussed above, the computer program may be stored / distributed on a suitable medium, for example, an optical storage medium or a solid state medium provided together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. If the term "adapted to" is used in the claims or the specification, it should be noted that the term "adapted to" is intended to be equivalent to the term "configured to". Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A method for encoding video data including one or more source views, each source view including a texture map and a depth map, the method comprises: receiving (110) the video data; processing the depth map of at least one source view to generate a processed depth map, the processing comprising: performing non - linear filtering (120) on the depth map to generate a non - linearly filtered depth map, and downsampling (130) the non - linearly filtered depth map to generate a processed depth map; and encoding (140) the processed depth map and the texture map of the at least one source view to generate a video bitstream, wherein the non - linear filtering includes enlarging the area of at least one foreground object in the depth map.

2. The method according to claim 1, wherein, the non - linear filtering includes applying a filter designed using a machine learning algorithm.

3. The method according to any one of the preceding claims, wherein, the non - linear filtering is performed by a neural network including multiple layers, and the downsampling is performed between two of the multiple layers.

4. The method according to any one of claims 1 - 2, wherein, the method includes processing (120a, 130a) the depth map according to multiple sets of processing parameters to generate corresponding multiple processed depth maps, wherein the processing parameters include at least one of the following: the definition of the non - linear filtering to be performed, the definition of the downsampling to be performed, and the definition of the processing operation to be performed at the decoder when reconstructing the depth map, the method further includes: selecting, after the corresponding processed depth maps have been encoded and decoded, a set of processing parameters that reduces the reconstruction error of the reconstructed depth map; and generating a metadata bitstream identifying the selected set of parameters.

5. A method for decoding video data including one or more source views, the method comprises: receiving (210) a video bitstream including an encoded depth map and an encoded texture map for at least one source view; decoding (226) the encoded depth map to produce a decoded depth map; decoding (224) the encoded texture map to produce a decoded texture map; and processing the decoded depth map to generate a reconstructed depth map, wherein the processing comprises: upsampling (230) the decoded depth map to generate an upsampled depth map, and performing non - linear filtering (240) on the upsampled depth map to generate a reconstructed depth map, wherein the non - linear filtering includes reducing the area of at least one foreground object in the depth map.

6. The method according to claim 5, further comprises: detecting that the resolution of the decoded depth map is lower than the resolution of the decoded texture map before the step of processing the decoded depth map to generate the reconstructed depth map.

7. The method according to any one of claims 5 - 6, wherein, the processing of the decoded depth map is at least partially based on the decoded texture map.

8. The method according to any one of claims 5-6, comprising: upsampling (230) the decoded depth map; identifying (242) peripheral pixels of at least one foreground object in the upsampled depth map; determining (244) whether the peripheral pixels are more similar to the foreground object or the background based on the decoded texture map; and applying non-linear filtering (240a) only to the peripheral pixels determined to be more similar to the background.

9. The method according to any one of claims 5-6, wherein, the non-linear filtering includes smoothing the edges of at least one foreground object (250).

10. The method according to any one of claims 5-6, further comprising receiving a metadata bitstream associated with the video bitstream, the metadata bitstream identifying a set of parameters including the definition of the non-linear filtering and / or the definition of the upsampling to be performed, the method further comprising processing the decoded depth map according to the identified set of parameters.

11. A computer program product comprising computer code which, when the computer program product is run on a processing system, causes the processing system to implement the method according to any one of claims 1 to 10.

12. A video encoder (300) configured to encode video data including one or more source views, each source view including a texture map and a depth map, the video encoder comprising: an input unit (310) configured to receive the video data; a video processor (320) configured to process the depth map of at least one source view to generate a processed depth map, the processing including: performing non-linear filtering (120) on the depth map to generate a non-linearly filtered depth map, and downsampling (130) the non-linearly filtered depth map to generate a processed depth map; an encoder (330) configured to encode the texture map and the processed depth map of the at least one source view to generate a video bitstream; and an output unit (360) configured to output the video bitstream, wherein the non-linear filtering includes expanding the area of at least one foreground object in the depth map.

13. A video decoder (400) configured to decode video data including one or more source views, the video decoder comprising: a bitstream input unit (410) configured to receive a video bitstream, wherein the video bitstream includes an encoded depth map and an encoded texture map for at least one source view; a first decoder (426) configured to decode the encoded depth map according to the video bitstream to produce a decoded depth map; a second decoder (424) configured to decode the encoded texture map according to the video bitstream to produce a decoded texture map; a reconstruction processor (450) configured to process the decoded depth map to generate a reconstructed depth map, wherein the processing includes: Upsample (230) the decoded depth map to generate an upsampled depth map, and Non-linearly filter (240) the upsampled depth map to generate the reconstructed depth map, and An output unit (470) configured to output the reconstructed depth map, wherein the non-linear filtering includes reducing the area of at least one foreground object in the depth map.

Citation Information

Patent Citations

  • Method for Generating High Resolution Depth Images from Low Resolution Depth Images Using Edge Layers

    US20120269458A1