Encoding scheme for video data
By nonlinear filtering and downsampling of the depth map of immersive video and encoding it in combination with texture maps, the problem of difficulty in effectively reducing the pixel rate without damaging the video quality in the prior art is solved, and more efficient encoding and better video quality are achieved.
Patent Information
- Application Number
- CN202510363876.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-18
- Filing Date
- 2020-12-17
- Publication Date
- 2025-06-06
AI Technical Summary
When encoding immersive videos, it is difficult to effectively reduce the pixel rate without damaging the video quality. Especially when processing three-dimensional scenes recorded by multiple cameras, traditional methods require detailed analysis, which can easily lead to quality reduction.
By nonlinear filtering and downsampling the depth map of the source view, the processed depth map is generated and encoded with the texture map to generate a video code stream. Nonlinear filtering may include enlarging the area of the foreground object to reduce errors introduced by downsampling.
This method can effectively reduce pixel rate while reducing damage to video quality, especially in retaining the details of the foreground object.
Smart Images

Figure CN120111254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video coding. In particular, the present invention relates to methods and apparatus for encoding and decoding immersive videos. Background Art
[0002] Immersive video, also known as six degrees of freedom (6DoF) video, is a three-dimensional (3D) scene video that allows the reconstruction of views of the scene for viewpoints that vary in position and orientation. It represents a further development of three degrees of freedom (3DoF) video, which allows the reconstruction of views for viewpoints with arbitrary orientations but only at fixed points in space. In 3DoF, the degrees of freedom are angles—
[0003] That is, pitch, roll, and yaw. 3DoF video supports head rotation - in other words, the user consuming the video content can look in any direction in the scene, but cannot move to a different position in the scene. 6DoF video supports head rotation, and additionally supports selecting the position in the scene from which to view the scene.
[0004] To generate 6DoF video, multiple cameras are required to record the scene. Each camera generates image data (often referred to as texture data in this context) and corresponding depth data. For each pixel, the depth data represents the depth at which the corresponding image pixel data is observed by a given camera. Each of the multiple cameras provides a corresponding view of the scene. In many applications, it may be impractical or inefficient to transmit all texture data and depth data for all views.
[0005] To reduce the redundancy between views, it has been proposed to prune the views and pack them into "texture atlases" for each frame of the video stream. This approach attempts to reduce or eliminate the overlap between multiple views, thereby improving efficiency. The non-overlapping parts of different views that remain after pruning can be called "tiles". Alvaro Collet et al. describe an example of this approach in "High-quality streamable free-viewpoint video" (ACM Trans. Graphics (SIGGRAPH), 34(4), 2015). Summary of the invention
[0006] It is desirable to improve the quality and coding efficiency of immersive video. As described above, methods for generating texture atlases using pruning (i.e., excluding redundant texture tiles) can help reduce pixel rate. However, pruning views typically requires detailed analysis that is not error-free and may result in reduced quality for the end user. Therefore, a robust and simple method for reducing pixel rate is needed.
[0007] The invention is defined by the claims.
[0008] According to an example of one aspect of the present invention, there is provided a method for encoding video data comprising one or more source views, each source view comprising a texture map and a depth map, the method comprising:
[0009] receiving the video data;
[0010] processing the depth map of at least one source view to generate a processed depth map, the processing comprising:
[0011] Nonlinear filtering, and
[0012] downsampling; and
[0013] The processed depth map and the texture map of the at least one source view are encoded to generate a video code stream.
[0014] Preferably, prior to downsampling, at least part of non-linear filtering is performed.
[0015] The inventors have discovered that nonlinear filtering of the depth map prior to downsampling can help avoid, reduce or mitigate errors introduced by downsampling. In particular, nonlinear filtering can help prevent small or thin foreground objects from partially or completely disappearing from the depth map due to downsampling. It has been found that nonlinear filtering can be superior to linear filtering in this regard because linear filtering introduces intermediate depth values at the boundary between foreground objects and the background. This makes it difficult for the decoder to distinguish object boundaries from large depth gradients.
[0016] The video data may include 6DoF immersive video.
[0017] The non-linear filtering includes enlarging an area of at least one foreground object in the depth map.
[0018] Scaling up foreground objects before downsampling can help ensure that foreground objects are better protected from damage by the downsampling process—in other words, foreground objects are better preserved in the processed depth map.
[0019] A foreground object can be identified as a group of local pixels at a relatively small depth. The background can be identified as pixels at a relatively large depth. The surrounding pixels of the foreground object can be locally distinguished from the background, for example, by applying a threshold to the depth values in the depth map.
[0020] Non-linear filtering may include morphological filtering, in particular grayscale morphological filtering, such as a maximum filter, a minimum filter or other sequential filters. When the depth map contains depth levels with special meanings (e.g., a depth level of "zero" indicates an invalid depth), such depth levels should preferably be considered as foreground despite their actual values. For this reason, these levels are preferentially retained after downsampling. Therefore, their area will also be enlarged.
[0021] The non-linear filtering may include applying filters designed using a machine learning algorithm.
[0022] The machine learning algorithm may be trained to reduce or minimize reconstruction errors of a reconstructed depth map after the processed depth map has been encoded and decoded.
[0023] The trained filters can similarly help preserve foreground objects in the processed (downsampled) depth map.
[0024] The method may further include designing a filter using a machine learning algorithm, wherein the filter is designed to reduce reconstruction errors of a reconstructed depth map after the processed depth map has been encoded and decoded, and wherein the nonlinear filtering includes applying the designed filter.
[0025] Non-linear filtering may include processing by a neural network, and designing of the filter may include training the neural network.
[0026] The nonlinear filtering may be performed by a neural network including a plurality of layers, and the downsampling may be performed between two layers among the plurality of layers.
[0027] Downsampling can be performed by a max-pooling (or min-pooling) layer of a neural network.
[0028] The method may include processing the depth map according to multiple sets of processing parameters to generate a corresponding plurality of processed depth maps, the method further comprising: selecting a set of processing parameters that reduces reconstruction errors of the reconstructed depth maps after the corresponding processed depth maps have been encoded and decoded; and generating a metadata codestream identifying the selected set of parameters.
[0029] This can allow optimizing parameters for a given application or a given video sequence.
[0030] The processing parameters may comprise a definition of the non-linear filtering performed and / or a definition of the down-sampling performed.Alternatively or additionally, the processing parameters may comprise a definition of processing operations to be performed at a decoder when reconstructing the depth map.
[0031] For each set of processing parameters, the method may include: generating a corresponding processed depth map based on the set of processing parameters; encoding the processed depth map to generate an encoded depth map; decoding the encoded depth map; reconstructing the depth map based on the decoded depth map; and comparing the reconstructed depth map with a depth map of at least one source view to determine a reconstruction error.
[0032] According to another aspect, there is provided a method of decoding video data comprising one or more source views, the method comprising:
[0033] Receiving a video stream comprising an encoded depth map and an encoded texture map for at least one source view;
[0034] decoding the encoded depth map to produce a decoded depth map;
[0035] decoding the encoded texture map to produce a decoded texture map; and
[0036] and processing the decoded depth map to generate a reconstructed depth map, wherein the processing comprises:
[0037] Upsampling, and
[0038] Nonlinear filtering.
[0039] The method may further comprise, before the step of processing the decoded depth map to generate the reconstructed depth map, detecting that a resolution of the decoded depth map is lower than a resolution of the decoded texture map.
[0040] In some encoding schemes, depth maps may be downsampled only in certain cases or only for certain views. By comparing the resolution of the decoded depth map with the resolution of the decoded texture map, the decoding method can determine whether to apply downsampling at the encoder. This can avoid the need for metadata in the metadata bitstream to signal which depth maps are downsampled and to what degree. (In this example, it is assumed that the texture map is encoded at full resolution.)
[0041] To generate a reconstructed depth map, the decoded depth map may be upsampled to the same resolution as the decoded texture map.
[0042] Preferably, the non-linear filtering in the decoding method is adapted to compensate for the effects of the non-linear filtering applied in the encoding method.
[0043] The non-linear filtering may include reducing the area of at least one foreground object in the depth map. This may be appropriate when the non-linear filtering during encoding includes increasing the area of at least one foreground object.
[0044] The non-linear filtering may include morphological filtering, in particular grey-scale morphological filtering, such as a maximum filter, a minimum filter or other sequential filters.
[0045] The non-linear filtering during decoding preferably compensates or reverses the effect of the non-linear filtering during encoding. For example, if the non-linear filtering during encoding includes a maximum filter (gray level enhancement), the non-linear filtering during decoding may include a minimum filter (gray level reduction), and vice versa, when the depth map contains depth levels with special meanings (for example, a depth level of "zero" indicates an invalid depth), then such depth levels should preferably be considered as foreground even though they have actual values.
[0046] Preferably, at least part of said non-linear filtering is performed after said upsampling. Optionally, all non-linear filtering is performed after upsampling.
[0047] The processing of the decoded depth map may be based at least in part on the decoded texture map. The inventors have recognised that the texture map contains useful information that aids in reconstructing the depth map. In particular, where boundaries of foreground objects are altered by non-linear filtering during encoding, analysis of the texture map can help to compensate for or reverse the changes.
[0048] The method may include: upsampling the decoded depth map; identifying surrounding pixels of at least one foreground object in the upsampled depth map; determining whether the surrounding pixels are more similar to the foreground object or to the background based on the decoded texture map; and applying non-linear filtering only to surrounding pixels determined to be more similar to the background.
[0049] In this way, the texture map is used to help identify pixels that have transitioned from background to foreground due to nonlinear filtering during encoding. Nonlinear filtering during decoding can help restore these identified pixels to part of the background.
[0050] The non-linear filtering may include smoothing the edges of at least one foreground object.
[0051] The smoothing process may include: identifying peripheral pixels of at least one foreground object in the upsampled depth map; for each peripheral pixel, analyzing the number and / or arrangement of foreground pixels and background pixels in a neighborhood around the peripheral pixel; identifying distant peripheral pixels projected from the object into the background based on the results of the analysis; and applying nonlinear filtering only to the identified peripheral pixels.
[0052] The analysis may include counting the number of background pixels in the neighborhood, wherein if the number of background pixels in the neighborhood is above a predefined threshold, the surrounding pixels are identified as outliers that deviate from the object.
[0053] Alternatively or additionally, the analysis may include identifying spatial patterns of foreground pixels and background pixels in a neighborhood, wherein a surrounding pixel is identified as an outlier if its neighborhood matches one or more predefined spatial patterns.
[0054] The method may further include receiving a metadata stream associated with the video stream, the metadata stream identifying a set of parameters, and the method optionally further includes processing the decoded depth map according to the identified set of parameters.
[0055] The processing parameters may comprise a definition of the non-linear filtering and / or a definition of the upsampling to be performed.
[0056] The non-linear filtering may include applying filters designed using a machine learning algorithm.
[0057] The machine learning algorithm may be trained to reduce or minimize reconstruction errors of a reconstructed depth map after the processed depth map has been encoded and decoded.
[0058] Filters may be defined in a metadata stream associated with a video stream.
[0059] There is also provided a computer program comprising computer code for causing a processing system to carry out the method as outlined above, when said program is run on the processing system.
[0060] The computer program may be stored on a computer readable storage medium. This may be a non-transitory storage medium.
[0061] According to another aspect, there is provided a video encoder configured to encode video data comprising one or more source views, each source view comprising a texture map and a depth map, the video encoder comprising:
[0062] an input unit configured to receive the video data;
[0063] a video processor configured to process the depth map of at least one source view to generate a processed depth map, the processing comprising:
[0064] Nonlinear filtering, and
[0065] Downsampling;
[0066] An encoder configured to encode the texture map and the processed depth map of the at least one source view to generate a video code stream; and
[0067] An output unit is configured to output the video code stream.
[0068] According to yet another aspect, there is provided a video decoder configured to decode video data comprising one or more source views, the video decoder comprising:
[0069] A bitstream input unit, configured to receive a video bitstream, wherein the video bitstream includes an encoded depth map and an encoded texture map for at least one source view;
[0070] A first decoder configured to decode the encoded depth map according to the video code stream to generate a decoded depth map;
[0071] a second decoder configured to decode the encoded texture map according to the video code stream to generate a decoded texture map;
[0072] a reconstruction processor configured to process the decoded depth map to generate a reconstructed depth map, wherein the processing comprises:
[0073] Upsampling, and
[0074] Nonlinear filtering, and
[0075] An output unit is configured to output the reconstructed depth map.
[0076] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] For a better understanding of the invention and to show more clearly how it may be put into practice, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0078] Figure 1 An example of encoding and decoding immersive video using an existing video codec is illustrated;
[0079] Figure 2is a flowchart illustrating a method of encoding video data according to an embodiment;
[0080] Figure 3 is a block diagram of a video encoder according to an embodiment;
[0081] Figure 4 is a flow chart illustrating a method of encoding video data according to further embodiments;
[0082] Figure 5 is a flow chart illustrating a method of decoding video data according to an embodiment;
[0083] Figure 6 is a block diagram of a video decoder according to an embodiment;
[0084] Figure 7 A method for selectively applying nonlinear filtering to specific pixels in a decoding method according to an embodiment is illustrated;
[0085] Figure 8 is a flowchart illustrating a method of decoding video data according to a further embodiment; and
[0086] Fig. 9 Illustrated is a process for encoding and decoding video data using neural network processing according to an embodiment. DETAILED DESCRIPTION
[0087] The present invention will be described with reference to the accompanying drawings.
[0088] It should be understood that although the detailed description and specific examples indicate exemplary embodiments of the device, system and method, the detailed description and specific examples are intended to be used for illustrative purposes only and are not intended to limit the scope of the invention. These and other features, aspects and advantages of the device, system and method of the present invention will be better understood from the following description, claims and drawings. It should be understood that these drawings are merely schematic and are not drawn to scale. It should also be understood that throughout all drawings, the same reference numerals are used to indicate the same or similar parts.
[0089] Methods for encoding and decoding immersive videos are disclosed. In one encoding method, source video data including one or more source views are encoded into a video bitstream. Depth data of at least one of the source views is nonlinearly filtered and downsampled before encoding. Downsampling the depth map helps reduce the amount of data to be transmitted, thereby helping to reduce the bit rate. However, the inventors have found that simple downsampling can cause thin or small foreground objects (e.g., cables) to disappear from the downsampled depth map. Embodiments of the present invention attempt to mitigate this effect and retain small and thin objects in the depth map.
[0090] Embodiments of the present invention may be suitable for use in implementing portions of a technical standard, such as ISO / IEC 23090-12 immersive video for MPEG-I Part 12. Where possible, the terminology used herein is selected to be consistent with that used in MPEG-I Part 12. Nevertheless, it should be understood that the scope of the present invention is neither limited to MPEG-I Part 12 nor to any other technical standard.
[0091] It would be helpful to elaborate on the following definitions / explanations:
[0092] A "3D scene" refers to the visual content in a global reference coordinate system.
[0093] An "atlas" is an aggregated content that aggregates tiles from one or more view representations after a packing process into picture pairs containing a texture component picture and a corresponding depth component picture.
[0094] An "atlas component" is a texture component or a depth component of an atlas.
[0095] The Camera Parameters define the projection used to generate a view representation from a 3D scene.
[0096] “De-cluttering” is the process of identifying and extracting occluded regions across views to obtain patches.
[0097] "Rendering" is an embodiment of the process of creating a viewport or omnidirectional view corresponding to a viewing position and orientation from a 3D scene representation.
[0098] A "source view" is the source video material prior to encoding corresponding to a format of a view representation, which may be acquired by capturing the 3D scene with a real camera, or by projection of a virtual camera onto a surface using source camera parameters.
[0099] A "target view" is defined as a perspective viewport or omnidirectional view at a desired viewing position and orientation.
[0100] A "view representation" consists of a 2D sample array of texture components and corresponding depth components, which represents the projection of the 3D scene onto a surface using the camera parameters.
[0101] A machine learning algorithm is any self-training algorithm that processes input data to generate or predict output data. In some embodiments of the present invention, the input data includes one or more views decoded from a bitstream, and the output data includes a prediction result / reconstruction result of a target view.
[0102] Machine learning algorithms suitable for use in the present invention are apparent to the skilled person. Examples of suitable machine learning algorithms include decision tree algorithms and artificial neural networks. Other machine learning algorithms (e.g., logistic regression, support vector machines, or naive Bayes models) are suitable alternatives.
[0103] The structure of an artificial neural network (or neural network for short) is inspired by the human brain. A neural network consists of multiple layers, each layer including multiple neurons. Each neuron includes a mathematical operation. In particular, each neuron can include a different weighted combination of a single type of transformation (e.g., the same type of transformation, sigmoid, etc., but weighted differently). In the process of processing input data, the mathematical operation of each neuron is performed on the input data to produce a numerical output, and the output of each layer in the neural network is fed into one or more other layers (e.g., sequentially). The last layer provides the output.
[0104] Methods for training machine learning algorithms are well known. Typically, such methods include obtaining a training data set, the training data set comprising training input data entries and corresponding training output data entries. An initialized machine learning algorithm is applied to each input data entry to generate a predicted output data entry. The error between the predicted output data entry and the corresponding training output data entry is used to modify the machine learning algorithm. The process can be repeated until the error converges and the predicted output data entry is sufficiently similar to the training output data entry (e.g., ±1%). This is generally referred to as a supervised learning technique.
[0105] For example, in the case where the machine learning algorithm is formed by a neural network, the mathematical operation (weighting) of each neuron can be modified until the error converges. Known methods of modifying a neural network include gradient descent, back propagation algorithms, and the like.
[0106] Convolutional Neural Networks (CNN or ConvNet) are a type of deep neural network that is most commonly used to analyze visual images. CNN is a regularized version of a multi-layer perceptron.
[0107] Figure 1A system for encoding and decoding immersive video is illustrated in simplified form. An array of cameras 10 is used to capture multiple views of a scene. Each camera captures a conventional image (referred to herein as a texture map) and a depth map of the view in front of the texture map. A set of views including texture data and depth data is provided to an encoder 300. The encoder encodes both texture data and depth data into a conventional video stream - in this case, a high-efficiency video coding (HEVC) stream. This is accompanied by a metadata stream to inform the decoder 400 of the meaning of different parts of the video stream. For example, metadata tells the decoder which parts of the video stream correspond to the texture map and which parts of the video stream correspond to the depth map. Depending on the complexity and flexibility of the encoding scheme, more or less metadata may be required. For example, a very simple scheme may very tightly specify the structure of the stream so that little or no metadata is required for unpacking the stream at the decoder end. In the case where the stream has more optional possibilities, a larger amount of metadata will be required.
[0108] The decoder 400 decodes the coded (texture and depth) view. The decoder 400 passes the decoded view to the compositor 500. The compositor 500 is coupled to a display device, such as a virtual reality head mounted device 550. The head mounted device 550 requests the compositor 500 to use the decoded view to synthesize and render a specific view of the 3D scene according to the current position and orientation of the head mounted device 550.
[0109] Figure 1 The advantage of the system shown is that it can use conventional 2D video codecs to encode and decode texture data and depth data. However, a disadvantage is that a large amount of data needs to be encoded, transmitted and decoded. Therefore, it is desirable to reduce the data rate while compromising the quality of the reconstructed view as little as possible.
[0110] Figure 2 The encoding method according to the first embodiment is illustrated. Figure 3 The diagram shows a Figure 2 The video encoder of the method. The video encoder includes: an input unit 310, which is configured to receive video data; a video processor 320, which is coupled to the input unit and configured to receive a depth map received by the input unit; an encoder 330, which is arranged to receive a processed depth map from the video processor 320; an output unit 370, which is arranged to receive and output a video code stream generated by the encoder 330. The video encoder 300 also includes a depth decoder 340, a reconstruction processor 350 and an optimizer 360. The following reference will be made to Figure 4 A second embodiment of the encoding method is described to describe these components in more detail.
[0111] refer to Figure 2 and Figure 3 The method of the first embodiment starts at step 110, where the input unit 310 receives video data including a texture map and a depth map. In steps 120 and 130, the video processor 320 processes the depth map to generate a processed depth map. This processing includes nonlinear filtering of the depth map in step 120 and downsampling of the filtered depth map in step 130. In step 140, the encoder 330 encodes the processed depth map and the texture map to generate a video stream. The generated video stream is then output via the output unit 370.
[0112] The source views received at the input 310 may be views captured by the array of cameras 10. However, this is not required, and the source views need not be the same as the views captured by the cameras. Some or all of the source views received at the input 310 may be synthesized or otherwise processed source views. The number of source views received at the input 310 may be greater or less than the number of views captured by the array of cameras 10.
[0113] exist Figure 2 In an embodiment, nonlinear filtering 120 is combined with downsampling 130 in a single step. The filter is downscaled using "max pooling 2×2". This means that each pixel in the processed depth map takes the maximum pixel value in a 2×2 neighborhood of four pixels in the original input depth map. This choice of nonlinear filtering and downsampling stems from the following two insights:
[0114] 1. The down-scaling result should not contain intermediate (ie "in between") depth levels. Such intermediate depth levels are generated when, for example, a linear filter is used. The inventors have realized that intermediate depth levels often produce erroneous results after view synthesis at the decoder side.
[0115] 2. Thin foreground objects represented in the depth map should be preserved. Otherwise, for example, relatively thin objects will disappear on the background. Note that it is assumed that the foreground (i.e., nearby objects) are encoded as high (bright) levels and the background (i.e., distant objects) are encoded as low (dark) levels (difference convention). Alternatively, the "min-pooling 2×2" downscaler will have the same effect when using the z-coordinate encoding convention (z-coordinate increases with distance from the camera).
[0116] This processing operation effectively increases the size of all local foreground objects and thus keeps objects small and thin. However, the decoder should preferably be aware of which operations are applied, as it preferably undoes the introduced bias and shrinks all objects to align the depth map with the texture again.
[0117] According to the current embodiment, the memory requirements for the video decoder are reduced. The original pixel rate is: 1Y+0.5CrCb+1D, where y=luminance channel, CrCb=chrominance channel, and D=depth channel. According to this example, by using 4 times (2×2) downsampling, the pixel rate becomes: 1Y+0.5CrCb+0.25D. Therefore, a 30% pixel rate reduction can be achieved. Most actual video decoders are 4:2:0 and do not include a monochrome mode. In this case, a 37.5% pixel reduction is achieved.
[0118] Figure 4 is a flowchart illustrating an encoding method according to the second embodiment. Figure 2 The method begins with receiving a source view at the input portion 310 of the video encoder in step 110. In steps 120a and 130a, the video processor 320 processes the depth map according to multiple sets of processing parameters to generate a corresponding plurality of processed depth maps (each depth map corresponding to a set of processing parameters). In this embodiment, the purpose of the system is to test each of these depth maps to determine which depth map will produce the best quality at the decoder end. Each of the processed depth maps is encoded by the encoder 330 in step 140a. In step 154, the depth decoder 340 decodes each encoded depth map. The decoded depth map is passed to the reconstruction processor 350. In step 156, the reconstruction processor 350 reconstructs the depth map based on the decoded depth map. Then, in step 158, the optimizer 360 compares each reconstructed depth map with the original depth map of the source view to determine the reconstruction error. The reconstruction error quantifies the difference between the original depth map and the reconstructed depth map. Based on the comparison result, the optimizer 360 selects a set of parameters that makes the reconstructed image have the minimum reconstruction error. This set of parameters is selected for generating a video bitstream. The output unit 370 outputs the video bitstream corresponding to the selected set of parameters.
[0119] Note that reference will be made to the decoding method below (see Figure 5-8 ) describes the operation of the depth decoder 340 and the reconstruction processor 350 in more detail.
[0120] Effectively, the video encoder 300 loops through the decoder to allow it to predict how the bitstream will be decoded at the far-end decoder. The video encoder 300 selects the set of parameters that gives the best performance (in terms of minimizing reconstruction errors for a given target bit rate or pixel rate) at the far-end decoder. Figure 4As shown in the flowchart of , the optimization can be performed iteratively, wherein the parameters of the nonlinear filtering 120a and / or the downsampling 130a are updated in each iteration after the comparison 158 performed by the optimizer 360. Alternatively, the video decoder can test a fixed plurality of parameter sets, and the above operations can be performed sequentially or in parallel. For example, in a highly parallel embodiment, there can be N encoders (and decoders) in the video encoder 300, each of which is configured to test a set of parameters for encoding a depth map. This can increase the number of parameter sets that can be tested in the available time, but at the expense of an increase in the complexity and / or size of the encoder 300.
[0121] The parameters tested may include parameters of nonlinear filtering 120a, parameters of downsampling 130a, or both. For example, the system may experiment with downsampling by various factors in one or two dimensions. Similarly, the system may experiment with different nonlinear filters. For example, instead of a maximum filter (which assigns the maximum value in a local neighborhood to each pixel), other types of sequential filters may be used. For example, a nonlinear filter may analyze a local neighborhood around a given pixel and may assign the second highest value in the neighborhood to a pixel. This may provide an effect similar to a maximum filter while helping to avoid sensitivity to a single outlier. The kernel size of the nonlinear filter is another parameter that may be varied.
[0122] Note that processing parameters at the video decoder may also be included in the parameter set (described in more detail below). In this way, the video encoder may select a set of parameters that helps optimize quality versus bitrate / pixel rate for both encoding and decoding. This optimization may be performed for a given scene or a given video sequence, or more generally on a training set of various scenes and video sequences. Thus, the optimal set of parameters may change for each sequence, each bitrate, and / or each allowed pixel rate.
[0123] Useful parameters or necessary parameters required by a video decoder to correctly decode a video bitstream can be embedded in a metadata bitstream associated with the video bitstream. The metadata bitstream can be sent / transmitted to the video decoder together with the video bitstream, or can be sent / transmitted to the video decoder separately from the video bitstream.
[0124] Figure 5 is a flowchart of a method of decoding video data according to an embodiment. Figure 64 is a block diagram of a corresponding video decoder 400. The video decoder 400 comprises an input 410, a texture decoder 424, a depth decoder 426, a reconstruction processor 450, and an output 470. The input 410 is coupled to the texture decoder 424 and the depth decoder 426. The reconstruction processor 450 is arranged to receive a decoded texture map from the texture decoder 424 and a decoded depth map from the depth decoder 426. The reconstruction processor 450 is arranged to provide a reconstructed depth map to the output 470.
[0125] Figure 5 The method begins at step 210, where an input unit 410 receives a video stream and optionally a metadata stream. In step 224, a texture decoder 424 decodes a texture map from the video stream. In step 226, a depth decoder 426 decodes a depth map from the video stream. In steps 230 and 240, a reconstruction processor 450 processes the decoded depth map to generate a reconstructed depth map. The processing includes upsampling 230 and non-linear filtering 240. The processing (particularly the non-linear filtering 240) may also depend on the content of the decoded texture map, which will be described in more detail below.
[0126] Now refer to Figure 8 To describe in more detail Figure 5 An example of a method of upsampling 230. In this embodiment, upsampling 230 includes nearest neighbor upsampling, where each pixel in a block of 2×2 pixels in the upsampled depth map is assigned the value of one of the pixels from the decoded depth map. This "nearest neighbor 2×2" upscaler scales the depth map to its original size. Just like the max pooling operation at the encoder, this process at the decoder avoids the generation of intermediate depth levels. The characteristics of the upscaled depth map compared to the original depth map at the encoder can be predicted in advance: the "max pooling" downscaling filter tends to enlarge the area of foreground objects. Therefore, some depth pixels in the upsampled depth map are foreground pixels, but should be changed to background pixels. However, there are usually no background depth pixels that should be changed to foreground pixels. In other words, after upscaling, objects are sometimes too large, but usually not too small.
[0127] In the current embodiment, to undo the bias (foreground object growth), nonlinear filtering 240 of the upscaled depth map includes color adaptation, conditional filtering, attenuation filtering ( Figure 8242, 244 and 240a in ). The attenuation part (minimum operator) ensures that the size of the object is reduced, while the color adaptation ensures that the depth edge ends up at the correct spatial location - that is, the transition in the full-scale texture map indicates where the edge should be. Due to the non-linear way in which the attenuation filter works (that is, whether a pixel is attenuated or not), the resulting object edges will be noisy. Neighboring edge pixels can give different results in the "attenuated or not" classification for different inputs to the minimization. This noise has an adverse effect on the smoothness of the object edges. The inventors have recognized that such smoothness is an important requirement for view synthesis results of sufficient perceptual quality. Therefore, the non-linear filtering 240 also includes contour smoothness filtering (step 250) to smooth the edges in the depth map.
[0128] The non-linear filtering 240 according to the present embodiment will now be described in more detail. Figure 7 A small magnified area of an upsampled depth map representing the filter kernel prior to nonlinear filtering 240 is shown. Grey squares indicate foreground pixels; black squares indicate background pixels. The surrounding pixels of the foreground object are marked as X. These pixels may represent an extended / magnified area of the foreground object caused by the nonlinear filtering at the encoder. In other words, there is uncertainty as to whether the surrounding pixel X is a true foreground pixel or a background pixel.
[0129] The steps taken to perform adaptive attenuation are:
[0130] 1. Find local foreground edges - that is, the surrounding pixels of the foreground object (in Figure 7 4 (marked with an X in the figure). This can be done by applying a local threshold to distinguish foreground pixels from background pixels. The surrounding pixels are then identified as those foreground pixels that are adjacent to the background pixels (in this example, in a 4-connected sense). This is done by the reconstruction processor 450 in step 242. The depth map may (for efficiency) contain packed regions from multiple camera views. Edges on the borders of such regions are ignored as these do not indicate object edges.
[0131] 2. For the identified edge pixels (e.g. Figure 7 The average foreground texture color and the average background texture color in the 5×5 kernel are determined based on the center pixel in the 5×5 kernel in the image. This is done based only on the "confident" pixels (marked with dots ●) - in other words, the calculations of the average foreground texture and the average background texture exclude the uncertain edge pixels X. They also exclude pixels from possible neighboring patch areas that may apply, for example, other camera views.
[0132] 3. Determine similarity with the foreground - i.e., foreground confidence:
[0133]
[0134] Where: D indicates the (e.g., Euclidean) color distance between the color of the center pixel and the average color of the background pixels or foreground pixels. If the center pixel is relatively more similar to the average foreground color in the neighborhood, the confidence will be close to 1. If the center pixel is relatively more similar to the average background color in the neighborhood, the confidence will be close to zero. In step 244, the reconstruction processor 450 determines the similarity of the identified surrounding pixels to the foreground.
[0135] 4. C 前景 All surrounding pixels < a threshold value (eg, 0.5) are marked with an X.
[0136] 5. Attenuate all marked pixels - ie, take the minimum in a local (eg, 3x3) neighborhood. In step 240a, the reconstruction processor 450 applies the nonlinear filter to the marked surrounding pixels (which are more similar to the background than to the foreground).
[0137] As mentioned above, this process can be noisy and can result in jagged edges in the depth map. The steps taken to smooth the edges of objects represented in the depth map are:
[0138] 1. Find local foreground edges - that is, the surrounding pixels of the foreground object (such as Figure 7 Those pixels marked with X in ).
[0139] 2. For these edge pixels (e.g. Figure 7 ), count the number of background pixels in a 3×3 kernel around the pixel of interest.
[0140] 3. Mark all edge pixels with count > threshold.
[0141] 4. Attenuate all marked pixels - that is, take the minimum value in a local (eg 3×3) neighborhood. This step is performed by the reconstruction processor 450 in step 250.
[0142] This smoothing process will tend to convert outlier or prominent foreground pixels into background pixels.
[0143] In the example above, the method uses the number of background pixels in a 3×3 kernel to identify whether a given pixel is an outlier surrounding pixel projected from a foreground object. Other methods can also be used. For example, as an alternative or in addition to counting the number of pixels, the location of the foreground and background pixels in the kernel can also be analyzed. If the background pixels are all on one side of the pixel in question, then the background pixel is more likely to be a foreground pixel. On the other hand, if the background pixels are all scattered around the pixel in question, then the pixel is likely an outlier or noise, and is more likely to be a true background pixel.
[0144] The pixels in the kernel can be classified as foreground or background in a binary manner. A binary flag is encoded for each pixel, where a logical "1" indicates background and a logical "0" indicates foreground. The neighborhood (i.e., the pixels in the kernel) can then be described by an n-bit binary number, where n is the number of pixels in the kernel surrounding the pixel of interest. An exemplary method of constructing the binary number is shown in the following table:
[0145] <![CDATA[b 7 =1]]> <![CDATA[b 6 =0]]> <![CDATA[b 5 =1]]> <![CDATA[b 4 =0]]> <![CDATA[b 3 =0]]> <![CDATA[b 2 =1]]> <![CDATA[b 1 =0]]> <![CDATA[b 0 =1]]>
[0146] In this example, b=b 7 b 6 b 5 b 4 b 3 b 2 b 1 b 0 =10100101 2 =165. (Note that the above reference Figure 5 The algorithm described corresponds to counting the number of non-zero bits in b (=4).
[0147] Training consists of counting how often the pixel of interest (the center pixel of the kernel) is foreground or background for each value of b. Assuming equal costs for false alarms and misses, a pixel (in the training set) is determined to be a foreground pixel if it is more likely to be a foreground pixel than a background pixel, and vice versa.
[0148] An implementation of the decoder will construct b and retrieve the answer (either the pixel of interest is foreground or the pixel of interest is background) from a lookup table (LUT).
[0149] The approach of nonlinear filtering of depth maps at both the encoder and decoder (e.g., enhancement and attenuation, respectively, as described above) is counterintuitive, as it is generally expected to remove information from the depth map. However, the inventors have surprisingly discovered that for a given bit rate, smaller depth maps produced by nonlinear downsampling methods can be encoded at higher quality (using conventional video codecs). This quality gain outweighs the loss in reconstruction; thus, the net effect is improved end-to-end quality while reducing the pixel rate.
[0150] As referenced above Figure 3 and Figure 4 As described, a decoder may be implemented within a video encoder to optimize the parameters of nonlinear filtering and downsampling to reduce reconstruction errors. In this case, the depth decoder 340 in the video encoder 300 is substantially the same as the depth decoder 426 in the video decoder 400; and the reconstruction processor 350 at the video encoder 300 is substantially the same as the reconstruction processor 450 at the video decoder 400. These corresponding components perform substantially the same process.
[0151] As described above, when parameters for nonlinear filtering and downsampling at the video encoder have been selected to reduce reconstruction errors, the selected parameters can be signaled in the metadata stream that is input to the video decoder. The reconstruction processor 450 can use the parameters signaled in the metadata stream to assist in correctly reconstructing the depth map. The parameters of the reconstruction process may include, but are not limited to, upsampling factors in one or two dimensions, kernel sizes for identifying surrounding pixels of foreground objects, kernel sizes for attenuation; the type of nonlinear filtering to be applied (e.g., whether to use a minimum filter or other type of filter), the kernel size for identifying foreground pixels for smoothing, and the kernel size for smoothing.
[0152] Now refer to Fig. 9 to describe an alternative embodiment. In this embodiment, instead of using hand-coded nonlinear filters for the encoder and decoder, a neural network architecture is used. The neural network is split to model the deep downscaling operation and the deep upscaling operation. The network is trained end-to-end and learns how to optimally downscale and optimally upscale. However, during deployment (i.e., encoding and decoding of real sequences), the first part is before the video encoder and the second part is after the video decoder. Therefore, the first part provides nonlinear filtering 120 for the encoding method; and the second part provides nonlinear filtering 240 for the decoding method.
[0153] The network parameters (weights) of the second part of the network can be transmitted as metadata with the bitstream. Note that different sets of neural network parameters may be created with different encoding configurations (different downscaling factors, different target bitrates, etc.) correspondingly. This means that for a given bitrate of the texture map, the upscaling filter for the depth map will work in an optimal way. This can improve performance because texture coding artifacts change the luminance and chrominance characteristics, and especially at object boundaries, such changes will cause different weights of the depth upscaling neural network.
[0154] Fig. 9 An example architecture for this embodiment is shown, where the neural network is a convolutional neural network (CNN). The symbols in the figure have the following meanings:
[0155] I = Input 3-channel full-resolution texture
[0156]
[0157] D = Input 1 channel full resolution depth map
[0158] D down = Down-scaled depth map
[0159]
[0160] C k = convolution with k×k kernel
[0161] P k = scaling down by factor k
[0162] U k = scaling by factor k
[0163] Each vertical black bar in the figure represents a tensor of input data or intermediate data - in other words, a tensor of input data to a layer of a neural network. The dimensions of each tensor are described by a triple (p, w, h), where w and h are the width and height of the image, respectively, and p is the number of planes or channels of data. Thus, the input texture map has dimensions (3, w, h) - three planes corresponding to the three color channels. The downsampled depth map has dimensions (1, w / 2, h / 2).
[0164] Downscaled P k This may include a down-scaled average by a factor k, or a max-pooling (or min-pooling) operation with a kernel size k. The down-scaled average operation may introduce some intermediate values, but later layers of the neural network may resolve this (e.g., based on texture information).
[0165] Note that during the training phase, the decoded depth map is not used. Instead, the uncompressed downscaled depth map D is used down The reason for this is that the training phase of the neural network requires the computation of derivatives, which is not possible for nonlinear video encoder functions. In practice, this approximation may be valid - especially for higher quality (higher bitrates). In the inference phase (i.e., for processing real video data), the uncompressed downscaled depth map D is obviously not available to the video decoder. down Therefore, using the decoded depth map Also note that the decoded full-resolution texture maps are used during both training and inference. There is no need to calculate derivatives, since this is auxiliary information and not data processed by the neural network.
[0166] Due to complexity constraints that may exist at the client device, the second part of the network (after video decoding) typically contains only a few convolutional layers.
[0167] Crucial to the use of deep learning methods is the availability of training data. In this case, these training data are easily available. Uncompressed texture images and full-resolution depth maps are used on the input side before video encoding. The second part of the network uses the decoded texture map and the downscaled depth map (via the first half of the network as input for training) and evaluates the error relative to the full-resolution depth map of the real situation, which is also used as input. So, in essence, the patches from the high-resolution source depth map serve as both input and output for the neural network. Therefore, the network has some aspects of both the autoencoder architecture and the UNet architecture. However, the proposed architecture is not just a combination of these methods. For example, the decoded texture map enters the second part of the network in the form of auxiliary data to optimally reconstruct the high-resolution depth map.
[0168] exist Fig. 9 In the example shown, the input to the neural network at the video encoder 300 includes a texture map I and a depth map D. 2 is performed between the other two layers of the neural network. There are three neural network layers before downsampling and two layers after downsampling. The output of the portion of the neural network at the video encoder 300 includes the downsampled depth map D down This is encoded by encoder 320 in step 140 .
[0169] The encoded depth map is transmitted in the video bitstream to the video decoder 400. The encoded depth map is decoded by the depth decoder 426 in step 226. This produces a downscaled decoded depth map This is the upsampling (U) to be used in the portion of the neural network at the video decoder 400. 2 ). Another input to this part of the neural network is the decoded full-resolution texture map generated by the texture decoder 424. This second part of the neural network has three layers. It produces a reconstructed estimate that is compared to the original depth map D As output, to produce the resulting error e.
[0170] As will be apparent from the above, neural network processing may be implemented at the video encoder 300 by the video processor 320, and may be implemented at the video decoder 400 by the reconstruction processor 450. In the example shown, non-linear filtering 120 and downsampling 130 are performed in an integrated manner by portions of the neural network at the video encoder 300. At the video decoder 400, upsampling 230 is performed separately prior to the non-linear filtering 240 performed by the neural network.
[0171] It should be understood that Fig. 9 The arrangement of the neural network layers shown is non-limiting and may vary in other embodiments. In the example, the network produces a 2×2 downsampled depth map. Of course, different scaling factors may also be used.
[0172] In some of the above embodiments, maximum filtering, maximum pooling, enhancement or similar operations are mentioned at the encoder. It should be understood that these embodiments assume that the depth is encoded as 1 / d (or other similar inverse relationship), where d is the distance from the camera. In this assumed case, high values in the depth map indicate foreground objects and low values in the depth map indicate background. Therefore, by applying a maximum operation or an enhanced operation, the method tends to enlarge the foreground object. The corresponding reverse process at the decoder can be to apply a minimum operation or an attenuated operation.
[0173] Of course, in other embodiments, depth can be encoded as d or log d (or another variable with a direct correlation to d). This means that the foreground object is represented by a low value of d, and the background is represented by a high value of d. In such embodiments, minimum filtering, minimum pooling, weakening or similar operations can be performed at the encoder. Once again, this will tend to expand the foreground object (which is the goal). The corresponding reverse process at the decoder can be to apply a maximum operation or an enhanced operation.
[0174] Figure 2 , Figure 4 , Figure 5 , Figure 8 and Fig. 9 The encoding method and the decoding method and Figure 3 and Figure 6The encoder and decoder may be implemented in hardware or software or a hybrid of the two (e.g., implemented as firmware running on a hardware device). To the extent that the embodiments are partially or completely implemented in software, the functional steps illustrated in the process flow chart may be performed by a suitably programmed physical computing device (e.g., one or more central processing units (CPUs), graphics processing units (GPUs), or neural network accelerators (NNAs)). Each process and its individual component steps as illustrated in the flow chart may be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including a computer program code, and when the program is run on one or more physical computing devices, the computer program code is configured to cause the one or more physical computing devices to perform the above-mentioned encoding method or decoding method.
[0175] Storage media may include volatile and nonvolatile computer memory, such as RAM, PROM, EPROM, and EEPROM. Various storage media may be fixed within a computing device or may be removable so that one or more programs stored on the storage media are loaded into a processor.
[0176] The metadata according to the embodiments may be stored on a storage medium. The code stream according to the embodiments may be stored on the same storage medium or on a different storage medium. The metadata may be embedded in the code stream, but this is not required. Likewise, the metadata and / or the code stream (with the metadata in the code stream or separated from the code stream) may be transmitted as a signal modulated onto an electromagnetic carrier.
[0177] The signal can be defined according to digital communication standards. The carrier can be an optical carrier, a radio frequency wave, a millimeter wave, or a near field communication wave. It can be wired or wireless.
[0178] To the extent that the embodiments are partially or fully implemented in hardware, Figure 3 and Figure 6 The blocks shown in the block diagram may be separate physical components, or may be the result of logical subdivision of a single physical component, or may all be implemented in one physical component in an integrated manner. In an implementation, the functions of a block shown in the figure may be split among multiple components, or the functions of multiple blocks shown in the figure may be combined in a single component. For example, although Figure 6 The texture decoder 424 and the depth decoder 46 are shown as separate components, but their functionality may be provided by a single unified decoder component.
[0179] Hardware components suitable for use in embodiments of the present invention include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs). One or more blocks may be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuits for performing other functions.
[0180] Those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention by studying the drawings, the disclosure and the claims. In the claims, the word "comprising" does not exclude other elements or steps, and the words "one" or "an" do not exclude multiple. A single processor or other unit can implement the functions of several items recorded in the claims. Although certain measures are recorded in mutually different dependent claims, this does not indicate that the combination of these measures cannot be used advantageously. If a computer program is discussed above, the computer program can be stored / distributed on a suitable medium, for example, an optical storage medium or solid-state medium provided with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. If the term "suitable for" is used in the claims or description, it should be noted that the term "suitable for" is intended to be equivalent to the term "configured to". Any figure mark in the claims should not be interpreted as limiting the scope.
Claims
1. A method for encoding video data comprising one or more source views, each source view comprising a texture map and a depth map, the method comprising: include: receiving (110) the video data; processing the depth map of at least one source view to generate a processed depth map, the processing comprising: Non-linear filtering (120), and downsample(130); and The processed depth map and the texture map of the at least one source view are encoded (140) to generate a video stream.
2. The method according to claim 1, in, The non-linear filtering includes enlarging an area of at least one foreground object in the depth map.
3. The method according to claim 1 or 2, in, The nonlinear filtering includes applying filters designed using machine learning algorithms.
4. The method according to any one of the preceding claims, in, The nonlinear filtering is performed by a neural network including a plurality of layers, and the downsampling is performed between two layers of the plurality of layers.
5. The method according to any one of the preceding claims, in, The method comprises processing (120a, 130a) the depth map according to a plurality of sets of processing parameters to generate a respective plurality of processed depth maps, The method further comprises: selecting a set of processing parameters that reduce reconstruction errors of a reconstructed depth map after a corresponding processed depth map has been encoded and decoded; and A metadata codestream is generated that identifies the selected set of parameters.
6. A method for decoding video data comprising one or more source views, the method comprising: include: Receiving (210) a video stream comprising an encoded depth map and an encoded texture map for at least one source view; decoding the encoded depth map (226) to produce a decoded depth map; decoding (224) the encoded texture map to produce a decoded texture map; and and processing the decoded depth map to generate a reconstructed depth map, wherein the processing comprises: Upsampling (230), and Non-linear filtering (240).
7. The method according to claim 6, further comprising: include: Prior to the step of processing the decoded depth map to generate the reconstructed depth map, it is detected that the resolution of the decoded depth map is lower than the resolution of the decoded texture map.
8. The method according to claim 6 or 7, in, The non-linear filtering includes reducing the area of at least one foreground object in the depth map.
9. The method according to any one of claims 6 to 8, in, The processing of the decoded depth map is based at least in part on the decoded texture map.
10. The method according to any one of claims 6 to 9, include: Upsampling the decoded depth map (230); identifying (242) surrounding pixels of at least one foreground object in the upsampled depth map; determining (244) whether the surrounding pixels are more similar to the foreground object or to the background based on the decoded texture map; and Non-linear filtering (240a) is applied only to surrounding pixels that are determined to be more similar to the background.
11. The method according to any one of claims 6 to 10, in, The non-linear filtering includes smoothing the edges of at least one foreground object (250).
12. The method according to any one of claims 6 to 11, further comprising receiving a metadata stream associated with the video stream, the metadata stream identifying a set of parameters, The method also includes processing the decoded depth map according to the identified set of parameters.
13. A computer program comprising computer code for causing a processing system to carry out the method according to any one of claims 1 to 12 when said program is run on the processing system.
14. A video encoder (300) configured to encode video data comprising one or more source views, each source view comprising a texture map and a depth map, the video encoder include: An input unit (310) configured to receive the video data; A video processor (320) configured to process the depth map of at least one source view to generate a processed depth map, the processing comprising: Non-linear filtering (120), and downsample(130); An encoder (330) configured to encode the texture map and the processed depth map of the at least one source view to generate a video stream; and An output unit (360) is configured to output the video code stream.
15. A video decoder (400) configured to decode video data comprising one or more source views, the video decoder include: A bitstream input unit (410) is configured to receive a video bitstream, wherein the video bitstream includes an encoded depth map and an encoded texture map for at least one source view; A first decoder (426) configured to decode the encoded depth map according to the video stream to generate a decoded depth map; A second decoder (424) configured to decode the encoded texture map according to the video code stream to generate a decoded texture map; A reconstruction processor (450) configured to process the decoded depth map to generate a reconstructed depth map, wherein the processing comprises: Upsampling (230), and Non-linear filtering (240), and An output unit (470) is configured to output the reconstructed depth map.