Down-sampling method and down-sampling rate based on super-resolution video coding and decoding
By adopting a neural network-based super-resolution method in video encoding and decoding, dynamically adjusting the downsampling rate and components, the problems of fixed downsampling rate and improper processing of chrominance components in the prior art are solved, and a more efficient video encoding and decoding effect is achieved.
Patent Information
- Application Number
- CN202380072839.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-16
- Publication Date
- 2025-05-27
AI Technical Summary
The existing neural network-based super-resolution design for video encoding and decoding has multiple problems, including that all frames are downsampled, the input downsampling rate is fixed, the traditional downsampling method is inefficient, and the downsampling of chroma components will affect the encoding and decoding performance.
The neural network-based super-resolution method is adopted to improve the encoding and decoding efficiency by determining the downsampling rate suitable for each frame and setting different downsampling rates according to different color components.
It realizes more efficient video encoding and decoding, and improves bit rate saving and codec performance by dynamically adjusting downsampling rate and component flexible processing.
Smart Images

Figure CN120051994A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority and benefit of International Patent Application PCT / CN2022 / 125383, filed on October 14, 2022. The entire content of the above - mentioned patent application is incorporated herein by reference. Technical field
[0003] This disclosure relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background art
[0004] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video may continue to grow. Summary of the invention
[0005] The first aspect relates to a method for processing video data, including: determining to apply neural network (NN) - based super - resolution (SR), wherein the input chroma format changes due to different downsampling rates of color components; and performing conversion between visual media data and a bitstream based on the chroma format.
[0006] The second aspect relates to an apparatus for processing video data, including: a processor; and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute any one of the foregoing aspects.
[0007] The third aspect relates to a non - transitory computer - readable medium, including a computer program product for use in a video codec device, the computer program product including computer - executable instructions stored on the non - transitory computer - readable medium, such that when the computer - executable instructions are executed by a processor, the video codec device executes the method according to any one of the foregoing aspects.
[0008] The fourth aspect relates to a non - transitory computer - readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to apply neural network (NN) - based super - resolution, wherein the input chroma format changes due to different downsampling rates of color components; and generating a bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of a video, including: determining to apply neural network (NN) - based super - resolution, wherein the input chroma format changes due to different downsampling rates of color components; generating a bitstream based on the determination; and storing the bitstream in a non - transitory computer - readable recording medium.
[0010] The sixth aspect relates to the methods, apparatuses or systems described in this document.
[0011] For clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0012] These features, as well as other features, will be more clearly understood from the following detailed description taken in conjunction with the drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more fully understand the present disclosure, reference is now made to the following brief description taken in conjunction with the drawings and the detailed description, in which like reference numerals represent like parts.
[0014] Figure 1 is a schematic diagram showing an example of reference picture resampling (RPR).
[0015] Figure 2 is a schematic diagram showing an example of deconvolution.
[0016] Figure 3 is a schematic diagram showing an example of pixel shuffle-based upsampling.
[0017] Figure 4 is a schematic diagram showing an example SR network, where RB represents a residual block, and R and M represent the number of feature maps after convolution.
[0018] Figure 5 is a schematic diagram showing an example of obtaining a residual block, where M represents the number of filters.
[0019] Figure 6 is a schematic diagram of an example of the inverse pixel shuffle process.
[0020] Figures 7A to 7D is a schematic diagram showing examples of different positions for upsampling.
[0021] Figure 8 is a schematic diagram showing an example downsampling network.
[0022] Figure 9 is a schematic diagram of an example model for luminance upsampling.
[0023] Figure 10 is a block diagram showing an example video processing system.
[0024] Figure 11 is a block diagram of an example video processing apparatus.
[0025] Figure 12 is a flowchart of an example method for video processing.
[0026] Figure 13 It shows a block diagram of an exemplary video codec system.
[0027] Figure 14 It shows a block diagram of an exemplary encoder.
[0028] Figure 15 It shows a block diagram of an exemplary decoder.
[0029] Figure 16 It is a schematic diagram of an exemplary encoder. Detailed implementation
[0030] First, it should be understood that although exemplary detailed implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of technologies, whether currently known or yet to be developed. The present disclosure should not in any way be limited to the exemplary detailed implementations, drawings, and technologies shown below, including the exemplary designs and specific implementations shown and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0031] The use of section headings in this document is for ease of understanding and does not limit the applicability of the technologies and embodiments disclosed in each section to that section. Additionally, the technologies described herein are also applicable to other video codec protocols and designs.
[0032] 1. Preliminary discussion
[0033] This document relates to video codec technology. Specifically, it relates to super-resolution-based upsampling technology in video coding and decoding. It can be applied to existing video codec standards such as High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC). It can also be applicable to other video codec standards or video codecs, or used as a post-processing method outside the encoding / decoding process.
[0034] 2. Abbreviations
[0035] This disclosure includes the following abbreviations: Adaptive Loop Filter (ALF), Deblocking Filter (DBF), Discrete Cosine Transform (DCT)-based Interpolation Filter (DCTIF), High-Resolution (HR), International Organization for Standardization (ISO), International Electrotechnical Commission (IEC), Low-Resolution (LR), Reference Picture Resampling (RRP), Sample Adaptive Offset (SAO), VVC Test Model (VTM), and Versatile Video Coding (VVC).
[0036] 3. Video Coding Standard
[0037] Video coding standards have evolved mainly through the development of the well-known International Telecommunication Union - Telecommunication Standardization Sector (ITU-T) and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET has adopted many additional methods and incorporated them into a reference software called the Joint Exploration Model (JEM). The Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to develop the VVC standard. Compared with HEVC, VVC version 1 achieves approximately 50% bitrate reduction.
[0038] An example of the VVC draft, namely Versatile Video Coding (Draft 10), can be found at the following URL: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=10399. The example reference software VTM for VVC can be found at the following URL: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags.
[0039] Figure 1 An example of the Reference Picture Resampling (RPR) 100 is shown. Reference Picture Resampling (RPR) is a mechanism in VVC where pictures in the reference list can be stored at a different resolution from the current picture and then resampled to perform regular decoding operations. The addition of this technology supports interesting application scenarios such as real-time communication with adaptive resolution and adaptive streaming with an open Group of Pictures (GOP) structure. As Figure 1 shown, the down-sampled (also known as downsampled or down sampled) sequence is encoded and then the reconstruction is up-sampled (also known as upsampled or up sampled) after decoding.
[0040] 3.1 Upsampling Techniques
[0041] Common or traditional upsampling techniques are discussed. In VTM 11.0, the upsampling filter is the Discrete Cosine Transform (DCT)-based Interpolation Filter (DCTIF). In addition, bicubic interpolation and bilinear interpolation are also commonly used. In these techniques, once the number of taps of the filter is given, the weight coefficients of the interpolation filter are fixed. Therefore, the weight coefficients of these methods may not be optimal.
[0042] Learning-based techniques are discussed. Figure 2 An example of deconvolution 200 is shown. Deconvolution and pixel shuffle layers are two example solutions in deep learning-based upsampling techniques. Deconvolution, also known as transposed convolution, is commonly used for upsampling in deep learning. In this method, the stride of the convolution is the same as the scaling ratio. The bottom matrix is the low-resolution input, where the white blocks are values filled with zeros and the gray blocks represent the original samples of the low resolution. The top matrix is the high-resolution output. In this example, stride = 2.
[0043] Figure 3 An example of upsampling 300 based on pixel shuffle is shown. The pixel shuffle layer is another upsampling method used in deep learning. As Figure 3As shown, pixel shuffle is usually placed after the convolutional layer. The number of filters for this convolution is M = C out r 2 , where C out is the number of output channels and r represents the magnification ratio. For example, given a low-resolution input of size H×W×3, if the size of the high-resolution output is 2H×2W×3, then the number of filters M = 3×2 2 = 12. The pixel shuffle technique will be described in further detail below with reference to Figure 6 .
[0044] 3.2 Convolutional Neural Network-Based Super-Resolution for Video Coding and Decoding
[0045] 3.2.1 Convolutional Neural Network
[0046] Super-resolution (SR) is the process of restoring a high-resolution (HR) image from a low-resolution (LR) image. SR can also be referred to as upsampling. In deep learning, a convolutional neural network (also known as CNN or ConvNet) is a class of deep neural networks used for analyzing visual images. CNN has had very successful applications in the fields of image and video recognition / processing, recommendation systems, image classification, medical image analysis, natural language processing, etc.
[0047] CNN is a regularized version of the multi-layer perceptron. The multi-layer perceptron usually refers to a fully connected network, that is, each neuron in one layer is connected to all neurons in the next layer. The "fully connectivity" of these networks makes them prone to overfitting the data. Regularization is used to alleviate overfitting, such as adding a form of weight magnitude metric to the loss function. CNN takes a different approach to regularization. That is, CNN exploits the hierarchical patterns in the data and uses smaller and simpler patterns to combine into more complex patterns. Therefore, on the scale of connectivity and complexity, CNN is at the lower extreme.
[0048] Compared with other image classification / processing algorithms, CNN uses relatively less preprocessing. This means that the network learns the filters that are hand-designed in traditional algorithms. In feature design, this property of being independent of prior knowledge and artificial feature design is a major advantage.
[0049] 3.2.2 Deep Learning for Image / Video Coding and Decoding
[0050] Deep learning-based image / video compression has two meanings: end-to-end compression purely based on neural networks (NN) and frameworks enhanced by neural networks. The first type adopts an autoencoder-like structure, which is implemented through convolutional neural networks or recurrent neural networks. Although relying solely on neural networks for image / video compression can avoid any manual optimization or hand design, the compression efficiency may not be satisfactory. Therefore, the work distributed in the second category uses neural networks as an aid and enhances the compression framework by replacing or enhancing some modules. In this way, they can inherit the advantages of highly optimized frameworks.
[0051] 3.2.3 Super-Resolution Based on Convolutional Neural Networks
[0052] In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, resulting in distortion of the reconstructed frame. In the context of the RPR of VVC, the input image / video can be downsampled. Therefore, the resolution of the original frame is twice that of the reconstructed resolution. To upsample the low-resolution reconstruction, a convolutional neural network can be trained to learn the mapping from the distorted low-resolution frame to the original high-resolution frame. In practice, the training must be carried out before deploying the NN-based loop filter. A CNN-based block upsampling method was proposed for HEVC. For each coding tree unit (CTU) block, this method determines whether to use the down / upsampling-based method or the full-resolution-based coding and decoding.
[0053] 3.2.3.1 Training
[0054] The purpose of the training process is to find the optimal values of the parameters including weights and biases. First, a codec (e.g., HEVC test model (HM), joint exploration model (JEM), VTM, etc.) is used to compress the training dataset to generate the distorted reconstructed frames. Then the reconstructed frames (low-resolution and compressed) are fed into the NN, and the cost is calculated using the output of the NN and the ground truth frame (also called the original frame). Commonly used cost functions include the sum of absolute differences (SAD) and the mean square error (MSE). Next, the gradient of the cost with respect to each parameter is derived through the backpropagation algorithm. With the gradient, the values of the parameters can be updated. The above process is repeated until the convergence criterion is met. After the training is completed, the derived optimal parameters are saved for the inference phase.
[0055] 3.2.3.2 Convolution Process
[0056] During convolution, the filter moves from left to right and from top to bottom over the image. When moving horizontally, one column of pixels changes, and when moving vertically, one row of pixels changes. The amount of movement between applying the filter to the input image is called the stride, and its height and width dimensions are almost always symmetric. The default stride in two dimensions for both height and width movement is (1,1).
[0057] In most deep convolutional neural networks, residual blocks are used as basic modules and stacked multiple times to build the final network. Figure 5 is a schematic diagram showing an example of obtaining a residual block 500, where M represents the number of filters. As Figure 5 shown in the example, the residual block is obtained by combining a convolutional layer, a rectified linear unit (ReLU) / parametric rectified linear unit (PReLU) activation function convolutional layer.
[0058] 3.2.3.3 Inference
[0059] During the inference phase, the distorted reconstructed frame is fed into the neural network and processed by the NN model whose parameters have been determined during the training phase. The input samples of the NN can be the reconstructed samples before or after deblocking (DB), or the reconstructed samples before or after sample adaptive offset (SAO), or the reconstructed samples before or after adaptive loop filter (ALF).
[0060] 4. Technical problems solved by the disclosed technical solutions
[0061] However, the existing designs of NN-based super-resolution for video coding and decoding have various problems or drawbacks.
[0062] First, all frames in the sequence are downsampled. However, some frames may be more preferably encoded at full resolution. It is desirable to design a mechanism that determines whether to perform downsampling or upsampling on each frame. This can contribute to random access or low-latency configurations.
[0063] Second, most of the input downsampling rates are fixed, such as 2x downsampling. It may be beneficial to provide different downsampling rates for different video units (e.g., a video unit can be a frame or a coding tree unit (CTU)).
[0064] Third, the method of downsampling the original input sequence is usually a traditional downsampling method, such as bilinear interpolation. Neural network-based downsampling can provide higher BD-rate savings.
[0065] Fourth, if the chrominance component is downsampled before encoding, the chrominance coding and decoding performance may degrade. By applying different downsampling rates to the luminance and chrominance components and only downsampling the luminance component, this problem can be solved.
[0066] 5. List of Solutions and Examples
[0067] To solve the above problems and other problems, the methods outlined below are disclosed. Detailed examples should be regarded as examples explaining general concepts and should not be interpreted narrowly. In addition, these examples can be applied individually or combined in any way. Moreover, the examples presented in this document can be applied together with the examples in other documents.
[0068] In this disclosure, NN-based SR can be any kind of NN-based method, such as SR based on a convolutional neural network (CNN). In the following discussion, NN-based SR can also be referred to as a non-CNN-based method, for example, using a machine learning-based solution.
[0069] Figure 4 is a schematic diagram showing an example SR network 400.
[0070] Figure 5 is a schematic diagram showing an example of a residual block 500.
[0071] Figure 6 is an example of an inverse pixel shuffle process 600.
[0072] In the following discussion, a video unit (also referred to as a video data unit) can be a picture sequence, a picture, a strip, a slice, a block, a sub-picture, a CTU / coding tree block (CTB), a CTU / CTB row, one or more coding units (CUs) / coding blocks (CBs), one or more CTUs / CTBs, one or more virtual pipelined data units (VPDUs), or a sub-region within a picture / strip / slice / block. In some examples, the video unit can be referred to as a video data unit.
[0073] Example 1
[0074] In one example, the downsampling method can employ a designed filter.
[0075] In one example, a discrete cosine transform interpolation filter (DCTIF) can be used for downsampling.
[0076] In one example, bilinear interpolation can be used for downsampling.
[0077] In one example, bicubic interpolation can be used for downsampling.
[0078] In one example, the downsampling method can be signaled from the encoder to the decoder. In one example, an index can be signaled to indicate the downsampling filter. In one example, at least one coefficient of the downsampling filter can be signaled directly or indirectly. The downsampling method can be signaled in the sequence header / sequence parameter set (SPS) / picture parameter set (PPS) / picture header / slice header / CTU / CTB or any rectangular region. Different downsampling methods can be signaled for different color components.
[0079] In one example, the downsampling method can be required on the decoder side and notified to the encoder side in an interactive application.
[0080] Example 2
[0081] In one example, the downsampling method can be a neural network (NN)-based method, such as a convolutional neural network (CNN)-based method.
[0082] The CNN-based downsampling method should include at least one downsampling layer. In one example, a convolution with a stride of K (e.g., K = 2) can be used as the downsampling layer, and the downsampling rate is K. In one example, a pixel shuffle method can be used, followed by a convolution with a stride of 1 for downsampling. Pixel shuffle is shown in Figure 6 In.
[0083] Example 3
[0084] A series of downsamplings can be used to achieve a specific downsampling rate. In one example, two convolutional layers with a stride of K (e.g., K = 2) are used in a network. In this case, the downsampling rate is 4. In one example, two traditional downsampling filters (e.g., each with a downsampling rate of 2) are used to achieve a downsampling rate of 4.
[0085] Example 4
[0086] In one example, a traditional filter and a CNN-based method can be combined to achieve a specific downsampling rate. In one example, a traditional filter is used, followed by a CNN-based method. The traditional filter achieves 2x downsampling, and the CNN-based method achieves 2x downsampling. Thus, the input is downsampled 4 times.
[0087] Example 5
[0088] When downsampling a specific input video unit level, different downsampling methods can be compared with each other to select the best or preferred downsampling method.
[0089] In one example, there are K (e.g., K = 3) CNN-based downsampling models. For a specific input, the three downsampling models will downsample the input respectively. The downsampled reconstruction will be upsampled to the original resolution. A quality metric (e.g., peak signal-to-noise ratio (PSNR)) is used to measure the three upsampling results. The model achieving the best performance will be used for actual downsampling. In one example, the quality metric is the multi-scale structural similarity index metric (MS-SSIM). In one example, the quality metric is PSNR.
[0090] The index of the downsampling method can be signaled to the encoder or decoder.
[0091] Example 6
[0092] The downsampling method can be signaled to the decoder.
[0093] In one example, a CNN-based downsampling method is used for downsampling. For a specific video unit (e.g., frame) level, the index of the selected model will be signaled to the decoder.
[0094] In one example, different CTUs within a frame use different downsampling methods. In this case, all the indices of the corresponding methods can be signaled to the decoder.
[0095] In one example, at least one coefficient of the downsampling filter can be signaled directly or indirectly.
[0096] For different color components, different downsampling methods can be signaled.
[0097] In one example, the downsampling method can be required on the decoder side and notified to the encoder side in an interactive application.
[0098] Example 7
[0099] The input of the downsampling method can be at all video unit (e.g., sequence / picture / strip / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) levels.
[0100] In one example, the input is at the frame level, and its size is at its original resolution.
[0101] In one example, the input is at the CTU level with a size of 128x128.
[0102] Example 8
[0103] In one example, the input is a block within a frame, and its size is not restricted.
[0104] In one example, it can be a block with a spatial domain size of (M, N). For example, M = 256 and N = 128.
[0105] Example 9
[0106] In one example, for all video unit (e.g., sequence / picture / strip / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) levels, the downsampling rate can be different.
[0107] In one example, for all frames of a sequence, the downsampling rate is 2.
[0108] In one example, for all CTUs of a frame, the downsampling rate is 2.
[0109] In one example, the downsampling rate is 2 for the first frame and can be 4 for the next frame.
[0110] Combinations of downsampling rates for different video unit levels can be used. In one example, the downsampling rate is 2 for a frame and can be 4 for a CTU in the same frame. In this case, the CTU will be downsampled 4 times.
[0111] Example 10
[0112] In one example, for all components at the input video unit level, the downsampling rate can be different.
[0113] In one example, the downsampling rate is 2 for both the luminance component and the chrominance component.
[0114] In another example, the downsampling rate is 2 for the luminance component and 4 for the chrominance component.
[0115] Example 11
[0116] In one example, the downsampling rate can be 1, which means no downsampling is performed.
[0117] Such a downsampling rate can be applied to all video unit (e.g., sequence / picture / strip / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) levels.
[0118] Example 12
[0119] The downsampling rate can be determined by comparison.
[0120] In one example, a frame can use downsampling rates of 2x and 4x. In this case, the encoder can compress the frame with 2x downsampling and then compress the frame with 4x downsampling. Subsequently, the low-resolution reconstruction can be upsampled using the same upsampling method. Then, a quality metric (e.g., PSNR) is calculated for each result, and the downsampling rate that achieves the best reconstruction quality is selected as the downsampling rate for compression. In one example, the quality metric is MS-SSIM.
[0121] Example 13
[0122] This determination can be performed at the encoder or at the decoder.
[0123] If the determination is made at the decoder, the distortion can be calculated based on samples other than the current picture / strip / / CTU / CTB or any rectangular region.
[0124] Example 14
[0125] Different quality metrics can be used as the metrics for comparison.
[0126] In one example, the quality metric is PSNR.
[0127] In one example, the quality metric is SSIM.
[0128] In one example, the quality metric is MS-SSIM.
[0129] In one example, the quality metric is Video Multimethod Assessment Fusion (VMAF).
[0130] Example 15
[0131] In one example, the downsampling rate can be signaled at the video unit level.
[0132] In one example, CNN information can be signaled in the SPS / PPS / picture header / strip header / CTU / CTB.
[0133] Example 16
[0134] In one example, the input chroma format will change due to different downsampling rates for the color components.
[0135] In one example, the input chroma format is YUV 4:2:0, and when the downsampling rate for the luminance component is 2 and the downsampling rate for the chrominance component is 1, the input chroma format will change to YUV 4:4:4.
[0136] Example 17
[0137] In one example, downsampling of the luma component can be performed at all video unit (e.g., sequence / picture / strip / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) levels.
[0138] In one example, the downsampled luma component is from a frame.
[0139] In one example, the downsampled luma component is from a CTU.
[0140] Example 18
[0141] In one example, the changed chroma format will be used for compression in the encoder.
[0142] In one example, the chroma format is changed from YUV 4:2:0 to YUV 4:4:4, and YUV 4:4:4 will be used as the chroma format for compression.
[0143] In one example, information needed to restore the original chroma format is signaled at the video unit level. In one example, the downsampling rate of all color components is signaled in the SPS / PPS / picture header / strip header / CTU / CTB. In one example, the original chroma format is signaled in the SPS / PPS / picture header / strip header / CTU / CTB.
[0144] Example 19
[0145] In one example, neural network-based tools can be used in the encoder and decoder.
[0146] In one example, to achieve better coding and decoding performance, a neural network-based loop filter can be used. In one example, two neural network-based loop filters are applied to the downsampled luma reconstruction and chroma reconstruction respectively. In one example, the downsampled luma reconstruction is concatenated with the chroma reconstruction as an input to the neural network-based loop filter for chroma.
[0147] In one example, to restore the original chroma format, neural network-based super-resolution can be applied to the reconstructed luma component of YUV and used as a post-processing filter.
[0148] In one example, to restore the original chroma format, neural network-based super-resolution can be applied before the loop filter.
[0149] Other examples
[0150] It is proposed that for two sub-regions within a video unit (e.g., picture / strip / slice / sub-picture), two different SR methods can be applied.
[0151] In one example, the SR method may include an NN-based solution.
[0152] In one example, the SR method may include a non-NN-based solution (e.g., via a traditional filter).
[0153] In one example, for a first sub-region, an NN-based solution is used, while for a second sub-region, a non-NN-based solution is used.
[0154] In one example, for a first sub-region, an NN-based solution with a first design / model is used, while for a second sub-region, an NN-based solution with a second design / model is used. In one example, the first / second design may have different inputs. In one example, the first / second design may have a different number of layers. In one example, the first / second design may have different strides.
[0155] In one example, an indication of the allowed SR method and / or which SR method will be used for the sub-region may be signaled in the bitstream or dynamically derived. In one example, it may be derived based on decoding information (e.g., how many samples / what proportion of the samples are intra-coded). In one example, it may be derived based on the SR solution used for a reference sub-region (e.g., a co-located sub-region).
[0156] A candidate set for a video unit may be predefined or signaled in the bitstream, where the candidate set may include multiple SR solutions for samples in the video unit for selection therefrom.
[0157] In one example, the candidate set may include multiple NN-based methods with different models / designs.
[0158] In one example, the candidate set may include NN-based methods and non-NN-based methods.
[0159] In one example, different candidate sets of the NN-based SR model are used for different situations, e.g., according to the decoded information. In one example, there are different sets of the NN-based SR model corresponding to different color components and / or different stripe types and / or different quantization parameters (QPs). In one example, QPs can be classified into several groups. For example, different NN-based SR models can be used for different groups [QP / M], where M is an integer such as 6. In one example, QPs are fed into the SR model, and one model can correspond to all QPs. In this case, only one QP group is used. In one example, the luminance component and the chrominance component can adopt different sets of the NN-based SR model. In one example, the first set of the NN-based SR model is applied to the luminance component, and the second set of the NN-based SR model is applied to at least one chrominance component. In one example, each color component is associated with its own set of the NN-based SR model. Additionally, alternatively, the number of sets of the NN-based SR model to be applied to the three color components can depend on the stripe / picture type and / or the split tree type (single tree or dual tree), etc. In one example, two stripe types (e.g., I stripe and B (or P) stripe) can utilize different sets of the NN-based SR model. In one example, for the first color component, two stripe types (e.g., I stripe and B (or P) stripe) can utilize different sets of the NN-based SR model; while for the second color component, two stripe types (e.g., I stripe and B (or P) stripe) can use the same set of the NN-based SR model. In one example, for each QP or QP group, an NN-based SR model is trained. The number of NN models is equal to the number of QPs or QP groups.
[0160] In one example, NN (e.g., CNN)-based SR and traditional filters can be used together.
[0161] In one example, for different video unit (e.g., sequence / picture / stripe / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) levels, different upsamplings can be used together. For example, for different CTUs in a picture, some CTUs can select traditional filters, while other CTUs can preferably use the NN-based SR method.
[0162] In one example, the selection between the NN-based SR and the traditional filter can be signaled from the encoder to the decoder. This selection can be signaled in the sequence header / SPS / PPS / picture header / stripe header / CTU / CTB or any rectangular region. Different selections can be signaled for different color components.
[0163] In the above example, a traditional filter can be used as an upsampling method.
[0164] In one example, a Discrete Cosine Transform Interpolation Filter (DCTIF) can be used as an upsampling method.
[0165] In one example, bilinear interpolation can be used as an upsampling method.
[0166] In one example, bicubic interpolation can be used as an upsampling method.
[0167] In one example, Lanczos interpolation can be used as an upsampling method.
[0168] In one example, the upsampling method can be signaled from the encoder to the decoder. In one example, an index can be signaled to indicate the upsampling filter. In one example, at least one coefficient of the upsampling filter can be signaled directly or indirectly. The upsampling method can be signaled in the sequence header / SPS / PPS / picture header / slice header / CTU / CTB or any rectangular region. Different upsampling methods can be signaled for different color components.
[0169] In one example, the upsampling method can be required on the decoder side and notified to the encoder side in an interactive application.
[0170] In one example, NN-based SR can be used as an upsampling method. In one example, the network of SR should include at least one upsampling layer. In one example, the neural network can be a CNN. In one example, deconvolution with a stride of K (e.g., K = 2) can be used as an upsampling layer, as Figure 2 shown. In one example, the pixel shuffle method can be used as an upsampling layer, as Figure 3 shown.
[0171] NN (e.g., CNN)-based SR can be applied to certain slices / pictures, certain temporal layers, or certain slices / pictures according to the reference picture list information.
[0172] Some selections of upsampling methods are discussed below.
[0173] Whether and / or how to use NN (e.g., CNN)-based SR (denoted as CNN information) can depend on the video standard profile or level.
[0174] Whether and / or how to use NN (e.g., CNN)-based SR (denoted as CNN information) can depend on the color component.
[0175] Whether and / or how to use NN (e.g., CNN)-based SR (denoted as CNN information) can depend on the picture / strip type.
[0176] Whether and / or how to use NN (e.g., CNN)-based SR (denoted as CNN information) can depend on the content or codec information of the video unit. In one example, when the variance of the reconstructed samples is greater than a predefined threshold, NN-based SR will be used. In one example, when the energy of the high-frequency components of the reconstructed samples is greater than a predefined threshold, NN-based SR will be used.
[0177] Whether and / or how to use NN (e.g., CNN)-based SR (denoted as CNN information) can be controlled at the video unit (e.g., sequence / picture / strip / slice / tile / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB) level. The CNN information can include an indication of enabling / disabling the CNN filter, which CNN filter to apply, CNN filtering parameters, the CNN model, the stride of the convolutional layer, and / or the precision of the CNN parameters.
[0178] In one example, the CNN information can be signaled at the video unit level. In one example, the CNN information can be signaled in the sequence header / SPS / PPS / picture header / strip header / CTU / CTB or any rectangular region.
[0179] The number of sets of different CNN SR models and / or CNN ensemble models can be signaled to the decoder. For different color components, the number of sets of different CNN SR models and / or CNN ensemble models can be different.
[0180] In one example, a rate-distortion optimization (RDO) strategy or a distortion minimization strategy is used to determine the upsampling for a video unit.
[0181] In one example, different CNN-based SR models will be used to upsample the current input (e.g., luminance reconstruction). Then, the PSNR values between the reconstructions upsampled by different CNN-based SR models and the corresponding original input (the original input without downsampling and compression) are calculated. The model that achieves the highest PSNR value will be selected as the model for upsampling. The index of this model can be signaled. In one example, the MS-SSIM value (instead of the PSNR value) is used as the metric for comparison.
[0182] In one example, different traditional upsampling filters are compared, and the filter that achieves the best quality metric is selected. In one example, the quality metric is PSNR.
[0183] In one example, different CNN-based SR models and traditional filters are compared, and the filter that achieves the best quality metric is selected. In one example, the quality metric is PSNR.
[0184] This determination can be performed at the encoder or the decoder. If the determination is made at the decoder, the distortion can be calculated based on samples other than the current picture / strip / / CTU / CTB or any rectangular region.
[0185] Different quality metrics can be used as the metric for comparison. In one example, the quality metric is PSNR. In one example, the quality metric is SSIM. In one example, the quality metric is MS-SSIM. In one example, the quality metric is VMAF.
[0186] The location of SR will be discussed in more detail below.
[0187] A super-resolution (SR) process, such as an NN-based or non-NN-based SR process, can be placed before the loop filter. In one example, the SR process can be invoked immediately after a block (e.g., CTU / CTB) is reconstructed. In one example, the SR process can be invoked immediately after a region (e.g., CTU row) is reconstructed.
[0188] A super-resolution (SR) process, such as an NN-based or non-NN-based SR process, can be placed at different positions in the loop filter chain.
[0189] Figures 7A to 7D An example 700 of the location for upsampling is shown.
[0190] In one example, the SR process can be applied before or after a given loop filter. In one example, the SR process is placed before the DBF, as Figure 7A shown. In one example, the SR process is placed between the DBF and the SAO, as Figure 7B shown. In one example, the SR process is placed between the SAO and the ALF, as Figure 7C shown. In one example, super-resolution is placed after the ALF, as Figure 7D shown. In one example, the SR process is placed before the SAO. In one example, the SR process is placed before the ALF.
[0191] In one example, whether to apply SR before a given loop filter can depend on whether the loop filter decision process considers the original image.
[0192] An indication of the location of the SR process can be signaled in the bitstream or determined dynamically based on the decoded information.
[0193] SR processes such as NN-based or non-NN-based SR processes can be used specifically with other codec tools such as loop filters, i.e., when an SR process is applied, one or more loop filters may no longer be applied, and vice versa.
[0194] In one example, an SR process can be used specifically with at least one loop filter. In one example, when an SR process is applied, the original loop filters such as DB, SAO, and ALF are all turned off. In one example, when ALF is disabled, an SR process can be applied. In one example, when the cross-component ALF (CC-ALF) is disabled, the SR process can be applied to the chrominance components.
[0195] In one example, the signaling of the side information of the loop filtering method can depend on whether / how the SR process is applied.
[0196] In one example, whether / how the SR process is applied can depend on the use of the loop filtering method.
[0197] The following examples relate to the SR network structure.
[0198] The proposed NN-based (e.g., CNN-based) SR network includes multiple convolutional layers. An upsampling layer is used in the proposed network to upsample the resolution.
[0199] In one example, deconvolution with a stride K greater than 1 (e.g., K = 2) can be used for upsampling. In one example, K can depend on the decoded information (e.g., color format).
[0200] In one example, pixel shuffle is used for upsampling, as Figure 4 shown. Assume the downsampling rate is K, where the resolution of the LR input is 1 / K of the original input. The first 3×3 convolution is used to fuse the information from the LR input and generate a feature map. Then, the output feature map from the first convolutional layer passes through a number of sequentially stacked residual blocks, which are labeled as RB. The feature maps are labeled as M and R. The last convolutional layer takes the feature map from the last residual block as input and produces R (e.g., R = K*K) feature maps. Finally, a shuffle layer is used to generate a filtered image with the same spatial resolution as the original resolution.
[0201] In one example, residual blocks can be used in the SR network. In one example, a residual block consists of three sequentially connected components, as Figure 5 shown: a convolutional layer, a PReLU activation function, and a convolutional layer. The input of the first convolutional layer is added to the output of the second convolutional layer.
[0202] The input of an NN-based (e.g., CNN-based) SR network can be at different video unit (e.g., sequence / picture / strip / slice / block / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTB, or any rectangular region) levels. In one example, the input of the SR network can be a downsampled CTU block. In one example, the input is a downsampled entire frame.
[0203] The input of an NN-based (e.g., CNN-based) SR network can be a combination of different color components. In one example, the input can be a reconstructed luminance component. In one example, the input can be a reconstructed chrominance component. In one example, the input can be both the same reconstructed luminance and chrominance components.
[0204] In one example, the luminance component can be used as the input, and the output of an NN-based (e.g., CNN-based) SR network is an upsampled chrominance component.
[0205] In one example, the chrominance component can be used as the input, and the output of an NN-based (e.g., CNN-based) SR network is an upsampled luminance component.
[0206] An NN-based (e.g., CNN-based) SR network is not limited to upsampling the reconstruction. In one example, the decoded side information can be used as the input of an NN-based (e.g., CNN-based) SR network for upsampling. In one example, the predicted picture can be used as the input for upsampling. The output of the network is an upsampled predicted picture.
[0207] It is proposed that the codec (encoding / decoding) information can be utilized during the super-resolution process.
[0208] In one example, the codec information can be used as the input of an NN-based SR solution.
[0209] In one example, the codec information can be used to determine which SR solution to apply.
[0210] In one example, the codec information may include segmentation information, prediction information, and intra prediction mode, etc. In one example, the input includes reconstructed low-resolution samples and other decoded information (e.g., segmentation information, prediction information, and intra prediction mode). In one example, the segmentation information has the same resolution as the reconstructed low-resolution frame. The sample values in the segmentation are derived by averaging the reconstructed samples in the codec unit. In one example, the prediction information may be predicted samples generated from intra prediction or intra block copy (IBC) prediction or inter prediction. In one example, the intra prediction mode has the same resolution as the reconstructed low-resolution frame. The sample values in the intra prediction mode are derived by filling the intra prediction mode in the corresponding codec unit. In one example, the QP value information can be used as auxiliary information to improve the quality of upsampled reconstruction. In one example, a QP map is constructed by filling a matrix with QP values, and its spatial domain size is the same as other input data. The QP map will be fed into the super-resolution network.
[0211] The following example relates to the color components of the SR network input.
[0212] During the SR process is applied to the second color component, the information related to the first color component can be utilized.
[0213] The information related to the first color component can be used as the input for the SR process applied to the second color component.
[0214] The chrominance information can be used as the input for the luminance upsampling process.
[0215] Luminance information can be used as an input to the chrominance upsampling process. In one example, the luminance reconstruction samples before the loop filter can be used. Alternatively, the luminance reconstruction samples after the loop filter can be used. In one example, the input to the NN contains both chrominance reconstruction samples and luminance reconstruction samples. In one example, the luminance information can be downsampled to the same resolution as the chrominance component. The downsampled luminance information will be concatenated with the chrominance component. In one example, the downsampling method is bilinear interpolation. In one example, the downsampling method is bicubic interpolation. In one example, the downsampling method is a convolution with a stride equal to the scaling ratio of the original frame. In one example, the downsampling method is the inverse method of pixel shuffling. A high-resolution block (HR block) of size 4×4×1 will be downsampled to a low-resolution block (LR block) of size 2×2×4. The font of the first element in each channel of the LR block and its corresponding position. In one example, the downsampling method can depend on the color format, such as 4:2:0 or 4:2:2. In one example, the downsampling method can be signaled from the encoder to the decoder. Additionally, alternatively, whether the downsampling process is applied can depend on the color format. In another example, the color format is 4:4:4, and no downsampling is performed on the luminance information. In one example, the chrominance reconstruction samples before the loop filter can be used. Alternatively, the chrominance reconstruction samples after the loop filter can be used. In one example, the input to the NN contains both chrominance reconstruction samples and luminance reconstruction samples. In one example, the input to the NN contains both chrominance reconstruction samples and luminance prediction samples.
[0216] In one example, the information of one chrominance component (e.g., Cb) can be used as an input to the upsampling process of another chrominance component (e.g., Cr).
[0217] In one example, the input includes reconstruction samples and decoded information (e.g., mode information and prediction information). In one example, the mode information is a binary frame, where each value indicates whether a sample belongs to a skipped coding / decoding unit. In one example, the prediction information is derived via motion compensation of the coding / decoding unit for inter-frame coding / decoding.
[0218] In one example, the prediction information can be used as an input to the SR process applied to the reconstruction.
[0219] In one example, the luminance information of the predicted picture can be used as an input to the SR process of the reconstructed luminance component.
[0220] In one example, the luminance information of the predicted picture can be used as an input to the SR process of the reconstructed chrominance component.
[0221] In one example, the chrominance information of the predicted picture can be used as an input to the SR process of the reconstructed chrominance component.
[0222] In one example, the luminance and chrominance information of the predicted picture can be used together as the input for the reconstruction SR process (e.g., luminance reconstruction).
[0223] In the case where the prediction information is not available (such as when the codec mode is palette or PCM), the predicted samples are filled.
[0224] In one example, the segmentation information can be used as the input applied to the reconstruction SR process.
[0225] In one example, the segmentation information has the same resolution as the reconstructed low-resolution frame. The sample values in the segmentation are derived by averaging the reconstructed samples in the codec unit.
[0226] In one example, the intra prediction mode information can be used as the input applied to the reconstruction SR process.
[0227] In one example, the intra prediction mode of the current sample via intra or inter prediction can be used. In one example, an intra prediction mode matrix with the same resolution as the reconstruction is constructed as an input for the SR process. For each sample in the intra prediction mode matrix, its value comes from the intra prediction mode of the corresponding CU.
[0228] In one example, the above method can be applied to a specific picture / strip type, such as an I-strip / picture. For example, train an NN-based SR model to upsample the reconstructed samples in the I-strip.
[0229] In one example, the above method can be applied to B / P strips / pictures. For example, train an NN-based SR model to upsample the reconstructed samples in the B-strip or P-strip.
[0230] Regarding the processing unit for SR
[0231] The super-resolution / upsampling process can be performed at the SR unit level, where the SR unit transforms more than one sample / pixel.
[0232] In one example, the SR unit can be the same as the video unit that calls the downsampling process.
[0233] In one example, the SR unit can be different from the video unit that calls the downsampling process. In one example, even if downsampling is performed at the picture / strip / slice level, the SR unit can be a block (e.g., CTU). In one example, even if downsampling is performed at the CTU / CTB level, the SR unit can be a CTU row or multiple CTUs / CTBs.
[0234] Furthermore, alternatively, for the NN-based SR method, the input of the network can be set to the SR unit.
[0235] In addition, alternatively, for an NN-based SR method, the input to the network can be set to a region that includes the SR unit to be upsampled and other samples / pixels.
[0236] In one example, the SR unit can be indicated or predefined in the bitstream.
[0237] For two SR units, the super-resolution method / upsampling method can be different. In one example, the super-resolution method / upsampling method can include an NN-based solution and a non-NN-based solution (e.g., a traditional upsampling filtering method).
[0238] The input to the SR network can be at different video unit levels (e.g., sequence / picture / strip / slice / block / sub-picture / CTU / CTU row / one or more CUs or CTUs / CTBs, or any region covering more than one sample / pixel). In one example, the input to the SR network can be a downsampled CTU block. In one example, the input is a downsampled entire frame.
[0239] A CNN-based SR model can be used to upsample at different video unit levels. In one example, the CNN-based SR model is trained on frame-level data and used to upsample frame-level input. In one example, the CNN-based SR model is trained on frame-level data and used to upsample CTU-level input. In one example, the CNN-based SR model is trained on CTU-level data and used to upsample frame-level input.
[0240] In one example, the CNN-based SR model is trained on CTU-level data and used to upsample CTU-level input.
[0241] The following examples relate to side information of the SR network input.
[0242] The downsampling rate of the video unit can be regarded as the input to the SR network.
[0243] In addition, alternatively, the convolutional layer can be configured with a stride depending on the downsampling rate.
[0244] The downsampling rate of the SR network input can be any positive integer. In addition, alternatively, the minimum spatial resolution of the input should be 1×1.
[0245] The downsampling rate of the SR network input can be a ratio of any two positive integers, such as 3:2.
[0246] The horizontal downsampling rate and the vertical downsampling rate can be the same, or they can also be different.
[0247] It is proposed that encoding / decoding information can be utilized during the upsampling process.
[0248] In one example, the encoding / decoding information can be used as the input to a super-resolution network.
[0249] In one example, the encoding / decoding information can include but is not limited to prediction signals, segmentation structures, and intra-prediction modes.
[0250] Figure 8 An example downsampling network 800 is shown. The embodiments are as described below. First, given a sequence for compression, downsampling is performed at the picture level. Second, before encoding, the current frame is downsampled at a downsampling rate of 2. (Assume that downsampling rates of 2 and 4 need to be determined). In one example, the NN-based downsampling method shown in Figure 8 can be used. Figure 8 A downsampling network for the luminance component is shown, but it can also be used for the chrominance component. Additionally, Figure 8 the downsampling rate in Figure 4 is 2, so in order to provide the performance of 4x downsampling, the network can be applied twice. Third, the downsampled frame is encoded. Fourth, the low-resolution reconstruction is upsampled to the original resolution. In one example, the upsampling network can use the network shown in
[0251] This embodiment relates to the example project outlined in Section 5 of the previous article.
[0252] Figure 9 An example model of luminance upsampling 900 is shown.
[0253] In the example model 900, only a rescaling operation is applied to the luminance component to avoid loss of the chrominance component, which is commonly observed in super-resolution methods. The downsampled luminance component and the unchanged chrominance component are encoded and decoded in the 4:4:4 color format.
[0254] The upsampling model for the luminance component is as shown in Figure 9As shown. The input to the model consists of three parts, namely, low-resolution luminance reconstruction samples, low-resolution luminance prediction samples, and a QP map filled with QP values. These three parts are concatenated together and then fed into the first convolutional layer. The output of the first layer further passes through several residual blocks and an additional convolutional layer. Then, the shuffle layer generates the high-resolution reconstruction from the output of the last convolutional layer. To maintain a tolerable complexity, in this example, N is set to be equal to 16, and M is set to be equal to 96 and 64, respectively, for processing intra-frame and inter-frame stripes.
[0255] Use the CNN-based loop filter of JVET-AA0111. To achieve a better trade-off between luminance and chrominance codec performance, this example sets chromaQpOffset = -7 and performs only luminance downsampling only when QP > 32.
[0256] Figure 10 FIG. is a block diagram showing an example video processing system 4000 in which various techniques disclosed herein can be implemented. Various specific implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0257] System 4000 may include a codec component 4004 that can implement various codec or encoding methods described in this document. The codec component 4004 may reduce the average bit rate of the video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via a connected communication, as represented by component 4006. The bitstream (or coded) representation of the stored or transmitted video received at input 4002 may be used by component 4008 to generate pixel values or a displayable video that is sent to the display interface 4010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, while the corresponding decoding tools or operations that reverse the codec result will be performed by the decoder.
[0258] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), a DisplayPort, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0259] Figure 11 is a block diagram of an example video processing apparatus 4100. The apparatus 4100 may be used to implement one or more of the methods described herein. The apparatus 4100 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The (one or more) processors 4102 may be configured to implement one or more of the methods described in this document. The (one or more) memories 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics co-processor.
[0260] Figure 12 is a flowchart of an example method 4200 for video processing. The method 4200 includes determining to apply neural network (NN)-based super-resolution at step 4202. The input chroma format changes due to different downsampling rates of color components. At step 4204, a conversion is performed between visual media data and a bitstream based on the chroma format. According to an example, the conversion of step 4204 may include encoding at an encoder or decoding at a decoder.
[0261] It should be noted that the method 4200 may be implemented in a device for processing video data, the device including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to perform the method 4200. Additionally, the method 4200 may be executed by a non-transitory computer-readable medium, which includes a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec device is caused to perform the method 4200.
[0262] Figure 13FIG. 0 is a block diagram illustrating an example video coding and decoding system 4300 that may utilize the techniques of the present disclosure. The video coding and decoding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, which may be referred to as a video coding device. The destination device 4320 may decode the encoded video data generated by the source device 4310, which may be referred to as a video decoding device.
[0263] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly transmitted to the destination device 4320 via the I / O interface 4316 over a network 4330. The encoded video data may also be stored on a storage medium / server 4340 for access by the destination device 4320.
[0264] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320, or may be external to the destination device 4320, and may be configured to interface with an external display device.
[0265] The video encoder 4314 and the video decoder 4324 may operate according to video compression standards such as HEVC, VVC, and other current and / or additional standards.
[0266] Figure 14 FIG. 13 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 13The video encoder 4314 in the system 4300 shown in [figure]. The video encoder 4400 may be configured to perform any or all of the techniques of this disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in this disclosure may be shared among various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0267] The functional components of the video encoder 4400 may include a splitting unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding / decoding unit 4414. The prediction unit may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406.
[0268] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0269] In addition, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 may be highly integrated, but are shown separately in the example of the video encoder 4400 for purposes of explanation.
[0270] The splitting unit 4401 may split a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0271] The mode selection unit 4403 may select, for example, one of the coding / decoding modes (intra or inter) based on error results, and provide the resulting intra or inter coded block to the residual generation unit 4407 to generate residual block data and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy).
[0272] To perform inter prediction on a current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures from buffer 4413 other than the picture associated with the current video block.
[0273] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0274] In some examples, the motion estimation unit 4404 may perform uni-directional prediction for the current video block, and the motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Then the motion estimation unit 4404 may generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0275] In other examples, the motion estimation unit 4404 may perform bi-directional prediction for the current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. Then, the motion estimation unit 4404 may generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0276] In some examples, the motion estimation unit 4404 may output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 may signal the motion information for the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0277] In one example, the motion estimation unit 4404 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0278] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0279] As discussed above, the video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0280] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0281] The residual generation unit 4407 may generate residual data for the current video block by subtracting the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0282] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform the subtraction operation.
[0283] The transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0284] After the transform processing unit 4408 generates the transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0285] The inverse quantization unit 4410 and the inverse transform unit 4411 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.
[0286] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0287] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0288] Figure 15 is a block diagram showing an example of a video decoder 4500, which can be Figure 13 the video decoder 4324 in the system 4300 shown in. The video decoder 4500 can be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0289] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 4400.
[0290] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy encoded video data, and based on the entropy decoded video data, the motion compensation unit 4502 can determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 4502 can determine such information by performing AMVP and Merge modes.
[0291] The motion compensation unit 4502 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. An identifier for the interpolation filter to be used with sub-pixel accuracy may be included in a syntax element.
[0292] The motion compensation unit 4502 may use the interpolation filter used by the video encoder 4400 during encoding of a video block to compute interpolations of sub-integer pixels of a reference block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to received syntax information and use the interpolation filter to generate a prediction block.
[0293] The motion compensation unit 4502 may use some syntax information to determine the size of a block for encoding frames and / or slices of an encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0294] The intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in a bitstream. The inverse quantization unit 4504 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501, i.e., dequantizes. The inverse transform unit 4505 applies an inverse transform.
[0295] The reconstruction unit 4506 may add a residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0296] Figure 16 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing the techniques of VVC. The encoder 4600 includes three loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter respectively, using side information for encoding / decoding the signal transmission offset and filter coefficients. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and repair the artifacts created by the previous stages.
[0297] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using reference pictures obtained from the reference picture buffer 4612. Residual blocks from inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction result and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting an image to the DF 4602, SAO 4604, and ALF 4606 for filtering before these images are stored in the reference picture buffer 4612.
[0298] Next, a list of preferred solutions for some examples is provided.
[0299] The following solutions show technical examples discussed herein.
[0300] 1. A method for processing video data, comprising: determining to apply neural network (NN)-based super-resolution, wherein an input chroma format changes due to different downsampling rates of color components; and performing a conversion between visual media data and a bitstream based on the chroma format.
[0301] 2. The method according to solution 1, wherein when the downsampling rate of the luminance component is 2 and the downsampling rate of the chrominance component is 1, the input chroma format changes from YUV 4:2:0 to YUV 4:4:4.
[0302] 3. The method according to solution 1 or 2, wherein the luminance component is downsampled at a video unit, and wherein the video unit is a sequence, a picture, a slice, a tile, a sub-picture, a coding tree unit (CTU), a CTU row, one or more coding units (CU), one or more CTUs, one or more coding tree blocks (CTBs), or a combination thereof.
[0303] 4. The method according to any one of solutions 1 to 3, wherein the downsampled luminance component is a frame or a CTU.
[0304] 5. The method according to any one of Solutions 1 to 4, wherein compression is performed in the encoder using the changed chrominance format.
[0305] 6. The method according to any one of Solutions 1 to 5, wherein the input chrominance format is changed from YUV4:2:0 to YUV4:4:4, and wherein YUV4:4:4 is used as the chrominance format for compression.
[0306] 7. The method according to any one of Solutions 1 to 6, wherein information for restoring the original chrominance format is signaled at the video unit level.
[0307] 8. The method according to any one of Solutions 1 to 7, wherein the downsampling rates of all color components are signaled in the sequence parameter set (SPS), picture parameter set (PPS), picture header, slice header, CTU, CTB, or a combination thereof.
[0308] 9. The method according to any one of Solutions 1 to 8, wherein the original chrominance format is signaled in the SPS, PPS, picture header, slice header, CTU, CTB, or a combination thereof.
[0309] 10. The method according to any one of Solutions 1 to 9, wherein neural network-based tools are used in the encoder.
[0310] 11. The method according to any one of Solutions 1 to 10, wherein neural network-based tools are used in the decoder.
[0311] 12. The method according to any one of Solutions 1 to 11, wherein a neural network-based loop filter is used.
[0312] 13. The method according to any one of Solutions 1 to 12, wherein two neural network-based loop filters are respectively applied to the downsampled luminance reconstruction and chrominance reconstruction.
[0313] 14. The method according to any one of Solutions 1 to 13, wherein the downsampled luminance reconstruction is connected to the chrominance reconstruction and used as the input to the neural network-based loop filter for chrominance.
[0314] 15. The method according to any one of Solutions 1 to 14, wherein neural network-based super-resolution is applied to reconstruct the luminance component of YUV and used as a post-processing filter to restore the original chrominance format.
[0315] 16. The method according to any one of Solutions 1 to 15, wherein neural network-based super-resolution is applied before the loop filter to restore the original chrominance format.
[0316] 17. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Solutions 1 to 16.
[0317] 18. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec device is caused to perform the method according to any one of Solutions 1 to 16.
[0318] 19. A non-transitory computer-readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method comprises: determining to apply neural network (NN)-based super-resolution, wherein an input chroma format is changed due to different downsampling rates of color components; and generating the bitstream based on the determination.
[0319] 20. A method of storing a bitstream of a video, comprising: determining to apply neural network (NN)-based super-resolution, wherein an input chroma format is changed due to different downsampling rates of color components; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0320] 21. The method, apparatus, or system described in this document.
[0321] In the solutions described herein, an encoder can conform to format rules by generating an encoded / decoded representation according to the format rules. In the solutions described herein, a decoder can use the format rules to parse syntax elements in the encoded / decoded representation to produce a decoded video in the case of knowing the presence or absence of the syntax elements according to the format rules.
[0322] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may correspond, for example, to bits that are co-located or distributed at different positions within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on the transform and the coded error residual values, and may also use bits in the header and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream based on this determination, knowing the presence or absence of some fields, as described in the above solution. Similarly, the encoder may determine whether to include certain syntax fields, and may generate the coded representation accordingly by including or excluding the syntax fields in the coded representation.
[0323] The disclosed solutions, examples, embodiments, modules, and functional operations described in this document, as well as other solutions, examples, embodiments, modules, and functional operations, may be implemented in digital electronic circuitry or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed embodiments, as well as other embodiments, may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to the hardware, the apparatus may also include code that creates an execution environment for the computer programs being considered, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver apparatus.
[0324] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that contains other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program being considered, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0325] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0326] By way of example, processors suitable for executing a computer program include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. In general, a computer will also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, from which it receives data or to which it transfers data, or both. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read only memory (CD ROM) and digital versatile disc read only memory (DVD-ROM) discs. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0327] Although this disclosure contains many details, these should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be excised from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0328] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the operations shown be performed to achieve the desired result. Additionally, the separation of various system components described in the embodiments of this disclosure should not be understood as required in all embodiments.
[0329] Only a few specific embodiments and examples are described, and other specific embodiments, enhancements, and variations can be made based on what is described and illustrated in this disclosure.
[0330] When there is no intermediate component between a first component and a second component other than a wire, trace, or another medium, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than a wire, trace, or another medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include direct coupling and indirect coupling. Unless otherwise stated, the use of the word "about" means a range that includes the subsequent number ±10%.
[0331] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The examples should be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.
[0332] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and shown as separate or discrete in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, include: Determining to apply a neural network (NN) based super-resolution, wherein the chroma format of the input is changed due to different downsampling rates of the color components; and Conversion is performed between the visual media data and the bitstream based on the chroma format. 2 . The method of claim 1 , wherein when a down-sampling rate of a luma component is 2 and a down-sampling rate of a chroma component is 1, the chroma format of the input is changed from YUV 4:2:0 to YUV 4:4:
4.
3. The method according to claim 1 or 2, in, The luma component is downsampled at a video unit, and wherein the video unit is a sequence, a picture, a slice, a tile, a sub-picture, a codec tree unit (CTU), a CTU row, one or more codec units (CU), one or more CTUs, one or more codec tree blocks (CTB), or a combination thereof. 4 . The method according to claim 1 , wherein the downsampled luminance component is a frame or a CTU.
5. The method according to any one of claims 1 to 4, in, The changed chroma format is used for compression in the encoder.
6. The method according to any one of claims 1 to 5, in, The chroma format of the input is changed from YUV 4:2:0 to YUV 4:4:4, and wherein YUV 4:4:4 is used as the chroma format for compression.
7. The method according to any one of claims 1 to 6, in, Information used to restore the original chroma format is signaled at the video unit level.
8. The method according to any one of claims 1 to 7, in, The downsampling rates of all color components are signaled in a sequence parameter set SPS, a picture parameter set PPS, a picture header, a slice header, a codec tree unit CTU, a codec tree block CTB or a combination thereof.
9. The method according to any one of claims 1 to 8, in, The original chroma format is signaled in a sequence parameter set SPS, a picture parameter set PPS, a picture header, a slice header, a codec tree unit CTU, a codec tree block CTB or a combination thereof.
10. The method according to any one of claims 1 to 9, in, Use neural network based tools in coders.
11. The method according to any one of claims 1 to 10, in, Use neural network based tools in the decoder.
12. The method according to any one of claims 1 to 11, in, Use a neural network based loop filter.
13. The method according to any one of claims 1 to 12, in, Two neural network-based loop filters are applied for downsampled luminance reconstruction and chrominance reconstruction, respectively.
14. The method according to any one of claims 1 to 13, in, The downsampled luma reconstruction is concatenated with the chroma reconstruction and used as input to a neural network-based loop filter for chroma.
15. The method according to any one of claims 1 to 14, in, A neural network based super-resolution is applied to reconstruct the luminance component of YUV and as a post-processing filter to recover the original chroma format.
16. The method according to any one of claims 1 to 15, in, A neural network-based super-resolution is applied before the loop filter to restore the original chroma format.
17. The method according to any one of claims 1 to 16, in, The converting includes decoding the visual media data from the bitstream.
18. The method according to any one of claims 1 to 16, in, The converting includes encoding the visual media data into the bitstream.
19. A device for processing video data, include: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer executable instructions are executed by a processor, the video codec device performs the method according to any one of claims 1 to 18.
21. A non-transitory computer-readable recording medium, in, The non-transitory computer-readable recording medium stores a bit stream of a video, the bit stream being generated by a method performed by a video processing device, wherein the method comprises: Determining to apply a neural network (NN) based super-resolution, wherein the chroma format of the input is changed due to different downsampling rates of the color components; and The bitstream is generated based on the determination.
22. A method for storing a bit stream of a video, include: Determining the application of a neural network (NN) based super-resolution, wherein the chroma format of the input is changed due to different downsampling rates of the color components; generating the bitstream based on the determination; and The bit stream is stored in a non-transitory computer-readable recording medium.