Methods and apparatus of restoring compressed video by neural networks with segmentation for video coding

By classifying blocks with similar residual patterns and applying individual neural networks as in-loop filters, the method improves PSNR in video coding systems, addressing the challenge of chaotic residual patterns and enhancing video quality.

WO2025130289A1PCT designated stage expired Publication Date: 2025-06-26MEDIATEK INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/124884
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-10-15
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing video coding systems face challenges in significantly improving Peak Signal-to-Noise Ratio (PSNR) between original and reconstructed frames, even with well-trained CNNs as in-loop or post-loop filters, due to chaotic residual patterns.

Method used

The proposed method classifies blocks with similar patterns in residual distribution into the same category, using features like average luma intensity, variance of luma samples, and prediction mode, and applies individual neural networks as in-loop filters for each category to generate filtered outputs.

Benefits of technology

This approach improves the PSNR by refining each category of blocks with tailored neural network filtering, effectively addressing the challenge of chaotic residual patterns and enhancing video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124884_26062025_PF_FP_ABST
    Figure CN2024124884_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding using coding tools including one or more cross component models related modes are disclosed. According to the method, target data to be restored is derived, wherein the target data comprises residual data, reconstructed data, filtered-reconstructed data or a combination thereof for one or more blocks of said one or more pictures. The target data is classified into multiple categories according to one or more features. For the target data in each of the multiple categories, the target data in said each of the multiple categories is processed using an individual NN (Neural Network) as an in-loop filter for said each of the multiple categories to generate NN in-loop filtered output for said each of the multiple categories. The NN in-loop filtered output for said each of the multiple categories is provided.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS OF RESTORING COMPRESSED VIDEO BY NEURAL NETWORKS WITH SEGMENTATION FOR VIDEO CODING

[0001] CROSS REFERENCE TO RELATED APPLICATIONS

[0002] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 612,385, filed on December 20, 2023. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0003] The present invention relates to video coding system. In particular, the present invention relates to using neural networks for in-loop filtering in a video coding system.

[0004] BACKGROUND AND RELATED ART

[0005] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0006] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video  data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0007] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0008] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0009] Neural Network (NN) , also referred as an 'A rtificial' Neural Network (ANN) , is an information-processing system that has certain performance characteristics in common with biological neural networks. A Neural Network system is made up of a number of simple and highly interconnected processing elements to process information by their dynamic state response to external inputs. The processing element can be considered as a neuron in the human brain, where each perceptron accepts multiple inputs and computes weighted sum of the inputs. In the field of neural network, the perceptron is considered as a mathematical model of a biological neuron. Furthermore, these interconnected processing elements are often organized in layers. For recognition applications, the external inputs may correspond to patterns are presented to the network, which communicates to one or more middle layers, also called 'hidden layers' , where the actual processing is done via a system of weighted 'connections' .

[0010] Artificial neural networks may use different architecture to specify what variables are involved in the network and their topological relationships. For example, the variables involved in a neural network might be the weights of the connections between the neurons, along with activities of  the neurons. Feed-forward network is a type of neural network topology, where nodes in each layer are fed to the next stage and there is connection among nodes in the same layer. Most ANNs contain some form of 'learning rule' , which modifies the weights of the connections according to the input patterns that it is presented with. In a sense, ANNs learn by example as do their biological counterparts. Backward propagation neural network is a more advanced neural network that allows backwards error propagation of weight adjustments. Consequently, the backward propagation neural network is capable of improving performance by minimizing the errors being fed backwards to the neural network.

[0011] The NN can be a deep neural network (DNN) , convolutional neural network (CNN) , recurrent neural network (RNN) , or other NN variations. Deep multi-layer neural networks or deep neural networks (DNN) correspond to neural networks having many levels of interconnected nodes allowing them to compactly represent highly non-linear and highly-varying functions. Nevertheless, the computational complexity for DNN grows rapidly along with the number of nodes associated with the large number of layers.

[0012] The CNN is a class of feed-forward artificial neural networks that is most commonly used for analysing visual imagery. A recurrent neural network (RNN) is a class of artificial neural network where connections between nodes form a directed graph along a sequence. Unlike feedforward neural networks, RNNs can use their internal state (memory) to process sequences of inputs. The RNN may have loops in them so as to allow information to persist. The RNN allows operating over sequences of vectors, such as sequences in the input, the output, or both.

[0013] In the present invention, methods and apparatus to classify reconstructed video data into categories and apply neural network in-loop filtering based on the classified categories are disclosed to improve the performance.

[0014] BRIEF SUMMARY OF THE INVENTION

[0015] A method and apparatus for video coding using coding tools including one or more cross component models related modes are disclosed. According to the method, input data at an encoder side or a video bitstream at a decoder side is received, wherein the input data comprise one or more pictures in a video sequence or the video bitstream comprises compressed data associated with said one or more pictures in the video sequence, and wherein each picture comprises one or more colour components. Target data to be restored is derived, wherein the target data comprises residual data, reconstructed data, filtered-reconstructed data or a combination thereof for one or more blocks of said one or more pictures. The target data is classified into multiple categories according to one or more features. For the target data in each of the multiple categories, the target data in said each of the multiple categories is processed using an individual NN (Neural Network) as an in-loop filter for said each of the multiple categories to generate NN in-loop filtered output for said each of the multiple categories. The NN in-loop filtered output for said each of the multiple categories is provided.

[0016] In one embodiment, the target data corresponds to signal right before deblocking filter, SAO (Sample Adaptive Offset) , CCSAO (Cross-Component SAO) , ALF (Adaptive Loop Filter) , or CCALF (Cross-Component ALF) .

[0017] In one embodiment, said one or more features are available at both the encoder side and the decoder side. In one embodiment, the target data is segmented into multiple MxN blocks and said classifying the target data is applied to the multiple MxN blocks, and wherein M and N are positive integers. In one embodiment, said one or more features comprise average or median luma intensity of a block, average or median weighted YUV value of the block, variance of luma samples of the block, average edge response of luma sample of the block, number of edge pixels of the block, magnitude of residual, or a combination thereof.

[0018] In one embodiment, said one or more features are available only at the encoder side. In one embodiment, the target data is segmented into multiple MxN blocks and said classifying the target data is applied to the multiple MxN blocks, and wherein M and N are positive integers. In one embodiment, said one or more features comprise prediction mode, QP value of each CU within a block, split depth of each CU within the block, or a combination thereof. In one embodiment, the prediction mode comprises intra prediction, uni-prediction with L0, uni-prediction with L1, bi-prediction, or combined prediction with inter and intra predictions.

[0019] In one embodiment, said one or more features are available at both the encoder side and the decoder side, or only at the encoder side. In one embodiment, said one or more features comprise average or median luma intensity of a block, average or median weighted YUV value of the block, variance of luma samples of the block, average edge response of luma sample of the block, number of edge pixels of the block, magnitude of residual, prediction mode, QP value of each CU within the block, split depth of each CU within the block, or a combination thereof.

[0020] In one embodiment, two shared neural networks are trained for all of the multiple categories, and wherein one of the two shared neural networks is used for luma component and another of the two shared neural networks is used for chroma component.

[0021] In one embodiment, one shared neural network is trained for all of the multiple categories, and wherein said one shared neural network is used for all colour components.

[0022] In one embodiment, two dedicated neural networks are trained for each of the multiple categories, and wherein one of the two dedicated neural networks is used for luma component and another of the two dedicated neural networks is used for chroma component.

[0023] In one embodiment, one dedicated neural network is trained for each of the multiple categories, and wherein said one dedicated neural networks is used for all colour components.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.

[0025] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0026] Fig. 2 illustrates an example of two shared neural networks trained for all of the multiple categories, and one of the two shared neural networks is used for luma component (Fig. 2A) and the other is used for chroma component (Fig. 2B) .

[0027] Fig. 3 illustrates an example of one shared neural network trained for all of the multiple categories, and the shared neural network is used for all colour components.

[0028] Fig. 4 illustrates a flowchart of an exemplary video coding system that classifies reconstructed data into multiple categories and trains neural network in-loop filtering for each of the multiple categories according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0029] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0030] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0031] Our recent research results have demonstrated that it is hard to significantly improve PSNR between original frame and reconstruction frame even if we applied a well-trained CNN as an in-loop filter or post-loop filter. The processing order of CNN in-loop filter can be before deblocking filter, before SAO, before CCSAO (Cross-Component SAO) , before ALF, before CCALF (Cross-Component ALF) or after ALF. The reason is that the residual between original frame and reconstruction frame is quite chaotic and lacks a regular pattern to follow. In order to solve this problem, we propose to classify blocks with similar patterns in the distribution of residuals into the same category, reducing the difficulty of training. We assume that blocks with similar features have similar patterns in the distribution of  residuals. Therefore, we will classify blocks with similar features into the same category and train them accordingly. During testing, blocks will also be classified based on certain features, and each category has its corresponding inference method. The features used as classification criteria can be information that only the encoder side can access, or information that both the encoder side and decoder side can access. Our proposed method can be applied to CNN post-loop filter or CNN in-loop filter.

[0032] The proposed method classifies blocks of input frame (s) into multiple categories according to some features derived from input frame (s) or from encoding information. As mentioned earlier, the location of CNN in-loop filter can be before deblocking filter, before SAO, before CCSAO, before ALF, before CCALF or after ALF. Therefore, the “input frame (s) ” refer to the signal at the corresponding locations. Each category is refined (i.e., in-loop filtered) in a different way. Each proposed method consists of three parts. The first part is how to classify blocks of input frame (s) into different categories when training neural networks. The second part is how to design architecture of neural network. The third part is how to merge the outputs of output branches or networks in the inference process.

[0033] In one embodiment of the first part, we classify all blocks of input frame (s) into G categories according to a feature or multiple features that are available at both the encoder side and the decoder side (s) , where G is a positive integer and larger than one. The unit of segmentation is denoted as an M × N block, where M and N are positive integer. Each M × N block will be classified into a certain category. The valid features that can be used to make decision can be one or a combination of the following:

[0034] 1. Average luma intensity of the block

[0035] 2. Median luma intensity of the block

[0036] 3. Average weighted YUV value of the block

[0037] 4. Median weighted YUV value of the block

[0038] 5. Variance of luma sample of the block

[0039] 6. Average edge response of luma sample of the block

[0040] 7. Number of edge pixels of the block

[0041] 8. Magnitude of residual

[0042] In another embodiment of the first part, we classify all blocks of input frame (s) into G categories according to a feature or multiple features that are only available at encoder side, where G is a positive integer and larger than one. Note that the features used to make decision must be signalled to decoder since they are only available at encoder side. The unit of segmentation is denoted as an M × N block, where M and N are positive integer. Each M × N block will be classified into a certain category. The valid features that can be used to make decision can be one or a combination of the following:

[0043] 1. Prediction mode, (intra prediction, uni-prediction with L0, uni-prediction with L1, bi-prediction, or combined prediction with inter and intra predictions)

[0044] 2. QP value of each CU within the block

[0045] 3. Split depth of each CU within the block

[0046] In another embodiment of the first part, we classify all blocks of input frame (s) into G categories according to a feature or multiple features. The features we used can be available at both the encoder side and the decoder side or can only be available at the encoder side, where G is a positive integer and larger than one. The unit of segmentation is denoted as an M × N block, where M and N are positive integers. Each M × N block will be classified into a certain category. The valid features that can be used to make decision can be one or a combination of the following:

[0047] 1. Average luma intensity of the block

[0048] 2. Median luma intensity of the block

[0049] 3. Average weighted YUV value of the block

[0050] 4. Median weighted YUV value of the block

[0051] 5. Variance of luma sample of the block

[0052] 6. Average edge response of luma sample of the block

[0053] 7. Number of edge pixels of the block

[0054] 8. Magnitude of residual

[0055] 9. Prediction mode, (intra prediction, uni-prediction with L0, uni-prediction with L1, bi-prediction, or combined prediction with inter and intra predictions)

[0056] 10. QP value of each CU within the block

[0057] 11. Split depth of each CU within the block

[0058] In one embodiment of the second part, we train two shared neural networks for all categories. One refines luma component and the other one refines chroma components.

[0059] We design a shared neural network for processing luma component in a way such that there are K1 different output branches, where K1 is a positive integer and larger than or equal to one. We design another shared neural network for processing chroma component in a way such that there are K2 different output branches, where K2 is a positive integer and larger than or equal to one. Note that K1 and K2 cannot both be equal to 1. An output branch can consist of multiple convolutional layers or just one output channel in the last layer. An example of overall architecture is illustrated in Fig. 2A (for Y component) and Fig. 2B (for U / V components) . In Fig. 2A, shared layers 210 of a shared neural network is used to generate K1 output branches 220. The K1 output ports 230 (i.e., Yi (x, y) , i=0, …, (K1-1) ) are weighted by individual weights ywi, j 240 and combined using adder 250 to generate the blended luma Yblend. Similarly, Fig. 2B illustrates the processing for chroma components (e.g. U and V) . In Fig. 2B, shared layers 212 of a shared neural network is used to generate K2 chroma output branches 222. The K2 output ports (i.e., Ui and Vi ) , i=0, …, (K2-1) ) 232 are weighted by individual weights uwi, j and  vwi, j 242 and combined using adders (252 for U and 254 for V) to generate the blended chroma Ublend and Vblend respectively.

[0060] As for the design of loss function, assume that the input block is classified into the i-th category, we calculate loss of an M × N block by following equations:

[0061] Loss= WY×L (Yblend, Ytarget) +WU×L (Ublend, Utarget) +WV×L (Vblend, Vtarget) (4)

[0062] where Yjj (x, y) represents for block located at coordinate (x , y ) of Y component output of j-th output branch or j-th neural network. L represents for loss function. Ytarget , Utarget and Vtarget are the learning target. In most of cases,

[0063] Vtarget= Voriginal-Yreconstruction,

[0064] Utarget= Uoriginal-Ureconstruction,

[0065] Vtarget= Voriginal-Vreconstruction.

[0066] As for the values of blending coefficients, ywi, j , uwi, j and vwi, j , they can be predefined constants or outputs of a module or multiple modules that are trainable. If they are outputs of a module or multiple modules that are trainable, it means that we learn adequate blending coefficients during the training process instead of defining fixed blending coefficients in advance. The modules used to generate blending coefficients can be convolutional neural networks or any other modules with learnable parameters.

[0067] In another embodiment of the second part, we train a shared neural network for all categories. We design a shared neural network 310 for processing all colour components in a way such that there are K1 different output branches 320 for luma and K2 different output branches 322 for chroma as shown in Fig. 3, where K1 and K2 are positive integers and larger than or equal to one. Note that K1 and K2 cannot both be equal to 1. An output branch can consist of multiple convolutional layers or just one output channel in the last layer. An example of overall architecture for processing all colour components is illustrated as Fig. 3. As for the design of loss function, we calculate loss of a M × N block by equations (1) , (2) and (3) . As for the values of blending coefficients, ywi, j 340, uwi, j and vwi, j 342, they can be predefined constants or outputs of a module or  multiple modules that are trainable. If they are outputs of a module or multiple modules that are trainable, it means that we learn the adequate blending coefficients during training process instead of defining a fixed blending coefficients in advance. The modules used to generate blending coefficients can be convolutional neural networks or any other modules with learnable parameters.

[0068] In another embodiment of the second part, we remove the shared layers. In other words, we train a neural network for each category and refine each category with the corresponding neural network.

[0069] In another embodiment of the second part, we train two dedicated neural networks for each category. One refines luma component and the other one refines chroma components. We design a dedicated neural network for processing luma component in a way such that there are K1 different output branches, where K1 is a positive integer and larger than or equal one. We design another dedicated neural network for processing chroma component in a way such that there are K2 different output branches, where K2 is a positive integer and larger than or equal one. Note that K1 and K2 cannot both be equal to 1. An output branch can consist of multiple convolutional layers or just one output channel in the last layer. The overall architecture for a specific category is illustrated as Fig. 2A and Fig. 2B. As for the design of loss function, we calculate loss of an M × N block by equations (1) , (2) and (3) . As for the values of blending coefficients, ywi, j , uwi, j and vwi, j , they can be predefined constants or outputs of a module or multiple modules that are trainable. If they are outputs of a module or multiple modules that are trainable, it means that we learn the adequate blending coefficients during training process instead of defining a fixed blending coefficients in advance. The modules used to generate blending coefficients can be convolutional neural networks or any other modules with learnable parameters.

[0070] In another embodiment of the second part, we train a dedicated neural network for each category. It refines luma and chroma components. We design a dedicated neural network for processing all color components in a way such that there are K1 different output branches for luma and K2 different output branches for chroma. where K1 and K2 are positive integers and larger than or equal to one. Note that K1 and K2 cannot both be equal to 1. An output branch can consist of multiple convolutional layers or just one output channel in the last layer. The overall architecture for processing all color components is illustrated as Fig. 3. As for the design of loss function, we calculate loss of a M × N block by equations (1) , (2) and (3) . As for the values of blending coefficients, ywi, j , uwi, j and vwi, j , they can be predefined constants or outputs of a module or multiple modules that are trainable. If they are outputs of a module or multiple modules that are trainable, it means that we learn the adequate blending coefficients during training process instead of defining a fixed blending coefficients in advance. The modules used to generate blending coefficients can be  convolutional neural networks or any other modules with learnable parameters.

[0071] In one embodiment of the third part, we merge the outputs of output branches or networks in a block level. More specifically speaking, each M × N block in the final output is computed by following equations:

[0072] Assume that the input block is classified into the i-th category

[0073] Yj(x, y) represents for block located at coordinate (x , y ) of Y component output of j-th output branch or j-th neural network. As for the values of blending coefficients, we use the same set of blending coefficients as what we used in training process. Note that the values of blending coefficients can be predefined constants or outputs of a module or multiple modules that are trainable.

[0074] In another embodiment of the third part, we merge the outputs of output branches or networks in a block level. The blended output can be obtained through equations (5) , (6) and (7) . As for the values of blending coefficients, we use a different set of blending coefficients as what we used in training process.

[0075] In another embodiment of the third part, we merge the outputs of output branches or networks in a block level and define multiple sets of blending coefficients. For each set of blending coefficients, we calculate the PSNR between blended output and target. The blended output can be obtained through equations (5) , (6) and (7) . We select the best blended output with the highest PSNR. This embodiment needs to signal the index of best set of blending coefficients to the decoder.

[0076] In another embodiment of the third part, we merge the outputs of output branches or networks in a frame level. More specifically speaking, the final output is computed by below equation:

[0077] Assume that the input frame is classified into the i-th category:

[0078] As for the values of blending coefficients, yci j, uci, j and vc, ,j , we use the same set of blending coefficients as what we used in the training process. Note that the values of blending coefficients can be predefined constants or outputs of a module or multiple modules that are trainable.

[0079] In another embodiment of the third part, we merge the outputs from the output branches or networks in a frame level. The blended output can be obtained through equations (8) , (9) and (10) . As for the values of blending coefficients, we use a different set of blending coefficients as what we used in the training process.

[0080] In another embodiment of the third part, we merge the outputs of output branches or networks in a frame level and define several sets of blending coefficients. For each set blending coefficients, we calculate the PSNR between blended output and target. The blended output can be obtained through equations (8) , (9) and (10) . We select the best blended output with the highest PSNR. This embodiment needs to signal the index of best set of blending coefficients to decoder.

[0081] Each of our proposed method consists of three parts. We have mentioned all embodiments of each part above. A proposed method can be obtained by combining one of the above-mentioned embodiments of each part.

[0082] The proposed methods of neural network based in-loop filtering with classified signals as described above can be implemented in an encoder side or a decoder side. For example, any of the proposed method of neural network based in-loop filtering with classified signals can be implemented in an Intra / Inter coding module (e.g. Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module is an encoder (e.g. Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . Any of the proposed methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required cross-component prediction processing. While the Intra Pred. units (e.g. unit 110 / 112 in Fig. 1A and unit 150 / 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0083] Fig. 4 illustrates a flowchart of an exemplary video coding system that classifies reconstructed data into multiple categories and trains neural network in-loop filtering for each of the multiple categories according to an embodiment of the present invention. The steps shown in the  flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data at an encoder side or a video bitstream at a decoder side is received in step 410, wherein the input data comprise one or more pictures in a video sequence or the video bitstream comprises compressed data associated with said one or more pictures in the video sequence, and wherein each picture comprises one or more colour components. Target data to be restored is derived in step 420, wherein the target data comprises residual data, reconstructed data, filtered-reconstructed data or a combination thereof for one or more blocks of said one or more pictures. The target data is classified into multiple categories according to one or more features in step 430. For the target data in each of the multiple categories, the target data in said each of the multiple categories is processed using an individual NN (Neural Network) as an in-loop filter for said each of the multiple categories to generate NN in-loop filtered output for said each of the multiple categories in step 440. The NN in-loop filtered output for said each of the multiple categories is provided in step 450.

[0084] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0085] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0086] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be  performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0087] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding for a video encoder or decoder, the method comprising:receiving input data at an encoder side or receiving a video bitstream at a decoder side, wherein the input data comprise one or more pictures in a video sequence or the video bitstream comprises compressed data associated with said one or more pictures in the video sequence, and wherein each picture comprises one or more colour components;deriving target data to be restored, wherein the target data comprises residual data, reconstructed data, filtered-reconstructed data or a combination thereof for one or more blocks of said one or more pictures;classifying the target data into multiple categories according to one or more features;for the target data in each of the multiple categories, processing the target data in said each of the multiple categories using an individual NN (Neural Network) as an in-loop filter for said each of the multiple categories to generate NN in-loop filtered output for said each of the multiple categories; andproviding the NN in-loop filtered output for said each of the multiple categories.2.The method of Claim 1, wherein the target data corresponds to signal right before deblocking filter, SAO (Sample Adaptive Offset) , CCSAO (Cross-Component SAO) , ALF (Adaptive Loop Filter) , or CCALF (Cross-Component ALF) .3.The method of Claim 1, wherein said one or more features are available at both the encoder side and the decoder side.4.The method of Claim 3, wherein the target data is segmented into multiple MxN blocks and said classifying the target data is applied to the multiple MxN blocks, and wherein M and N are positive integers.5.The method of Claim 4, wherein said one or more features comprise average or median luma intensity of a block, average or median weighted YUV value of the block, variance of luma samples of the block, average edge response of luma sample of the block, number of edge pixels of the block, magnitude of residual, or a combination thereof.6.The method of Claim 1, wherein said one or more features are available only at the encoder side.7.The method of Claim 6, wherein the target data is segmented into multiple MxN blocks and said classifying the target data is applied to the multiple MxN blocks, and wherein M and N are positive integers.8.The method of Claim 7, wherein said one or more features comprise prediction mode, QP value of each CU within a block, split depth of each CU within the block, or a combination thereof.9.The method of Claim 8, wherein the prediction mode comprises intra prediction, uni-prediction with L0, uni-prediction with L1, bi-prediction, or combined prediction with inter and intra predictions.10.The method of Claim 1, wherein said one or more features are available at both the encoder side and the decoder side, or only at the encoder side.11.The method of Claim 10, wherein said one or more features comprise average or median luma intensity of a block, average or median weighted YUV value of the block, variance of luma samples of the block, average edge response of luma sample of the block, number of edge pixels of the block, magnitude of residual, prediction mode, QP value of each CU within the block, split depth of each CU within the block, or a combination thereof.12.The method of Claim 1, wherein two shared neural networks are trained for all of the multiple categories, and wherein one of the two shared neural networks is used for luma component and another of the two shared neural networks is used for chroma component.13.The method of Claim 1, wherein one shared neural network is trained for all of the multiple categories, and wherein said one shared neural network is used for all colour components.14.The method of Claim 1, wherein two dedicated neural networks are trained for each of the multiple categories, and wherein one of the two dedicated neural networks is used for luma component and another of the two dedicated neural networks is used for chroma component.15.The method of Claim 1, wherein one dedicated neural network is trained for each of the multiple categories, and wherein said one dedicated neural networks is used for all colour components.16.An apparatus for video coding in a video encoder or decoder, the apparatus comprising one or more electronics or processors arranged to:receive input data at an encoder side or receiving a video bitstream at a decoder side, wherein the input data comprise one or more pictures in a video sequence or the video bitstream comprises compressed data associated with said one or more pictures in the video sequence;derive target data to be restored, wherein the target data comprises residual data, reconstructed data, filtered-reconstructed data or a combination thereof for one or more blocks of said one or more pictures;classify the target data into multiple categories according to one or more features available at both the encoder side and the decoder side;for the target data in each of the multiple categories, process the target data in said each of the multiple categories using an NN (Neural Network) as an in-loop filter for said each of the multiple categories to generate NN in-loop filtered output for said each of the multiple categories; andprovide the NN in-loop filtered output for said each of the multiple categories.

Citation Information

Patent Citations

  • Method and apparatus of neural network based processing in video coding

    CN107925762A

  • Method and apparatus of neural network for video coding

    CN111133756A

  • System and method for video coding and decoding

    CN114845101A

  • Sample offset with predetermined filter

    CN115428453A

  • Multiple neural network models for filtering during video coding

    US20220215593A1