Loop filtering and video encoding and decoding method, device and system based on neural network

CN120077656APending Publication Date: 2025-05-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100847.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing digital video compression technology still has shortcomings in reducing video transmission bandwidth and traffic pressure. Especially in the context of pursuing high video definition, more efficient encoding methods are needed to optimize the compression efficiency of video data.

Method used

A loop filtering method based on neural network is adopted, and the neural network is used for skip connection processing in the NNLF filter of the decoding end and the encoding end. Different modes are selected according to the residual adjustment flag to optimize the rate distortion cost, so that in the encoding and decoding process Make appropriate residual adjustments to improve video encoding performance.

Benefits of technology

Through the rate-distortion cost optimization of the neural network, the lag between training and testing can be compensated to a certain extent, the effect of loop filtering can be improved, and the performance of video encoding and decoding can be improved, especially the adaptation under the influence of inter-frame prediction technology. performance and coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077656A_ABST
    Figure CN120077656A_ABST
Patent Text Reader

Abstract

The invention discloses a loop filtering and video encoding and decoding method, device and system based on a neural network, and the method comprises the steps: selecting an NNLF mode for carrying out residual adjustment or an NNLF mode without carrying out residual adjustment when an encoding end carries out NNLF on a reconstructed image; and the decoding end selects one mode to carry out NNLF on the reconstructed image according to the mark. The coding performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Neural network-based loop filtering, video encoding and decoding method, device and system Technical Field

[0001] The embodiments of the present disclosure relate to, but are not limited to, video technology, and more specifically, to a neural network-based loop filtering method, video encoding and decoding method, device, and system. Background Art

[0002] Digital video compression technology primarily compresses large amounts of digital video data for easier transmission and storage. The original video sequence consists of luminance and chrominance components. During the digital video encoding process, the encoder reads a black-and-white or color image and divides each frame into largest coding units (LCUs) of equal size (e.g., 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) according to a rule, and further divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module can use intra-frame prediction and inter-frame prediction. Intra-frame prediction predicts the pixel information within the current block based on the information of the same image to eliminate spatial redundancy; inter-frame prediction can refer to the information of different images and use motion estimation to search for the motion vector information that best matches the current block to eliminate temporal redundancy; the transformation can convert the predicted residual into the frequency domain to redistribute its energy. Combined with quantization, it can remove information that the human eye is not sensitive to to eliminate visual redundancy; entropy coding can eliminate character redundancy based on the current context model and the probability information of the binary code stream to generate a code stream.

[0003] With the surge in Internet videos and people's increasing demand for video clarity, although existing digital video compression standards can save a lot of video data, there is still a need to pursue better digital video compression technology to reduce the bandwidth and traffic pressure of digital video transmission.

[0004] SUMMARY OF THE INVENTION

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] An embodiment of the present disclosure provides a neural network based loop filter (NNLF) method, which is applied to an NNLF filter at a decoding end. The NNLF filter includes a neural network and a jump connection branch from the input to the output of the NNLF filter. The method includes:

[0007] The residual adjustment flag roflag is used for decoding the reconstructed image. The roflag is used to indicate whether residual adjustment is required when performing NNLF on the reconstructed image.

[0008] When it is determined according to the roflag that residual adjustment is not required, the first mode is used to perform NNLF on the reconstructed image; when it is determined according to the roflag that residual adjustment is required, the second mode is used to perform NNLF on the reconstructed image;

[0009] The first mode is an NNLF mode that does not perform residual adjustment on the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the residual image.

[0010] An embodiment of the present disclosure further provides a neural network-based loop filtering method, which is applied to a NNLF filter at a decoding end. The NNLF filter includes a neural network and a skip connection branch from the input to the output of the NNLF filter. The method includes performing the following processing on each component when performing NNLF on a reconstructed image including three components input to the neural network:

[0011] A residual adjustment flag roflag is used for decoding the component of the reconstructed image, where the roflag is used to indicate whether residual adjustment is required when performing NNLF on the component of the reconstructed image;

[0012] When it is determined according to the roflag that residual adjustment is not required, performing NNLF on the component of the reconstructed image in the first mode; when it is determined according to the roflag that residual adjustment is required, performing NNLF on the component of the reconstructed image in the second mode;

[0013] The first mode is an NNLF mode that does not perform residual adjustment on the component of the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the component of the residual image.

[0014] An embodiment of the present disclosure also provides a video decoding method, which is applied to a video decoding device and includes: when performing neural network-based loop filtering (NNLF) on a reconstructed image, perform the following processing: when NNLF allows residual adjustment, perform NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure applied to the NNLF filter at the decoding end.

[0015] An embodiment of the present disclosure also provides a neural network-based loop filtering method, which is applied to the NNLF filter at the encoding end. The NNLF filter includes a neural network and a skip connection branch from the input to the output of the NNLF filter. The method includes:

[0016] Input the reconstructed image into the neural network to obtain the residual image output by the neural network;

[0017] Calculate the rate-distortion cost cost1 of performing NNLF on the reconstructed image in the first mode and the rate-distortion cost cost2 of performing NNLF on the reconstructed image in the second mode; wherein, the first mode is the NNLF mode without performing residual adjustment on the residual image, and the second mode is the NNLF mode of performing residual adjustment on the residual image;

[0018] When cost1 < cost2, select the first mode to perform NNLF on the reconstructed image; when cost2 < cost1, select the second mode to perform NNLF on the reconstructed image; when cost1 = cost2, select the first mode or the second mode to perform NNLF on the reconstructed image.

[0019] An embodiment of the present disclosure also provides a neural network-based loop filtering method, which is applied to the NNLF filter at the encoding end. The NNLF filter includes a neural network and a skip connection branch from the input to the output of the NNLF filter. The method includes:

[0020] Input the reconstructed image into the neural network to obtain the residual image output by the neural network; both the reconstructed image and the residual image include 3 components; and

[0021] Perform the following processing on each of the 3 components:

[0022] Calculate the rate-distortion cost cost1 of performing NNLF on this component of the reconstructed image in the first mode, and the rate-distortion cost cost2 of performing NNLF on this component of the reconstructed image in the second mode; the first mode is the NNLF mode without performing residual adjustment on this component of the residual image, and the second mode is the NNLF mode of performing residual adjustment on this component of the residual image;

[0023] In the case where cost1 < cost2, select the first mode to perform NNLF on this component of the reconstructed image; in the case where cost2 < cost1, select the second mode to perform NNLF on this component of the reconstructed image; in the case where cost1 = cost2, select the first mode or the second mode to perform NNLF on the reconstructed image.

[0024] An embodiment of the present disclosure further provides a video encoding method, which is applied to a video encoding device and includes: when performing neural network-based loop filtering NNLF on a reconstructed image, perform the following processing:

[0025] In the case where NNLF allows residual adjustment, perform NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure applied to the NNLF filter at the encoding end;

[0026] Encode a flag for residual adjustment of the reconstructed image to indicate whether residual adjustment is required when performing NNLF on the reconstructed image. <00000​​​​​​​​​​​​An embodiment of the present disclosure further provides a video encoding and decoding system, comprising the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0032] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, can implement the neural network-based loop filtering method as described in any embodiment of the present disclosure, or implement the video decoding method as described in any embodiment of the present disclosure, or implement the video encoding method as described in any embodiment of the present disclosure.

[0033] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0034] Summary of the Figures

[0035] The accompanying drawings are used to provide an understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.

[0036] FIG1A is a schematic diagram of an encoding and decoding system according to an embodiment, FIG1B is a schematic diagram of an encoding end in FIG1A , and FIG1C is a schematic diagram of a decoding end in FIG1A ;

[0037] FIG2 is a block diagram of a filter unit according to an embodiment;

[0038] FIG3A is a network structure diagram of an NNLF filter according to an embodiment; FIG3B is a structure diagram of a residual block in FIG3A ; FIG3C is a schematic diagram of the input and output of the NNLF filter in FIG3A ;

[0039] FIG4A is a structural diagram of a backbone network in an NNLF filter according to another embodiment; FIG4B is a structural diagram of a residual block in FIG4A ; FIG4C is a schematic diagram of the input and output of the NNLF filter in FIG4 ;

[0040] FIG5 is a schematic diagram of the basic network structure of the residual network;

[0041] FIG6A is a schematic diagram of iterative training of an NNLF model according to an embodiment; FIG6B is a schematic diagram of encoding testing of the NNLF model in FIG6A ;

[0042] FIG7 is a flowchart of an NNLF method applied to an encoding end according to an embodiment of the present disclosure;

[0043] FIG8 is a flowchart of an NNLF method applied to an encoding end according to another embodiment of the present disclosure;

[0044] FIG9 is a flowchart of a video encoding method according to an embodiment of the present disclosure;

[0045] FIG10 is a flowchart of an NNLF method applied to a decoding end according to an embodiment of the present disclosure;

[0046] FIG11 is a schematic diagram of a NNLF capable of mode selection according to an embodiment of the present disclosure;

[0047] FIG12 is a flowchart of an NNLF method applied to a decoding end according to another embodiment of the present disclosure;

[0048] FIG13 is a flowchart of a video decoding method according to an embodiment of the present disclosure;

[0049] FIG14 is a schematic structural diagram of a filter unit according to an embodiment of the present disclosure;

[0050] FIG15 is a schematic diagram of an NNLF filter according to an embodiment of the present disclosure;

[0051] FIG16 is a schematic diagram of residual value adjustment according to an embodiment of the present disclosure.

[0052] Details

[0053] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure.

[0054] In the description of the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment described as "exemplary" or "for example" in the present disclosure should not be interpreted as being more preferred or advantageous than other embodiments. "And / or" in this article is a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "Multiple" refers to two or more than two. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present disclosure, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0055] When describing representative exemplary embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific sequence of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art can readily understand that these sequences can vary and still remain within the spirit and scope of the disclosed embodiments.

[0056] The neural network-based loop filtering method and video coding and decoding method of the disclosed embodiments can be applied to various video coding and decoding standards, such as: H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), AVS (Audio Video coding Standard), and other standards developed by MPEG (Moving Picture Experts Group), AOM (Alliance for Open Media), JVET (Joint Video Experts Team) and extensions of these standards, or any other customized standards.

[0057] Figure 1A is a block diagram of a video encoding and decoding system that can be used in embodiments of the present disclosure. As shown, the system is divided into an encoder 1 and a decoder 2. Encoder 1 generates a bitstream. Decoder 2 can decode the bitstream. Decoder 2 can receive the bitstream from encoder 1 via link 3. Link 3 includes one or more media or devices capable of moving the bitstream from encoder 1 to decoder 2. In one example, link 3 includes one or more communication media that enable encoder 1 to send the bitstream directly to decoder 2. Encoder 1 modulates the bitstream according to a communication standard and sends the modulated bitstream to decoder 2. The one or more communication media may include wireless and / or wired communication media and may form part of a packet network. In another example, the bitstream can also be output from output interface 15 to a storage device, and decoder 2 can read the stored data from the storage device via streaming or downloading.

[0058] As shown in the figure, encoding end 1 includes a data source 11, a video encoding device 13, and an output interface 15. Data source 11 may include a video capture device (e.g., a camera), an archive containing previously captured data, a feed interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. Video encoding device 13, also known as a video encoder, encodes data from data source 11 and outputs it to output interface 15. Output interface 15 may include at least one of a modulator, a modem, and a transmitter. Decoding end 2 includes an input interface 21, a video decoding device 23, and a display device 25. Input interface 21 includes at least one of a receiver and a modem. Input interface 21 can receive a bitstream via link 3 or from a storage device. Video decoding device 23, also known as a video decoder, decodes the received bitstream. Display device 25 displays the decoded data. Display device 25 may be integrated with other devices in decoding end 2 or provided separately; display device 25 is optional for the decoding end. In other examples, the decoding end may include other devices or equipment for applying the decoded data.

[0059] FIG1B is a block diagram of an exemplary video encoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video encoding device 10 includes:

[0060] The division unit 101 is configured to cooperate with the prediction unit 100 to divide the received video data into slices, coding tree units (CTUs) or other larger units. The received video data may be a video sequence including video frames such as I frames, P frames or B frames.

[0061] The prediction unit 100 is configured to divide a CTU into coding units (CUs) and perform intra-frame prediction coding or inter-frame prediction coding on the CU. When performing intra-frame prediction and inter-frame prediction on the CU, the CU can be divided into one or more prediction units (PUs).

[0062] The prediction unit 100 includes an inter-frame prediction unit 121 and an intra-frame prediction unit 126 .

[0063] The inter-frame prediction unit 121 is configured to perform inter-frame prediction on the PU and generate prediction data for the PU, wherein the prediction data includes the prediction block of the PU, the motion information of the PU, and various syntax elements. The inter-frame prediction unit 121 may include a motion estimation (ME) unit and a motion compensation (MC) unit. The motion estimation unit can be used to perform motion estimation to generate a motion vector, and the motion compensation unit can be used to obtain or generate a prediction block based on the motion vector.

[0064] The intra prediction unit 126 is configured to perform intra prediction on a PU and generate prediction data for the PU. The prediction data for the PU may include a prediction block and various syntax elements of the PU.

[0065] The residual generating unit 102 (indicated by the circle with a plus sign after the dividing unit 101 in the figure) is configured to subtract the prediction block of the PU into which the CU is divided from the original block of the CU to generate a residual block of the CU.

[0066] The transform processing unit 104 is configured to partition a CU into one or more transform units (TUs). The partitioning of prediction units and transform units may be different. A TU-associated residual block is a sub-block obtained by partitioning the residual block of the CU. A TU-associated coefficient block is generated by applying one or more transforms to the TU-associated residual block.

[0067] The quantization unit 106 is configured to quantize the coefficients in the coefficient block based on a quantization parameter. The quantization degree of the coefficient block can be changed by adjusting the quantization parameter (QP: Quantizer Parameter).

[0068] The inverse quantization unit 108 and the inverse transform unit 110 are configured to apply inverse quantization and inverse transform to the coefficient block, respectively, to obtain a reconstructed residual block associated with the TU.

[0069] The reconstruction unit 112 (represented by the circle with a plus sign after the inverse transform processing unit 110 in the figure) is configured to add the reconstructed residual block and the prediction block generated by the prediction unit 100 to generate a reconstructed image.

[0070] The filter unit 113 is configured to perform loop filtering on the reconstructed image.

[0071] The decoded image buffer 114 is configured to store the reconstructed image after loop filtering. The intra prediction unit 126 can extract reference images of blocks adjacent to the current block from the decoded image buffer 114 to perform intra prediction. The inter prediction unit 121 can use the reference image of the previous frame cached in the decoded image buffer 114 to perform inter prediction on the PU of the current frame image.

[0072] The entropy coding unit 115 is configured to perform entropy coding operations on received data (such as syntax elements, quantized coefficient blocks, motion information, etc.) to generate a video bitstream.

[0073] In other examples, the video encoding apparatus 10 may include more, fewer, or different functional components than those in this example, for example, the transform processing unit 104 and the inverse transform processing unit 110 may be eliminated.

[0074] FIG1C is a block diagram of an exemplary video decoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video decoding device 15 includes:

[0075] Entropy decoding unit 150 is configured to perform entropy decoding on the received encoded video stream, extracting syntax elements, quantized coefficient blocks, and motion information for PUs. Prediction unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, and filter unit 159 may each perform corresponding operations based on the syntax elements extracted from the stream.

[0076] The inverse quantization unit 154 is configured to perform inverse quantization on the coefficient block associated with the quantized TU.

[0077] The inverse transform processing unit 156 is configured to apply one or more inverse transforms to the inverse quantized coefficient block to generate a reconstructed residual block for the TU.

[0078] Prediction unit 152 includes an inter-prediction unit 162 and an intra-prediction unit 164. If the current block is encoded using intra prediction, intra prediction unit 164 determines an intra prediction mode for the PU based on syntax elements decoded from the codestream, and performs intra prediction in conjunction with reconstructed reference information adjacent to the current block obtained from decoded image buffer 160. If the current block is encoded using inter prediction, inter prediction unit 162 determines a reference block for the current block based on motion information of the current block and corresponding syntax elements, and performs inter prediction on the reference block obtained from decoded image buffer 160.

[0079] The reconstruction unit 158 ​​(represented by a circle with a plus sign after the inverse transform processing unit 155 in the figure) is set to perform intra-frame prediction or inter-frame prediction on the current block based on the reconstructed residual block associated with the TU and the prediction unit 152 to obtain a reconstructed image.

[0080] The filter unit 159 is configured to perform loop filtering on the reconstructed image.

[0081] The decoded image buffer 160 is configured to store the reconstructed image after loop filtering as a reference image for subsequent motion compensation, intra-frame prediction, inter-frame prediction, etc. The filtered reconstructed image can also be output as decoded video data for presentation on a display device.

[0082] In other embodiments, the video decoding device 15 may include more, fewer, or different functional components. For example, the inverse transform processing unit 155 may be eliminated in some cases.

[0083] Herein, the current block may be a block-level coding unit such as a current coding tree unit (CTU), a current coding unit (CU), or a current prediction unit (PU) in the current image.

[0084] Based on the above-described video encoding and decoding devices, the following basic encoding and decoding process can be performed. On the encoding side, a frame of an image is divided into blocks. Intra-frame prediction, inter-frame prediction, or other algorithms are performed on the current block to generate a predicted block for the current block. The predicted block is subtracted from the original block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain quantized coefficients, and the quantized coefficients are entropy encoded to generate a bitstream. On the decoding side, intra-frame prediction or inter-frame prediction is performed on the current block to generate a predicted block for the current block. The quantized coefficients obtained from the decoded bitstream are then inversely quantized and inversely transformed to obtain a residual block. The predicted block and residual block are added together to obtain a reconstructed block. The reconstructed block is then loop-filtered on the reconstructed image based on an image or block basis to obtain a decoded image. The encoding side also performs similar operations as the decoding side to obtain a decoded image, also known as a reconstructed image after loop filtering. The reconstructed image after loop filtering can be used as a reference frame for inter-frame prediction of subsequent frames. Block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode and parameter information determined by the encoding side can be written into the bitstream. The decoding end determines the block division information, prediction, transformation, quantization, entropy coding, loop filtering and other mode information and parameter information used by the encoding end by decoding the code stream or analyzing according to the setting information, so as to ensure that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end.

[0085] Although the above is an example of a block-based hybrid coding framework, the embodiments of the present disclosure are not limited thereto. With the development of technology, one or more modules in the framework and one or more steps in the process can be replaced or optimized.

[0086] The embodiments of the present disclosure relate to, but are not limited to, the filter unit (the filter unit may also be referred to as a loop filtering module) in the above-mentioned encoding end and decoding end and the corresponding loop filtering method.

[0087] In one embodiment, the filter units at the encoding and decoding ends include tools such as a deblocking filter (DBF) 20, a sample adaptive offset filter (SAO) 22, and an adaptive loop filter (ALF) 26. Between the SAO and ALF, a neural network-based loop filter (NNLF) 26 is also included, as shown in Figure 2. The filter units perform loop filtering on the reconstructed image to compensate for distortion and provide a better reference for subsequent pixel encoding.

[0088] In an exemplary embodiment, a neural network-based loop filtering NNLF solution is provided, and the model used adopts the filtering network shown in Figure 3A. The NNLF is denoted as NNLF1 in the text, and the filter that executes NNLF1 is called NNLF1 filter. As shown in the figure, the backbone network (backbone) of the filtering network includes a plurality of residual blocks (ResBlock) connected in sequence, and also includes a convolution layer (represented by Conv in the figure), an activation function layer (ReLU in the figure), a merging (concat) layer (represented by Cat in the figure), and a pixel reorganization layer (represented by PixelShuffle in the figure). The structure of each residual block is shown in Figure 3B, including a convolution layer with a convolution kernel size of 1×1, a ReLU layer, a convolution layer with a convolution kernel size of 1×1, and a convolution layer with a convolution kernel size of 3×3 connected in sequence.

[0089] As shown in Figure 3A, the input of the NNLF1 filter includes the luminance information (i.e., Y component) and chrominance information (i.e., U component and V component) of the reconstructed image (rec_YUV), as well as a variety of auxiliary information, such as the luminance information and chrominance information of the predicted image (pred_YUV), QP information, and frame type information. The QP information includes the default baseline quantization parameter (BaseQP: Base Quantization Parameter) in the encoding profile and the slice quantization parameter (SliceQP: Slice Quantization Parameter) of the current slice. The frame type information includes the slice type (SliceType), that is, the type of frame to which the current slice belongs. The output of the model is the filtered image (output_YUV) after NNLF1 filtering. The filtered image output by the NNLF1 filter can also be used as the reconstructed image input to subsequent filters.

[0090] NNLF1 uses a model to filter the YUV components of the reconstructed image (rec_YUV) and output the YUV components of the filtered image (out_YUV), as shown in Figure 3C. Auxiliary input information such as the YUV components of the predicted image are omitted in this figure. The filtering network of this model has a skip connection branch between the input reconstructed image and the output filtered image, as shown in Figure 3A.

[0091] Another exemplary embodiment provides another NNLF scheme, denoted as NNLF2. NNLF2 uses two models, one model is used to filter the luminance component of the reconstructed image, and the other model is used to filter the two chrominance components of the reconstructed image. The two models can use the same filtering network, and there is also a jump connection branch between the reconstructed image input to the NNLF2 filter and the filtered image output by the NNLF2 filter. As shown in Figure 4A, the backbone network of the filtering network includes a plurality of residual blocks (AttRes Block) with attention mechanisms connected in sequence, a convolutional layer (Conv 3×3) for implementing feature mapping, and a reorganization layer (Shuffle). The structure of each residual block with an attention mechanism is shown in Figure 4B, including a convolutional layer (Conv 3×3), an activation layer (PReLU), a convolutional layer (Conv 3×3) and an attention layer (Attintion) connected in sequence. M represents the number of feature maps, and N represents the number of samples in one dimension.

[0092] Model 1 of NNLF2 for filtering the luminance component of the reconstructed image is shown in Figure 4C. Its input information includes the luminance component of the reconstructed image (rec_Y), and the output is the luminance component of the filtered image (out_Y). Auxiliary input information such as the luminance component of the predicted image is omitted in the figure. Model 2 of NNLF2 for filtering the two chrominance components of the reconstructed image is shown in Figure 4C. Its input information includes the two chrominance components of the reconstructed image (rec_UV) and the luminance component of the reconstructed image as auxiliary input information (rec_Y). The output of model 2 is the two chrominance components of the filtered image (out_UV). Model 1 and model 2 can also include other auxiliary input information, such as QP information, block partition image, deblocking filter boundary strength information, etc.

[0093] The above-mentioned NNLF1 and NNLF2 schemes can be implemented in neural network based video coding (NNVC: Neural Network based Video Coding) using neural network based common software (NCS: Neural Network based Common Software), serving as the baseline tool in the NNVC reference software testing platform, namely baseline NNLF.

[0094] In the field of deep learning, the concept of residual learning has been proposed. Through a simple skip connection structure from the input to the output, the network focuses on learning the residual information of the image, improving the network's learning ability and prediction performance. The basic structure of the residual network (ResNet) is shown in Figure 5. NNLF1 and NNLF2 draw on the concept of residual learning. See Figure 5. Their filtering networks include a neural network (NN) and a skip connection branch from the reconstructed image of the input filter to the filtered image output by the filter. The filtered image output by NNLF1 and NNLF2 is cnn = rec + res, where rec represents the input reconstructed image and res represents the residual image output by the neural network. The neural network includes the other parts of the filtering network except the skip connection branch mentioned above, and the neural network has the function of predicting residual information. NNLF1 and NNLF2 use a neural network to predict the residual information of the input reconstructed image relative to the original image, that is, the residual image, and then superimpose the residual image on the input reconstructed image (that is, add it to the reconstructed image) to obtain the filtered image output by the filter, which can make the quality of the filtered image closer to the original image.

[0095] In video coding, inter-frame prediction technology allows the current frame to reference the image information of the previous frame, improving coding performance. However, the coding performance of the previous frame also affects the coding performance of subsequent frames. In the NNLF1 and NNLF2 schemes, to adapt the filter network to the influence of inter-frame prediction technology, the model training process includes an initial training phase and an iterative training phase, using multiple rounds of training. In the initial training phase, the model to be trained is not yet deployed in the encoder. The model is trained in the first round using sample data of reconstructed images, resulting in a trained model. In the iterative training phase, the model is deployed in the encoder. The trained model is first deployed in the encoder, and sample data of reconstructed images is re-collected. The trained model is then trained in the second round, resulting in a trained model. The trained model is then deployed in the encoder again, and sample data of reconstructed images is re-collected. The trained model is then trained in the third round, resulting in a trained model. The training process repeats itself. Finally, each trained model is tested on a validation set to identify the model with the best coding performance for deployment.

[0096] However, this multi-round training operation still lags behind the encoding test. Analysis is as follows: Figure 6A is a schematic diagram of the N+1 round of training. As shown in the figure, during the N+1 round of training, the model model_N trained after the Nth round is deployed in the encoder, and training data of multiple frames of reconstructed images is collected. The boxes labeled 0, 1, 2, ... in the figure represent the reconstructed images of the first, second, third, and so on frames. The model model_N+1 trained after the N+1 round is obtained through training. Assuming that the encoding test shows the best performance of model_N+1, training is complete.

[0097] When model_N+1 is subjected to coding test, model model_N+1 is deployed in the encoder or decoder. As shown in FIG6B , the preceding frame referenced by the current frame using inter-frame prediction coding is generated based on loop filtering of model_N+1, and there is a lag in training relative to testing. However, for model_N+1, its applicable environment is the environment during the N+1 round of training, and the preceding frame referenced by its current frame is loop filtered using model_N, which is different from the environment in which model_N+1 is subjected to coding test. Since the performance of model_N+1 is better than that of model_N, during the coding test, the performance of the preceding frame is further improved after the preceding frame referenced by the current frame is filtered using model_N+1. This improves the quality of the reconstructed image input when model_N+1 is subjected to coding test (the residual with the original image becomes smaller), which is different from the quality expected in the training environment. However, model_N+1 still predicts the residual according to its trained ability, which may cause the residual output by the neural network in model_N+1 to be too large. Currently, there is no solution to try to adjust this residual.

[0098] In this article, for the residual value of the residual image, residual adjustment that reduces the residual of the residual image means that the residual value in the residual image is closer to 0, that is, the absolute value of the residual value becomes smaller, such as 3 becomes 2, -3 becomes -2, and does not mean a change such as -3 becomes -4. The reduction of the residual refers to the entire residual image, which can be that the absolute value of the residual value of some pixels becomes smaller, while the residual value of some pixels remains unchanged, except for the residual value of zero. It can be that the absolute value of all non-zero residual values ​​becomes smaller, or the absolute value of some non-zero residual values ​​becomes smaller. For example, the residual values ​​in the value range [1,2] and [-1,-2] can remain unchanged, while the absolute value of the residual values ​​greater than or equal to 3 and less than or equal to -3 becomes smaller.

[0099] An embodiment of the present disclosure provides a neural network-based loop filtering method, which is applied to a NNLF filter at an encoding end. The NNLF filter includes a neural network and a jump connection branch from the input to the output of the NNLF filter. The method includes:

[0100] Step S110: Input the reconstructed image into the neural network to obtain the residual image output by the neural network.

[0101] Step S120: Calculate the rate - distortion cost cost1 of performing NNLF on the reconstructed image in the first mode, and the rate - distortion cost cost2 of performing NNLF on the reconstructed image in the second mode.

[0102] Among them, the first mode is the NNLF mode without residual adjustment for the residual image, and the second mode is the NNLF mode with residual adjustment for the residual image.

[0103] Step S130: In the case of cost1 < cost2, select the first mode to perform NNLF on the reconstructed image; in the case of cost2 < cost1, select the second mode to perform NNLF on the reconstructed image; in the case of cost1 = cost2, select the first mode or the second mode to perform NNLF on the reconstructed image.

[0104] In this embodiment of the loop filtering method based on a neural network, the encoding end can select a mode with a smaller rate - distortion cost from the mode with residual adjustment and the mode without residual adjustment to perform NNLF, compensating to a certain extent for the performance loss caused by the lag of the training of the NNLF mode relative to the encoding test, thereby improving the effect of NNLF filtering and enhancing the encoding performance.

[0105] Unless otherwise specified, the residual adjustment herein refers to the residual adjustment made to the residual image when performing loop filtering based on a neural network on the reconstructed image.

[0106] In an exemplary embodiment of the present disclosure, the reconstructed image is the reconstructed image of the current frame or the current slice or the current block, but it can also be the reconstructed image of other coding units. The reconstructed image for performing the NNLF filter herein can be coding units at different levels such as image - level (including frames, slices), block - level, etc.

[0107] In an exemplary embodiment of the present disclosure, the residual adjustment makes the residual in the residual image smaller.

[0108] In an exemplary embodiment of the present disclosure, the calculation of the rate-distortion cost cost1 of performing NNLF on the reconstructed image using the first mode includes: adding the residual image to the reconstructed image to obtain a first filtered image; and, calculating the cost1 based on the difference between the first filtered image and the corresponding original image; when the reconstructed image and the residual image both include 3 components, such as the Y component, the U component, and the V component. The cost1 can be obtained by calculating the sum of squared errors (SSD) of the first filtered image and the original image on the 3 components: and weighting the SSD on the 3 components and adding them together. In this embodiment, the selection of the first mode to perform NNLF on the reconstructed image includes: using the first filtered image obtained by adding the residual image to the reconstructed image as the filtered image output after performing NNLF on the reconstructed image. When the first mode is used to perform NNLF on the reconstructed image in this embodiment, the above-mentioned NNLF1 or NNLF2 filtering method can be used, or other filtering methods that do not perform residual adjustment on the residual image can be used.

[0109] In an exemplary embodiment of the present disclosure, the calculating a rate-distortion cost cost2 of performing NNLF on the reconstructed image using the second mode includes:

[0110] According to each residual adjustment method set, the residual image is subjected to residual adjustment and the residual image is added to the reconstructed image to obtain a second filtered image, and a rate-distortion cost is calculated based on the difference between the second filtered image and the original image; and

[0111] The minimum rate-distortion cost among all the calculated rate-distortion costs is taken as cost2.

[0112] In one example of this embodiment, if there is only one residual adjustment method, a rate-distortion cost is calculated, and this rate-distortion cost is cost2. In another example of this embodiment, if there are multiple residual adjustment methods, assuming there are two, two rate-distortion costs are calculated, and the minimum of the two rate-distortion costs is used as cost2.

[0113] In one example of this embodiment, residual adjustment is performed on the residual image and added to the reconstructed image. This can be done by performing residual adjustment on the residual image and then adding the result of the residual adjustment to the reconstructed image. The residual adjustment method, for example, is to subtract 1 from positive residual values ​​in the residual image and add 1 to negative residual values. That is, residual values ​​that are 0 are not adjusted, so that the residual in the residual image is reduced overall. However, in a specific implementation, the calculations do not necessarily follow this order. For example, the residual image and the reconstructed image are first added, and the resulting image also contains the residual image. Residual adjustment can also be performed on the residual image. For example, for any pixel in the image, assuming that the value of the pixel in the residual image (i.e., the residual value) is x and the value of the pixel in the reconstructed image (i.e., the reconstructed value) is y, and the residual adjustment is to subtract 1 from the residual value of the pixel. When calculating the value of the pixel in the second filtered image, whether x is subtracted by 1 and then added to y, or x is added to y and then subtracted by 1, the result is the same. The same is true for other embodiments of the present disclosure, including embodiments at the decoding end, in the specific implementation of performing residual adjustment on the residual image (or its components) and adding the residual image (or its components) to the reconstructed image (or its components).

[0114] In an example of this embodiment, the reconstructed image and the residual image both include three components, and the rate-distortion cost is calculated based on the difference between the second filtered image and the original image, including: calculating the square error and SSD of the second filtered image and the original image on the three components, and then weighting the SSD on the three components and adding them to obtain the rate-distortion cost.

[0115] In an exemplary embodiment of the present disclosure, the set residual adjustment method includes one or more of the following types of residual adjustment methods:

[0116] A fixed value is added to or subtracted from the non-zero residual value in the residual image to reduce the absolute value of the non-zero residual value; for example, 1 is subtracted from the positive residual value in the residual image and 1 is added to the negative residual value.

[0117] The non-zero residual values ​​in the residual image are added or subtracted from the corresponding adjustment value according to the interval in which they are located, so that the absolute value of the non-zero residual value becomes smaller. There are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value. For example, the residual values ​​in the residual image that are greater than or equal to 1 and less than or equal to 5 are subtracted by 1, the residual values ​​that are greater than 5 are subtracted by 2, the residual values ​​that are less than or equal to -1 and greater than or equal to -5 are added by 1, and the residual values ​​that are less than -5 are added by 2.

[0118] The above-mentioned embodiment improves the coding performance by adjusting the residual information output by the filtering network. As mentioned above, the residual of the residual image output by the neural network may be too large. The residual adjustment can reduce the residual and help improve the coding performance. For a residual image, the residual value of each pixel point therein may be positive or negative. When reducing the residual, the positive residual value can be subtracted from the fixed value (positive number), and the negative residual value can be added to the fixed value. When the residual value is 0, no adjustment is made, so that the overall residual value becomes smaller, that is, closer to 0. As shown in Figure 16, assuming that the fixed value is (+1), the residual value of each pixel point in the original residual image is shown on the left, and the residual value of each pixel point in the residual image after residual adjustment is shown on the right. The residual image after residual adjustment is superimposed on the input reconstructed image to obtain a filtered image, which can also be called a reconstructed image after NNLF. Multiple fixed values ​​can be set to correspond to multiple residual adjustment methods. By calculating the rate-distortion cost under each residual adjustment method, the fixed value used by the residual adjustment method with the smallest rate-distortion cost is selected, and the index corresponding to the residual adjustment method is encoded into the bitstream for reading and processing by the decoding end.

[0119] In addition to using a fixed value residual adjustment method, other types of residual adjustment methods can also be used. For example, the residual value can be segmented according to its size, and adjustment operations with different precisions can be tried. For example, for a residual value with a larger absolute value, an adjustment value with a larger absolute value is set; for a residual value with a smaller absolute value, an adjustment value with a smaller absolute value is set.

[0120] An example pseudo code is as follows:

[0121] Assume that for the current pixel point of the current frame, the corresponding residual value is res, and the adjustment value to be determined is RO_FACTOR. The specific adjustment value derivation strategy is as follows.

[0122] if (res=0) RO_FACTOR=0;

[0123] else if(0 <res<=x1) RO_FACTOR=a1;

[0124] else if(x1 <res<=x2) RO_FACTOR=a2;

[0125] else if(x2 <res) RO_FACTOR=a3;

[0126] else if(y1<=res<0) RO_FACTOR=b1;

[0127] else if(y2<=res <y1) RO_FACTOR=b2;

[0128] else if(res <y2) RO_FACTOR=b3;

[0129] Among them, {x1, x2, x3} represents a positive residual value, {y1, y2, y3} represents a negative residual value, and {a1, a2, a3} and {b1, b2, b3} are preset candidate fixed values.

[0130] The above scheme does not adjust the residual value of zero. For non-zero residual values, it finds the interval it falls into (a total of 6 intervals are set) to determine the adjustment value to be used.

[0131] In this embodiment, the set multiple residual adjustment methods may include one type of residual adjustment method, or may include multiple types of residual adjustment methods.

[0132] The above embodiment uniformly adjusts all three components of the residual image using the same residual adjustment method. This residual adjustment method achieves the optimal overall result for all three components, but it may not be the optimal residual adjustment method for a specific component in the residual image. To further optimize coding performance, the decision to perform residual adjustment and the appropriate residual adjustment method can be made for each component individually.

[0133] An embodiment of the present disclosure provides a neural network-based loop filtering method, which is applied to a NNLF filter at an encoding end. The NNLF filter includes a neural network and a jump connection branch from the input to the output of the NNLF filter. As shown in FIG8 , the method includes:

[0134] Step S210, inputting the reconstructed image into the neural network to obtain a residual image output by the neural network; the reconstructed image and the residual image both include three components, such as a Y component, a U component, and a V component; and

[0135] In step S210, the following processing is performed on each of the three components, which may be referred to as mode selection processing:

[0136] Calculating a rate-distortion cost cost1 of performing NNLF on the component of the reconstructed image using a first mode, and a rate-distortion cost cost2 of performing NNLF on the component of the reconstructed image using a second mode; the first mode is an NNLF mode in which residual adjustment is not performed on the component of the residual image, and the second mode is an NNLF mode in which residual adjustment is performed on the component of the residual image;

[0137] When cost1 < cost2, select the first mode to perform NNLF on this component of the reconstructed image; when cost2 < cost1, select the second mode to perform NNLF on this component of the reconstructed image; when cost1 = cost2, select the first mode or the second mode to perform NNLF on the reconstructed image.

[0138] In this embodiment, the mode selection process can be performed separately for each component, which can further optimize the coding performance based on the foregoing embodiment of uniformly adjusting the three components. Since the relevant operations are performed at the output end of the NNLF filter, the impact on the computational complexity is small.

[0139] In an exemplary embodiment of the present disclosure, the reconstructed image is the reconstructed image of the current frame or the current slice or the current block.

[0140] In an exemplary embodiment of the present disclosure, the residual adjustment makes the residual in the component of the residual image smaller.

[0141] In an exemplary embodiment of the present disclosure, calculating the rate-distortion cost cost1 of performing NNLF on this component of the reconstructed image in the first mode includes: adding this component of the residual image to this component of the reconstructed image to obtain the filtered component; and calculating the cost1 according to the difference between the filtered component and this component of the corresponding original image. In one example, the difference is represented by SSD, that is, the SSD between the filtered component and this component of the corresponding original image is used as the cost1. In other examples, the difference can also be represented by other metrics such as mean squared error (MSE) and mean absolute error (MAE). The present disclosure is not limited thereto. The same applies to other embodiments of the present disclosure.

[0142] In an example of this embodiment, selecting the first mode to perform NNLF on this component of the reconstructed image includes using the filtered component obtained by adding this component of the residual image to this component of the reconstructed image as this component of the filtered image output after performing NNLF on the reconstructed image. The component of the filtered image obtained by performing NNLF on this component of the reconstructed image according to the first mode in this embodiment does not perform residual adjustment on this component of the residual image. Therefore, this component of the filtered image can be the same as this component of the filtered image obtained by the NNLF scheme without residual adjustment (such as NNLF1 and NNLF2).

[0143] In an exemplary embodiment of the present disclosure, the calculating of the rate-distortion cost cost2 of performing NNLF on the component of the reconstructed image using the second mode includes:

[0144] According to each residual adjustment mode set for the component, the residual adjustment is performed on the component of the residual image and the component is added to the component of the reconstructed image to obtain a filtered component; and a rate-distortion cost of the component is calculated based on the difference between the filtered component and the component of the corresponding original image; for example, the SSD between the filtered component and the component of the corresponding original image is used as the rate-distortion cost of the component; and

[0145] The minimum rate-distortion cost among all rate-distortion costs of the component is calculated as cost2 of the component; wherein one or more residual adjustment methods are set for the component.

[0146] In an example of this embodiment, selecting the second mode to perform NNLF on the component of the reconstructed image includes:

[0147] The component of the residual image is subjected to residual adjustment in accordance with the residual adjustment method corresponding to the cost2 of the component, and the filtered component obtained by adding the residual adjustment to the component of the reconstructed image is used as the component of the filtered image output after NNLF is performed on the reconstructed image.

[0148] In an example of this embodiment, the residual adjustment is performed on the component of the residual image and the component is added to the component of the reconstructed image. The residual adjustment can be performed on the component of the residual image and then the result of the residual adjustment is added to the component of the reconstructed image. However, in specific implementation, it is not necessary to calculate in this order.

[0149] In an example of this embodiment, the residual adjustment methods set for the three components are the same or different; for example, the residual adjustment method set for the Y component is: subtract 1 from the positive residual value in the residual image, and add 1 to the negative residual value; and the residual adjustment method set for the U component and the V component is: subtract 1 from the residual value in the residual image that is greater than 2, and add 1 to the residual value that is less than -2.

[0150] In an example of this embodiment, the residual adjustment method set for at least one of the three components includes one or more of the following types of residual adjustment methods:

[0151] Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller;

[0152] The non-zero residual value in the residual image is added or subtracted by the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

[0153] An embodiment of the present disclosure further provides a video encoding method, which is applied to a video encoding device, including: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image, as shown in FIG8 :

[0154] Step S310, when NNLF allows residual adjustment, performing NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure;

[0155] Step S320 : Encode a residual adjustment usage flag of the reconstructed image to indicate whether residual adjustment is required when performing NNLF on the reconstructed image.

[0156] When performing neural network-based loop filtering on the reconstructed image, the embodiment of the present disclosure can choose to adjust or not adjust the residual image according to the rate-distortion cost, which can compensate for the performance loss caused by the lag of NNLF mode training relative to the coding test and improve the coding performance.

[0157] In an exemplary embodiment of the present disclosure, the residual adjustment usage flag is a picture-level syntax element or a block-level syntax element.

[0158] In an exemplary embodiment of the present disclosure, it is determined that the NNLF allows residual adjustment when one or more of the following conditions are met:

[0159] Decoding a sequence-level residual adjustment permission flag, and determining NNLF residual adjustment permission according to the value of the residual adjustment permission flag;

[0160] The residual adjustment permission flag of the decoded picture level is determined, and the NNLF allows residual adjustment according to the value of the residual adjustment permission flag.

[0161] In addition to the above conditions, other conditions may be added, for example, the frame containing the input reconstructed image being an inter-frame coded frame is a necessary condition for NNLF to allow residual adjustment, etc.

[0162] In other embodiments of the present disclosure, the residual adjustment of NNLF may also be enabled all the time. In this case, there is no need to judge by a flag, and the residual adjustment of NNLF is allowed by default.

[0163] In an exemplary embodiment of the present disclosure, the method further includes: if it is determined that NNLF does not allow residual adjustment, skipping encoding the residual use flag, adding the reconstructed image input to the neural network and the residual image output by the neural network, to obtain a filtered image output after performing NNLF on the reconstructed image. In other words, in this case, the reconstructed image can be filtered using NNLF without residual adjustment.

[0164] In an exemplary embodiment of the present disclosure, the method is to perform NNLF on the reconstructed image according to any of the above-mentioned embodiments of uniformly performing residual adjustment on the three components of the residual image; the number of flags roflag used for the residual adjustment of the reconstructed image is 1; when the first mode is selected to perform NNLF on the reconstructed image, the roflag is set to a value indicating that no residual adjustment is required, such as 0; when the second mode is selected to perform NNLF on the reconstructed image, the roflag is set to a value indicating that residual adjustment is required, such as 1.

[0165] In an example of this embodiment, the method further includes: when the roflag is set to a value indicating that residual adjustment is required, and there are multiple residual adjustment methods set, continuing to encode a residual adjustment method index of the reconstructed image, wherein the residual adjustment method index is used to indicate the residual adjustment method based on which the residual adjustment is performed. For example, when there are three residual adjustment methods set, the residual adjustment method index can be a 2-bit flag, and when the values ​​of the flag are 0, 1, and 2, they represent the three residual adjustment methods respectively. The correspondence between the value and the residual adjustment method is pre-agreed on the encoding end and the decoding end, for example, as defined in a standard or protocol.

[0166] This embodiment uses two flags, namely the residual adjustment usage flag and the residual adjustment method index, to respectively indicate whether residual adjustment is required and the residual adjustment method based on which the residual adjustment is performed (when multiple residual adjustment methods are set). However, in another exemplary embodiment of the present disclosure, when there are multiple residual adjustment methods set, the residual adjustment usage flag is also used to indicate the residual adjustment method based on which the residual adjustment is performed, that is, this embodiment uses the residual adjustment usage flag to simultaneously indicate whether residual adjustment is required and the residual adjustment method based on which the residual adjustment is performed. For example, when there are three residual adjustment methods set, a 2-bit residual adjustment usage flag roflag is used, and the four values ​​of roflag can respectively indicate that residual adjustment is not required, residual adjustment is performed using the first residual adjustment method, residual adjustment is performed using the second residual adjustment method, and residual adjustment is performed using the third residual adjustment method.

[0167] If three residual adjustment methods are set and residual adjustment is not required, the embodiment using two flags only needs to encode a 1-bit flag, the residual adjustment use flag, and does not need to encode the residual adjustment method index. However, the embodiment using one flag needs to encode a 2-bit residual adjustment use flag. If three residual adjustment methods are set and residual adjustment is required, the embodiment using two flags needs to encode a 1-bit residual adjustment use flag and a 2-bit residual adjustment method index, while the embodiment using one flag needs to encode a 2-bit residual adjustment use flag.

[0168] In an exemplary embodiment of the present disclosure, the method is to perform NNLF on the reconstructed image according to any one of the above-mentioned embodiments of performing residual adjustment on the three components of the residual image respectively; the number of residual adjustment flags roflag(j) used for the reconstructed image is 3, j=1, 2, 3, and roflag(j) is used to indicate whether residual adjustment is required when performing NNLF on the j-th component of the reconstructed image; when the first mode is selected to perform NNLF on the j-th component of the reconstructed image, roflag(j) is set to a value indicating that residual adjustment is not required, such as 0; when the second mode is selected to perform NNLF on the j-th component of the reconstructed image, roflag(j) is set to a value indicating that residual adjustment is required, such as 1.

[0169] In an example of this embodiment, the method further includes: when the roflag(j) is set to a value indicating that residual adjustment is required and there are multiple residual adjustment methods set for the j-th component, continuing to encode the residual adjustment method index index(j) of the j-th component of the reconstructed image to indicate the residual adjustment method based on which the residual adjustment is performed on the j-th component of the residual image.

[0170] In another exemplary embodiment of the present disclosure, when there are multiple residual adjustment methods set for the j-th component, the residual adjustment usage flag of the j-th component is also used to indicate the residual adjustment method based on which the residual adjustment is performed, that is, this embodiment uses the residual adjustment usage flag of the j-th component to simultaneously indicate whether residual adjustment is required and the residual adjustment method based on which the residual adjustment is performed.

[0171] An embodiment of the present disclosure further provides a neural network-based loop filtering method, which is applied to a NNLF filter at a decoding end. The NNLF filter includes a neural network and a jump connection branch from the input to the output of the NNLF filter, as shown in FIG10 . The method includes:

[0172] Step S410, decoding the residual adjustment flag roflag of the reconstructed image, wherein the roflag is used to indicate whether residual adjustment is required when performing NNLF on the reconstructed image;

[0173] Step S420: If it is determined according to the roflag that residual adjustment is not required, the first mode is used to perform NNLF on the reconstructed image; if it is determined according to the roflag that residual adjustment is required, the second mode is used to perform NNLF on the reconstructed image;

[0174] The first mode is an NNLF mode that does not perform residual adjustment on the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the residual image.

[0175] Figure 11 is a schematic diagram of the NNLF performed on the reconstructed image at the decoding end of this embodiment. As shown in the figure, there are two NNLF paths. One is to add the residual image output by the neural network (NN) to the reconstructed image to obtain the filtered image output by the NNLF; the other requires residual adjustment of the residual image (RO in the figure represents the residual adjustment module, which is used to perform residual adjustment) and then add it to the reconstructed image. The choice of these two paths is determined by the decoded residual adjustment usage flag. In this embodiment, mode selection is performed at the NNLF output end, and flag decoding can be performed before obtaining the filtered image, but is not limited to this.

[0176] This embodiment is based on a loop filtering method of a neural network. By using a decoding residual adjustment use flag, a better mode is selected from two modes of residual adjustment and no residual adjustment, which can enhance the filtering effect of NNLF and improve the quality of the decoded image.

[0177] In an exemplary embodiment of the present disclosure, roflag can be a 1-bit flag. The value of roflag can indicate whether residual adjustment is required. For example, when the value of roflag is 1, residual adjustment is determined to be required, while when the value of roflag is 0, residual adjustment is determined not to be required. The same applies to other embodiments of the present disclosure.

[0178] In an exemplary embodiment of the present disclosure, the residual adjustment reduces the residual of the residual image.

[0179] In an exemplary embodiment of the present disclosure, the reconstructed image is a reconstructed image of a current frame, a current slice, or a current block.

[0180] In an exemplary embodiment of the present disclosure, the residual adjustment usage flag is a picture-level syntax element or a block-level syntax element.

[0181] In an exemplary embodiment of the present disclosure, performing NNLF on the reconstructed image using the first mode includes: adding the residual image output by the neural network to the reconstructed image input to the neural network to obtain a filtered image output after performing NNLF on the reconstructed image. Performing NNLF on the reconstructed image using the second mode includes: performing residual adjustment on the residual image according to one of the set residual adjustment methods and adding the residual image to the reconstructed image to obtain a filtered image output after performing NNLF on the reconstructed image.

[0182] In an example of this embodiment, there are multiple set residual adjustment methods, and the residual adjustment of the residual image according to one of the set residual adjustment methods includes: continuing to decode the residual adjustment method index index of the reconstructed image, the index is used to indicate the residual adjustment method based on which the residual adjustment is performed; and, performing residual adjustment on the residual image according to the residual adjustment method indicated by the index.

[0183] This embodiment is based on the use of two flags to indicate whether residual adjustment is required and the residual adjustment method based on which the residual adjustment is based. In another exemplary embodiment, when multiple residual adjustment methods are set, the encoder uses a flag, i.e., the residual adjustment flag roflag, to simultaneously indicate whether residual adjustment is required and the residual adjustment method based on which the residual adjustment is based. In this case, the decoder continues to determine the residual adjustment method based on which the residual adjustment is based based on the roflag of the reconstructed image, and performs residual adjustment on the residual image according to the determined residual adjustment method.

[0184] In an exemplary embodiment of the present disclosure, the set residual adjustment method includes one or more of the following types of residual adjustment methods:

[0185] Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller;

[0186] The non-zero residual value in the residual image is added or subtracted by the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

[0187] An embodiment of the present disclosure further provides a neural network-based loop filtering (NNLF) method, which is applied to an NNLF filter at a decoding end. The NNLF filter includes a neural network and a skip connection branch from the input to the output of the NNLF filter. The method includes performing the following processing on each component of a reconstructed image including three components input to the neural network when performing NNLF, as shown in FIG12 :

[0188] Step S510, decoding the residual adjustment use flag roflag of the component of the reconstructed image, wherein the roflag is used to indicate whether residual adjustment is required when performing NNLF on the component of the reconstructed image;

[0189] Step S520: If it is determined according to the roflag that residual adjustment is not required, perform NNLF on the component of the reconstructed image using the first mode; if it is determined according to the roflag that residual adjustment is required, perform NNLF on the component of the reconstructed image using the second mode;

[0190] The first mode is an NNLF mode that does not perform residual adjustment on the component of the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the component of the residual image.

[0191] This embodiment is based on a loop filtering method of a neural network. By using a decoding residual adjustment flag, for each component, a better mode is selected from two NNLF modes, one with residual adjustment and one without residual adjustment, to perform NNLF on the component. The unified mode selection for multiple components can further enhance the filtering effect of NNLF and improve the quality of the decoded image.

[0192] In an exemplary embodiment of the present disclosure, the residual adjustment reduces the residual in the components of the residual image.

[0193] In an exemplary embodiment of the present disclosure, the reconstructed image is a reconstructed image of a current frame, a current slice, or a current block.

[0194] In an exemplary embodiment of the present disclosure, the residual adjustment usage flag is a picture-level syntax element or a block-level syntax element.

[0195] In an exemplary embodiment of the present disclosure, performing NNLF on the component of the reconstructed image using the first mode includes: adding the component of the residual image to the component of the reconstructed image to obtain the component of the filtered image output after performing NNLF on the reconstructed image;

[0196] The use of the second mode to perform NNLF on the component of the reconstructed image includes: performing residual adjustment on the component of the residual image and adding it to the component of the reconstructed image according to one of the residual adjustment methods set for the component, to obtain the component of the filtered image output after NNLF of the reconstructed image; there are one or more residual adjustment methods set for the component.

[0197] In an example of this embodiment, there are multiple residual adjustment methods set for the component, and the residual adjustment of the component of the residual image according to one of the set residual adjustment methods includes: continuing to decode the residual adjustment method index index of the component of the reconstructed image, the index is used to indicate the residual adjustment method based on which the residual adjustment is performed; and, performing residual adjustment on the component of the residual image according to the residual adjustment method indicated by the index.

[0198] In an example of this embodiment, the image header is shown in the following table:

[0199]

[0200] In the table, ro_enable_flag indicates the sequence-level residual adjustment enable flag. When ro_enable_flag is 1, the following semantics are defined:

[0201] Residual adjustment using flag picture_ro_enable_flag (equivalent to roflag in other embodiments);

[0202] When picture_ro_enable_flag is 1, the following semantics are defined:

[0203] Residual adjustment method index picture_ro_index.

[0204] In the table above, compIdx represents the color component number. For YUV images, it is usually 0 / 1 / 2.

[0205] In other examples, NNLF may be performed in units of blocks (such as CTUs), and in this case, a residual adjustment flag and a residual adjustment mode index are defined as block-level syntax elements.

[0206] This embodiment is based on the use of two flags to respectively indicate whether a component requires residual adjustment and the residual adjustment method based on which the residual adjustment is performed. In another exemplary embodiment, when multiple residual adjustment methods are set for the component, the encoder uses a single flag, namely, a residual adjustment flag, roflag, to simultaneously indicate whether the component requires residual adjustment and the residual adjustment method based on which the residual adjustment is performed. In this case, the decoder determines the residual adjustment method based on the roflag of the component and performs residual adjustment on the component of the residual image according to the determined residual adjustment method.

[0207] In an exemplary embodiment of the present disclosure, the residual adjustment methods set for the three components are the same or different; and the residual adjustment method set for at least one of the three components includes one or more of the following types of residual adjustment methods:

[0208] Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller;

[0209] The non-zero residual value in the residual image is added or subtracted by the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

[0210] An embodiment of the present disclosure further provides a video decoding method, which is applied to a video decoding device, including: performing the following processing when performing neural network-based loop filtering on a reconstructed image, as shown in FIG13 :

[0211] Step S610, determining whether NNLF allows residual adjustment;

[0212] Step S620: When NNLF allows residual adjustment, perform NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure applied to the NNLF filter at the decoding end.

[0213] The video decoding method of this embodiment uses a decoding residual adjustment flag to select a better mode for each component from two NNLF modes, one with residual adjustment and one without residual adjustment, to perform NNLF on the component. The unified mode selection for multiple components can further enhance the filtering effect of NNLF and improve the quality of the decoded image.

[0214] In an exemplary embodiment of the present disclosure, it is determined that the NNLF allows residual adjustment when one or more of the following conditions are met:

[0215] Decoding a sequence-level residual adjustment permission flag, and determining that NNLF allows residual adjustment according to the residual adjustment permission flag;

[0216] The residual adjustment permission flag of the decoded picture level is used to determine whether the NNLF allows residual adjustment according to the residual adjustment permission flag.

[0217] In addition to the above conditions, other conditions may be added, for example, the frame containing the input reconstructed image being an inter-frame coded frame is a necessary condition for NNLF to allow residual adjustment, etc.

[0218] In an example of using the sequence-level residual adjustment permission flag, the sequence header of the video sequence is shown in the following table:

[0219]

[0220] The ro_enable_flag in the table is the sequence-level residual adjustment enable flag.

[0221] In an exemplary embodiment of the present disclosure, the method further includes: when NNLF does not allow residual adjustment, adding the reconstructed image input to the neural network and the residual image output by the neural network to obtain a filtered image output after NNLF is performed on the reconstructed image.

[0222] In an exemplary embodiment of the present disclosure, the NNLF filter is arranged after the deblocking filter or the sample adaptive offset filter and before the adaptive correction filter. In an example of this embodiment, the structure of the filter unit (or loop filter module, see Figures 1B and 1C) is shown in Figure 14, where DBF represents the deblocking filter, SAO represents the sample adaptive offset filter, and ALF represents the adaptive correction filter. NN represents the neural network used for loop filtering in the NNLF filter, which can be the same as the neural network of NNLF filters such as NNLF1 and NNLF2 that do not perform residual adjustment. The NNLF filter also includes a residual offset module (RO) and two jump connection branches from the NN input to the two path outputs, respectively. The RO is used to perform residual adjustment on the residual network output of the neural network. These filters are all components of the filter unit for reconstructing the image. During loop filtering, some or all of the DBF, SAO, and ALF may be disabled. The location where the NNLF filter is deployed is not limited to the location described in this embodiment. It is easy to understand that the implementation of the NNLF method of the present disclosure is not limited to its deployment location. In addition, the filters in the filter unit are not limited to those shown in FIG. 14 , and there may be more or fewer filters, or other types of filters.

[0223] An embodiment of the present disclosure provides a neural network-based loop filtering method. When the encoder performs loop filtering on the reconstructed image, the encoding process is performed in the order of the deployed filters. When entering the NNLF, the following processing is performed:

[0224] The first step is to determine whether residual adjustment is allowed for the current sequence based on the sequence-level residual adjustment enable flag ro_enable_flag. If ro_enable_flag is "1", it means that residual adjustment is allowed for the current sequence, and jump to the second step; if ro_enable_flag is "0", it means that residual adjustment is not allowed for the current sequence, and the process ends (skipping subsequent processing);

[0225] In the second step, the reconstructed image of the current frame is input into the NNLF neural network for prediction. The residual image is obtained from the output of the NNLF and superimposed on the input reconstructed image to obtain the first filtered image.

[0226] The third step is to perform residual adjustment on the residual image and then superimpose it on the input reconstructed image to obtain the second filtered image;

[0227] The fourth step is to compare the first filtered image with the original image of the current frame and calculate the rate-distortion cost C NNLF ; Compare the second filtered image with the original image of the current frame and calculate the rate-distortion cost C RO .

[0228] Step 5: Compare the two costs. If C RO <C NNLF , the second filtered image is used as the filtered image output by the NNLF filter, that is, the second mode is selected to perform NNLF on the reconstructed image; if C RO ≥C NNLF , taking the first filtered image as the filtered image output by the filter, that is, selecting the first mode to perform NNLF on the reconstructed image;

[0229] The calculation formula of the rate-distortion cost in this embodiment is:

[0230] cost=Wy*SSD(Y)+Wu*SSD(U)+Wv*SSD(V)

[0231] Among them, SSD(*) indicates the SSD for a certain color component; Wy, Wu, and Wv represent the weight values ​​of the SSD of the Y component, U component, and V component, respectively, and can be 10:1:1 or 8:1:1.

[0232] The calculation formula for SSD is as follows:

[0233]

[0234] Where M represents the length of the reconstructed image of the current frame, N represents the width of the reconstructed image of the current frame, and rec(x,y) and org(x,y) represent the pixel values ​​of the reconstructed image and the original image at the pixel point (x,y), respectively.

[0235] Step 6: Encode the residual adjustment flag picture_ro_enable_flag and the residual adjustment mode index picture_ro_index of the current frame into the bitstream;

[0236] Step 7: If all blocks in the current frame have been processed, the processing of the current frame is terminated, and the next frame can be loaded for processing. If there are still blocks in the current frame that have not been processed, return to step 2.

[0237] In this embodiment, NNLF processing is performed based on the reconstructed image of the current frame. In other embodiments, NNLF processing may also be performed based on other coding units such as blocks (such as CTU) and slices in the current frame.

[0238] This example uses the NNLF1 baseline tool for comparison. Based on NNLF1, mode selection processing is performed on inter-coded frames (i.e., non-I frames). Two residual adjustment methods using fixed values ​​of 1 and 2 are set. Under the general test conditions of Random Access and Low Delay B, tests are conducted on the common sequence specified by the Joint Video Experts Group (JVET). NNLF1 is used as the anchor for comparison. The results are shown in Tables 1 and 2.

[0239] Table 1: Performance of this example compared to the baseline NNLF1 under the Random Access configuration

[0240]

[0241] Table 2 Performance of this embodiment compared with baseline NNLF1 under Low Delay B configuration

[0242]

[0243]

[0244] The meanings of the parameters in the table are as follows:

[0245] EncT: Encoding Time, 10X% means that after the reference row sorting technology is integrated, the encoding time is 10X% compared with before integration, which means that the encoding time increases by X%.

[0246] DecT: Decoding Time, 10X% means that after the reference row sorting technology is integrated, the decoding time is 10X% compared with before integration, which means that the decoding time increases by X%.

[0247] ClassA1 and ClassA2 are test video sequences with a resolution of 3840x2160, ClassB is a test sequence with a resolution of 1920x1080, ClassC is 832x480, ClassD is 416x240, and ClassE is 1280x720; ClassF is a screen content sequence with several different resolutions.

[0248] Y, U, V are the three color components, and the columns where Y, U, V are located represent the BD-rate of the test results on Y, U, V. rate) indicator, the smaller the value, the better the encoding performance.

[0249] Analyzing the data in the two tables, we can see that by introducing the optimization method of residual adjustment, the coding performance can be further improved on the basis of NNLF1, especially in the chrominance component. The residual adjustment of this embodiment has little effect on the decoding complexity.

[0250] The method of this embodiment can also be used to select the NNLF mode for intra-coded frames (I frames).

[0251] An embodiment of the present disclosure further provides a code stream, wherein the code stream is generated by the video encoding method described in any embodiment of the present disclosure.

[0252] One embodiment of the present disclosure also provides a neural network-based loop filter, as shown in Figure 15 , comprising a processor and a memory storing a computer program. When the processor executes the computer program, it is capable of implementing the neural network-based loop filtering method described in any embodiment of the present disclosure. As shown in the figure, the processor and memory are connected via a system bus. The loop filter may also include other components such as memory and a network interface.

[0253] An embodiment of the present disclosure further provides a video decoding device, see FIG15 , comprising a processor and a memory storing a computer program, wherein the processor can implement the video decoding method as described in any embodiment of the present disclosure when executing the computer program.

[0254] An embodiment of the present disclosure further provides a video encoding device, see FIG15 , comprising a processor and a memory storing a computer program, wherein the processor can implement the video encoding method as described in any embodiment of the present disclosure when executing the computer program.

[0255] The processor of the above-mentioned embodiment of the present disclosure may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a microprocessor, etc., or other conventional processors; the processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), discrete logic or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or other equivalent integrated or discrete logic circuits, or a combination of the above devices. That is, the processor of the above-mentioned embodiment may be any processing device or device combination that implements the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. If the embodiments of the present disclosure are partially implemented in software, the instructions for the software may be stored in a suitable non-volatile computer-readable storage medium, and one or more processors may be used to execute the instructions in hardware to implement the methods of the embodiments of the present disclosure. The term "processor" used herein may refer to the above-mentioned structure or any other structure suitable for implementing the technology described herein.

[0256] An embodiment of the present disclosure further provides a video encoding and decoding system, see FIG1A , which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0257] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, can implement the neural network-based loop filtering method as described in any embodiment of the present disclosure, or implement the video decoding method as described in any embodiment of the present disclosure, or implement the video encoding method as described in any embodiment of the present disclosure.

[0258] In one or more exemplary embodiments above, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium that facilitates the transfer of a computer program from one place to another, such as according to a communication protocol. In this way, a computer-readable medium may generally correspond to a non-transitory tangible computer-readable storage medium or a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this disclosure. A computer program product may include a computer-readable medium.

[0259] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection may also be referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient (transient) media, but rather refer to non-transient tangible storage media. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, or Blu-ray disc, among others, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

Claims

1. A neural network-based loop filtering (NNLF) method, applied to a NNLF filter at a decoding end, wherein the NNLF filter comprises a neural network and a jump connection branch from the input to the output of the NNLF filter. include: The residual adjustment of the decoded reconstructed image uses a flag roflag, wherein roflag is used to indicate whether residual adjustment is required when performing NNLF on the reconstructed image; When it is determined according to the roflag that residual adjustment is not required, the first mode is used to perform NNLF on the reconstructed image; when it is determined according to the roflag that residual adjustment is required, the second mode is used to perform NNLF on the reconstructed image; The first mode is an NNLF mode that does not perform residual adjustment on the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the residual image.

2. The method according to claim 1, Features: The residual adjustment makes the residual of the residual image smaller; The reconstructed image is a reconstructed image of a current frame or a current slice or a current block; the residual adjustment usage flag is a picture-level syntax element or a block-level syntax element.

3. The method according to claim 1, Features: The adopting the first mode to perform NNLF on the reconstructed image includes: adding the residual image output by the neural network to the reconstructed image input to the neural network to obtain a filtered image output after performing NNLF on the reconstructed image; The adopting the second mode to perform NNLF on the reconstructed image includes: performing residual adjustment on the residual image according to one of the set residual adjustment methods and adding the residual image to the reconstructed image to obtain a filtered image output after performing NNLF on the reconstructed image; wherein the set residual adjustment method includes one or more.

4. The method according to claim 3, Features: There are multiple residual adjustment methods set, and the residual adjustment of the residual image according to one of the set residual adjustment methods includes: continuing to decode the residual adjustment method index index of the reconstructed image, the index is used to indicate the residual adjustment method based on which the residual adjustment is performed; and, performing residual adjustment on the residual image according to the residual adjustment method indicated by the index.

5. The method according to claim 1, Features: The set residual adjustment method includes one or more of the following types of residual adjustment methods: Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller; The non-zero residual value in the residual image is added or subtracted with the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

6. A neural network-based loop filtering (NNLF) method, applied to a NNLF filter at a decoding end, wherein the NNLF filter comprises a neural network and a jump connection branch from the input to the output of the NNLF filter, wherein when performing NNLF on a reconstructed image comprising three components input to the neural network, the following processing is performed on each component: A residual adjustment flag roflag is used for decoding the component of the reconstructed image, wherein the roflag is used to indicate whether residual adjustment is required when performing NNLF on the component of the reconstructed image; When it is determined according to the roflag that residual adjustment is not required, the first mode is used to perform NNLF on the component of the reconstructed image; when it is determined according to the roflag that residual adjustment is required, the second mode is used to perform NNLF on the component of the reconstructed image; in, The first mode is an NNLF mode that does not perform residual adjustment on the component of the residual image output by the neural network, and the second mode is an NNLF mode that performs residual adjustment on the component of the residual image.

7. The method according to claim 6, Features: The residual adjustment makes the residual in the components of the residual image smaller; The reconstructed image is a reconstructed image of a current frame or a current slice or a current block; the residual adjustment usage flag is a picture-level syntax element or a block-level syntax element.

8. The method according to claim 6, Features: The adopting the first mode to perform NNLF on the component of the reconstructed image includes: adding the component of the residual image to the component of the reconstructed image to obtain the component of the filtered image output after performing NNLF on the reconstructed image; The adopting the second mode to perform NNLF on the component of the reconstructed image includes: performing residual adjustment on the component of the residual image and adding the residual to the component of the reconstructed image according to one of the residual adjustment methods set for the component, so as to obtain the component of the filtered image output after the NNLF of the reconstructed image; there are one or more residual adjustment methods set for the component.

9. The method according to claim 8, Features: There are multiple residual adjustment methods set for the component, and the residual adjustment of the component of the residual image according to one of the set residual adjustment methods includes: continuing to decode the residual adjustment method index index of the component of the reconstructed image, the index is used to indicate the residual adjustment method based on which the residual adjustment is performed; and, residual adjustment of the component of the residual image according to the residual adjustment method indicated by the index.

10. The method according to claim 6, Features: The residual adjustment methods set for the three components are the same or different; The residual adjustment method set for at least one of the three components includes one or more of the following types of residual adjustment methods: Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller; The non-zero residual value in the residual image is added or subtracted with the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

11. A video decoding method, applied to a video decoding device, include: When the neural network-based loop filter NNLF is performed on the reconstructed image, the following processing is performed: When NNLF allows residual adjustment, NNLF is performed on the reconstructed image according to the method according to any one of claims 1 to 10.

12. The method according to claim 11, Features: The NNLF allows residual adjustment when one or more of the following conditions are met: Decoding a sequence-level residual adjustment permission flag, and determining that the NNLF allows residual adjustment according to the residual adjustment permission flag; The residual adjustment permission flag of the decoded picture level is used to determine whether the NNLF allows residual adjustment according to the residual adjustment permission flag.

13. The method according to claim 11, Features: The method also includes: when NNLF does not allow residual adjustment, adding the reconstructed image input to the neural network and the residual image output by the neural network to obtain a filtered image output after NNLF is performed on the reconstructed image.

14. The method according to claim 10, Features: The NNLF filter is arranged after the deblocking filter or the sample adaptive compensation filter and before the adaptive correction filter.

15. A neural network-based loop filtering (NNLF) method, applied to a NNLF filter at an encoding end, wherein the NNLF filter comprises a neural network and a jump connection branch from the input to the output of the NNLF filter. include: Inputting the reconstructed image into the neural network to obtain a residual image output by the neural network; Calculate the rate distortion cost cost of performing NNLF on the reconstructed image using the first mode 1 , and the rate-distortion cost cost of performing NNLF on the reconstructed image using the second mode 2 ; Wherein, the first mode is an NNLF mode in which residual adjustment is not performed on the residual image, and the second mode is an NNLF mode in which residual adjustment is performed on the residual image; At cost 1 <cost 2 In the case of cost , the first mode is selected to perform NNLF on the reconstructed image; 2 <cost 1 In the case of cost, the second mode is selected to perform NNLF on the reconstructed image; 1 =cost 2 In this case, the first mode or the second mode is selected to perform NNLF on the reconstructed image.

16. The method of claim 15, Features: The reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the residual adjustment reduces the residual in the residual image.

17. The method of claim 15, Features: The calculation uses the first mode to perform NNLF on the reconstructed image. 1 , including: adding the residual image and the reconstructed image to obtain a first filtered image; and calculating the cost according to the difference between the first filtered image and the corresponding original image 1 ; The selecting the first mode to perform NNLF on the reconstructed image includes: using a first filtered image obtained by adding the residual image to the reconstructed image as a filtered image output after performing NNLF on the reconstructed image.

18. The method of claim 17, Features: The calculation uses the second mode to perform NNLF on the reconstructed image. 2 ,include: According to each set residual adjustment method, the residual image is subjected to residual adjustment and added to the reconstructed image to obtain a second filtered image, and a rate distortion cost is calculated according to the difference between the second filtered image and the original image; and The minimum rate-distortion cost among all the calculated rate-distortion costs is taken as cost 2 ; Among them, there are one or more residual adjustment methods set.

19. The method of claim 18, Features: The selecting the second mode to perform NNLF on the reconstructed image includes: Will be based on cost 2 The corresponding residual adjustment method performs residual adjustment on the residual image and adds the obtained second filtered image to the reconstructed image, which is used as the filtered image output after NNLF is performed on the reconstructed image.

20. The method of claim 18, Features: The reconstructed image and the residual image each include three components; The cost 1 The obtained value is obtained by calculating the square error and SSD of the first filtered image and the original image on three components, and weighting and adding the SSD on the three components; The calculating of a rate-distortion cost according to the difference between the second filtered image and the original image includes: calculating the square error and SSD of the second filtered image and the original image on three components, and then weighting and adding the SSD on the three components to obtain the rate-distortion cost.

21. The method of claim 18, Features: The set residual adjustment method includes one or more of the following types of residual adjustment methods: Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller; The non-zero residual value in the residual image is added or subtracted with the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

22. A neural network-based loop filtering (NNLF) method, applied to a NNLF filter at an encoding end, wherein the NNLF filter comprises a neural network and a jump connection branch from the input to the output of the NNLF filter. include: Inputting the reconstructed image into the neural network to obtain a residual image output by the neural network; the reconstructed image and the residual image both include three components; and The following processing is performed on each of the three components: Calculate the rate-distortion cost cost of performing NNLF on the component of the reconstructed image using the first mode 1 , and the rate-distortion cost cost of NNLF is performed on the component of the reconstructed image using the second mode 2 ; The first mode is an NNLF mode in which residual adjustment is not performed on the component of the residual image, and the second mode is an NNLF mode in which residual adjustment is performed on the component of the residual image; At cost 1 <cost 2 In the case of cost, the first mode is selected to perform NNLF on the component of the reconstructed image; 2 <cost 1 In the case of cost, the second mode is selected to perform NNLF on the component of the reconstructed image; 1 =cost 2 In this case, the first mode or the second mode is selected to perform NNLF on the reconstructed image.

23. The method of claim 22, Features: The reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the residual adjustment reduces the residual in the components of the residual image.

24. The method of claim 24, Features: The calculation uses the first mode to perform NNLF on the component of the reconstructed image. 1 , including: adding the component of the residual image to the component of the reconstructed image to obtain the filtered component; and, calculating the cost according to the difference between the filtered component and the component of the corresponding original image. 1 ; The selecting the first mode to perform NNLF on the component of the reconstructed image includes: adding the component of the residual image and the component of the reconstructed image to obtain the filtered component, as the component of the filtered image output after performing NNLF on the reconstructed image.

25. The method of claim 22, Features: The calculation uses the second mode to perform NNLF on the component of the reconstructed image. 2 ,include: According to each residual adjustment mode set for the component, the component of the residual image is residually adjusted and added to the component of the reconstructed image to obtain the filtered component; and a rate-distortion cost of the component is calculated according to the difference between the filtered component and the component of the corresponding original image; The minimum rate-distortion cost among all the rate-distortion costs of the component is calculated as the cost of the component. 2 ; Among them, there are one or more residual adjustment methods set for this component.

26. The method of claim 25, Features: The selecting the second mode to perform NNLF on the component of the reconstructed image comprises: According to the cost of this component 2 The corresponding residual adjustment method performs residual adjustment on the component of the residual image, and adds the filtered component obtained by adding it to the component of the reconstructed image as the component of the filtered image output after NNLF is performed on the reconstructed image.

27. The method of claim 25, Features: The residual adjustment methods set for the three components are the same or different; The residual adjustment method set for at least one of the three components includes one or more of the following types of residual adjustment methods: Adding or subtracting a fixed value from a non-zero residual value in the residual image so that the absolute value of the non-zero residual value becomes smaller; The non-zero residual value in the residual image is added or subtracted with the adjustment value corresponding to the interval according to the interval in which the non-zero residual value is located, so that the absolute value of the non-zero residual value becomes smaller; wherein there are multiple intervals, and the larger the value in the interval, the larger the corresponding adjustment value.

28. A video encoding method, applied to a video encoding device, include: When the neural network-based loop filter NNLF is performed on the reconstructed image, the following processing is performed: In case that NNLF allows residual adjustment, performing NNLF on the reconstructed image according to any one of the methods of claims 15 to 27; A residual adjustment usage flag of the reconstructed image is encoded to indicate whether residual adjustment is required when performing NNLF on the reconstructed image.

29. The method of claim 28, Features: The residual adjustment usage flag is a picture level syntax element or a block level syntax element.

30. The method of claim 28, Features: The NNLF allows residual adjustment when one or more of the following conditions are met: Decoding a sequence-level residual adjustment permission flag, and determining that the NNLF allows residual adjustment according to the residual adjustment permission flag; The residual adjustment permission flag of the decoded picture level is used to determine whether the NNLF allows residual adjustment according to the residual adjustment permission flag.

31. The method of claim 28, Features: The method also includes: when it is determined that NNLF does not allow residual adjustment, skipping the encoding of the residual usage flag, adding the reconstructed image input to the neural network and the residual image output by the neural network, to obtain a filtered image output after NNLF is performed on the reconstructed image.

32. The method of claim 28, Features: The method is to perform NNLF on the reconstructed image according to the method of claims 15 to 21; The number of flags roflag used for residual adjustment of the reconstructed image is 1; when the first mode is selected to perform NNLF on the reconstructed image, the roflag is set to a value indicating that residual adjustment is not required; when the second mode is selected to perform NNLF on the reconstructed image, the roflag is set to a value indicating that residual adjustment is required.

33. The method of claim 32, Features: The method is to perform NNLF on the reconstructed image according to the method of claims 18 to 21; The method further includes: when the roflag is set to a value indicating that residual adjustment is required and there are multiple residual adjustment methods set, continuing to encode the residual adjustment method index of the reconstructed image to indicate the residual adjustment method based on which the residual adjustment is performed.

34. The method of claim 28, Features: The method is to perform NNLF on the reconstructed image according to the method of claims 22 to 27; The number of flags roflag(j) used for residual adjustment of the reconstructed image is 3, j=1, 2, 3, roflag(j) is used to indicate whether residual adjustment is required when NNLF is performed on the j-th component of the reconstructed image; when the first mode is selected to perform NNLF on the j-th component of the reconstructed image, roflag(j) is set to a value indicating that residual adjustment is not required, and when the second mode is selected to perform NNLF on the j-th component of the reconstructed image, roflag(j) is set to a value indicating that residual adjustment is required.

35. The method of claim 34, Features: The method is to perform NNLF on the reconstructed image according to the method of claims 25 to 27; The method also includes: when the roflag(j) is set to a value indicating that residual adjustment is required and there are multiple residual adjustment methods set for the j-th component, continuing to encode the residual adjustment method index index(j) of the j-th component of the reconstructed image to indicate the residual adjustment method based on which the residual adjustment is performed on the j-th component of the residual image.

36. A code stream, in, The code stream is generated by the video encoding method as described in any one of claims 28 to 35.

37. A loop filter based on a neural network, comprising a processor and a memory storing a computer program, in, When the processor executes the computer program, it can implement the neural network-based loop filtering method as described in any one of claims 1 to 10 and 15 to 27.

38. A video decoding device, comprising a processor and a memory storing a computer program, in, When the processor executes the computer program, the video decoding method according to any one of claims 11 to 14 can be implemented.

39. A video encoding device comprising a processor and a memory storing a computer program, in, When the processor executes the computer program, it can implement the video encoding method as described in any one of claims 28 to 35.

40. A video encoding and decoding system, in, It comprises the video encoding device as claimed in claim 39 and the video decoding device as claimed in claim 38.

41. A non-transitory computer-readable storage medium storing a computer program, in, When the computer program is executed by a processor, it can implement the neural network-based loop filtering method as described in any one of claims 1 to 10, 15 to 22, or implement the video decoding method as described in any one of claims 11 to 14, or implement the video encoding method as described in any one of claims 28 to 35.