Loop filtering and video encoding and decoding method, device and system based on neural network

CN120077665APending Publication Date: 2025-05-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100905.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing digital video compression technology still has shortcomings in reducing video transmission bandwidth and traffic pressure. Especially with the growth of Internet videos and the improvement of video definition, existing technology is difficult to meet the demand for higher compression efficiency.

Method used

A loop filtering method based on neural networks is used to optimize the rate-distortion cost by allowing chroma information to be fused during the video encoding and decoding process and specifying adjustments to the chroma information in the training data, thereby selecting the best filtering mode. Loop filtering.

Benefits of technology

It improves video encoding performance, reduces encoding time and decoding time, improves image quality, and shows excellent encoding performance under different resolutions and configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077665A_ABST
    Figure CN120077665A_ABST
Patent Text Reader

Abstract

According to the loop filtering and video encoding and decoding method, device and system based on the neural network, when an encoding end carries out NNLF on a reconstructed image, an optimal mode in a chrominance information fusion mode and other modes can be selected to carry out NNLF, and a corresponding mark is set; the NNLF model used in the chromaticity information fusion mode is trained by using training data obtained by chromaticity information adjustment. And a decoding end selects one mode to carry out NNLF on the reconstructed image according to the mark, so that the coding performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Neural network-based loop filtering, video encoding and decoding method, device and system Technical Field

[0001] The embodiments of the present disclosure relate to, but are not limited to, video technology, and more specifically, to a neural network-based loop filtering method, video encoding and decoding method, device, and system. Background Art

[0002] Digital video compression technology primarily compresses large amounts of digital video data for easier transmission and storage. The original video sequence consists of luminance and chrominance components. During the digital video encoding process, the encoder reads a black-and-white or color image and divides each frame into largest coding units (LCUs) of equal size (e.g., 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) according to a set of rules, which can be further divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module can use intra-frame prediction and inter-frame prediction. Intra-frame prediction predicts the pixel information within the current block based on the information of the same image to eliminate spatial redundancy; inter-frame prediction can refer to the information of different images and use motion estimation to search for the motion vector information that best matches the current block to eliminate temporal redundancy; the transformation can convert the predicted residual into the frequency domain to redistribute its energy. Combined with quantization, it can remove information that the human eye is not sensitive to to eliminate visual redundancy; entropy coding can eliminate character redundancy based on the current context model and the probability information of the binary code stream to generate a code stream.

[0003] With the surge in Internet videos and people's increasing demand for video clarity, although existing digital video compression standards can save a lot of video data, there is still a need to pursue better digital video compression technology to reduce the bandwidth and traffic pressure of digital video transmission.

[0004] SUMMARY OF THE INVENTION

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] An embodiment of the present disclosure provides a neural network based loop filtering (NNLF) method, which is applied to a video decoding device. The method includes:

[0007] Decode a first flag of a reconstructed image, where the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image;

[0008] Determining, according to the first flag, an NNLF mode to be used when performing NNLF on the reconstructed image; performing NNLF on the reconstructed image according to the determined NNLF mode;

[0009] Among them, the NNLF mode includes a first mode and a second mode, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: expanded data obtained after specifying adjustments to the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

[0010] An embodiment of the present disclosure further provides a video decoding method, which is applied to a video decoding device, comprising: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image:

[0011] In the case where NNLF allows chrominance information fusion, the reconstructed image is subjected to NNLF according to the NNLF method described in any embodiment of the present disclosure that is applied to the decoding end and can use the chrominance information fusion mode.

[0012] An embodiment of the present disclosure further provides a neural network-based loop filtering method applied to a video encoding device, the method comprising:

[0013] Calculating a rate-distortion cost of performing NNLF on the input reconstructed image using the first mode, and a rate-distortion cost of performing NNLF on the reconstructed image using the second mode;

[0014] Determine to use a mode with the lowest rate-distortion cost between the first mode and the second mode to perform NNLF on the reconstructed image;

[0015] Among them, the first mode and the second mode are both set NNLF modes, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: the expanded data obtained after the specified adjustment of the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

[0016] An embodiment of the present disclosure further provides a video encoding method, which is applied to a video encoding device, comprising: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image:

[0017] In the case where NNLF allows chroma information fusion, performing NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure applied to the encoding end and capable of using a chroma information fusion mode;

[0018] A first flag for encoding the reconstructed image, wherein the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image.

[0019] An embodiment of the present disclosure further provides a code stream, wherein the code stream is generated by the video encoding method described in any embodiment of the present disclosure.

[0020] An embodiment of the present disclosure further provides a neural network-based loop filter, comprising a processor and a memory storing a computer program, wherein the processor, when executing the computer program, can implement the neural network-based loop filtering method described in any embodiment of the present disclosure.

[0021] An embodiment of the present disclosure further provides a video decoding device, comprising a processor and a memory storing a computer program, wherein the processor can implement the video decoding method as described in any embodiment of the present disclosure when executing the computer program.

[0022] An embodiment of the present disclosure further provides a video encoding device, including a processor and a memory storing a computer program, wherein the processor can implement the video encoding method as described in any embodiment of the present disclosure when executing the computer program.

[0023] An embodiment of the present disclosure further provides a video encoding and decoding system, comprising the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0024] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, can implement the neural network-based loop filtering method as described in any embodiment of the present disclosure, or implement the video decoding method as described in any embodiment of the present disclosure, or implement the video encoding method as described in any embodiment of the present disclosure.

[0025] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0026] Summary of the Figures

[0027] The accompanying drawings are used to provide an understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.

[0028] FIG1A is a schematic diagram of an encoding and decoding system according to an embodiment, FIG1B is a schematic diagram of an encoding end in FIG1A , and FIG1C is a schematic diagram of a decoding end in FIG1A ;

[0029] FIG2 is a block diagram of a filter unit according to an embodiment;

[0030] FIG3A is a network structure diagram of an NNLF filter according to an embodiment; FIG3B is a structure diagram of a residual block in FIG3A ; FIG3C is a schematic diagram of the input and output of the NNLF filter in FIG3A ;

[0031] FIG4A is a structural diagram of a backbone network in a NNLF filter according to another embodiment; FIG4B is a structural diagram of a residual block in FIG4A ; FIG4C is a schematic diagram of input and output of the NNLF filter in FIG4A ;

[0032] FIG5 is a schematic diagram of an arrangement of feature maps of input information;

[0033] FIG6 is a flowchart of an NNLF method applied to an encoding end according to an embodiment of the present disclosure;

[0034] FIG7 is a module diagram of a filter unit according to an embodiment of the present disclosure;

[0035] FIG8 is a flowchart of a video encoding method according to an embodiment of the present disclosure;

[0036] FIG9 is a flowchart of an NNLF method applied to a decoding end according to an embodiment of the present disclosure;

[0037] FIG10 is a flowchart of a video decoding method according to an embodiment of the present disclosure;

[0038] FIG11 is a schematic diagram of the hardware structure of the NNLF filter according to an embodiment of the present disclosure;

[0039] FIG12A is a schematic diagram of the input and output of the NNLF filter when the order of chrominance information is not adjusted according to an embodiment of the present disclosure;

[0040] FIG12B is a schematic diagram of the input and output of the NNLF filter when the order of chrominance information is not adjusted according to an embodiment of the present disclosure;

[0041] FIG13 is a flowchart of an NNLF method applied to a decoding end according to an embodiment of the present disclosure;

[0042] FIG14 is a flowchart of a video decoding method according to an embodiment of the present disclosure;

[0043] FIG15 is a flowchart of an NNLF method applied to an encoding end according to an embodiment of the present disclosure;

[0044] FIG16 is a flowchart of a video encoding method according to an embodiment of the present disclosure;

[0045] FIG17 is a flowchart of a training method for a NNLF model according to an embodiment of the present disclosure.

[0046] Details

[0047] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure.

[0048] In the description of the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment described as "exemplary" or "for example" in the present disclosure should not be interpreted as being more preferred or advantageous than other embodiments. "And / or" in this article is a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "Multiple" refers to two or more than two. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present disclosure, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not limit them to be necessarily different.

[0049] When describing representative exemplary embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific sequence of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art can readily understand that these sequences can vary and still remain within the spirit and scope of the disclosed embodiments.

[0050] The neural network-based loop filtering method and video coding and decoding method of the disclosed embodiments can be applied to various video coding and decoding standards, such as: H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), AVS (Audio Video coding Standard), and other standards developed by MPEG (Moving Picture Experts Group), AOM (Alliance for Open Media), JVET (Joint Video Experts Team) and extensions of these standards, or any other customized standards.

[0051] Figure 1A is a block diagram of a video encoding and decoding system that can be used in embodiments of the present disclosure. As shown in the figure, the system is divided into an encoding end 1 and a decoding end 2. The encoding end 1 generates a bitstream. The decoding end 2 can decode the bitstream. The decoding end 2 can receive the bitstream from the encoding end 1 via a link 3. The link 3 includes one or more media or devices capable of moving the bitstream from the encoding end 1 to the decoding end 2. In one example, the link 3 includes one or more communication media that enable the encoding end 1 to send the bitstream directly to the decoding end 2. The encoding end 1 modulates the bitstream according to a communication standard and sends the modulated bitstream to the decoding end 2. The one or more communication media may include wireless and / or wired communication media and may form part of a packet network. In another example, the bitstream can also be output from the output interface 15 to a storage device, and the decoding end 2 can read the stored data from the storage device via streaming or downloading.

[0052] As shown in the figure, encoding end 1 includes a data source 11, a video encoding device 13, and an output interface 15. Data source 11 may include a video capture device (e.g., a camera), an archive containing previously captured data, a feed interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. Video encoding device 13, also known as a video encoder, encodes data from data source 11 and outputs it to output interface 15. Output interface 15 may include at least one of a modulator, a modem, and a transmitter. Decoding end 2 includes an input interface 21, a video decoding device 23, and a display device 25. Input interface 21 includes at least one of a receiver and a modem. Input interface 21 can receive a bitstream via link 3 or from a storage device. Video decoding device 23, also known as a video decoder, decodes the received bitstream. Display device 25 displays the decoded data. Display device 25 may be integrated with other devices in decoding end 2 or provided separately; display device 25 is optional for the decoding end. In other examples, the decoding end may include other devices or equipment for applying the decoded data.

[0053] FIG1B is a block diagram of an exemplary video encoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video encoding device 10 includes:

[0054] The division unit 101 is configured to cooperate with the prediction unit 100 to divide the received video data into slices, coding tree units (CTUs) or other larger units. The received video data may be a video sequence including video frames such as I frames, P frames or B frames.

[0055] The prediction unit 100 is configured to divide a CTU into coding units (CUs) and perform intra-frame prediction coding or inter-frame prediction coding on the CU. When performing intra-frame prediction and inter-frame prediction on the CU, the CU can be divided into one or more prediction units (PUs).

[0056] The prediction unit 100 includes an inter-frame prediction unit 121 and an intra-frame prediction unit 126 .

[0057] The inter-frame prediction unit 121 is configured to perform inter-frame prediction on the PU and generate prediction data for the PU, wherein the prediction data includes the prediction block of the PU, the motion information of the PU, and various syntax elements. The inter-frame prediction unit 121 may include a motion estimation (ME) unit and a motion compensation (MC) unit. The motion estimation unit can be used to perform motion estimation to generate a motion vector, and the motion compensation unit can be used to obtain or generate a prediction block based on the motion vector.

[0058] The intra prediction unit 126 is configured to perform intra prediction on a PU and generate prediction data for the PU. The prediction data for the PU may include a prediction block and various syntax elements of the PU.

[0059] The residual generating unit 102 (indicated by the circle with a plus sign after the dividing unit 101 in the figure) is configured to subtract the prediction block of the PU into which the CU is divided from the original block of the CU to generate a residual block of the CU.

[0060] The transform processing unit 104 is configured to partition a CU into one or more transform units (TUs). The partitioning of prediction units and transform units may be different. A TU-associated residual block is a sub-block obtained by partitioning the residual block of the CU. A TU-associated coefficient block is generated by applying one or more transforms to the TU-associated residual block.

[0061] The quantization unit 106 is configured to quantize the coefficients in the coefficient block based on a quantization parameter. The quantization degree of the coefficient block can be changed by adjusting the quantization parameter (QP: Quantizer Parameter).

[0062] The inverse quantization unit 108 and the inverse transform unit 110 are configured to apply inverse quantization and inverse transform to the coefficient block, respectively, to obtain a reconstructed residual block associated with the TU.

[0063] The reconstruction unit 112 (represented by the circle with a plus sign after the inverse transform processing unit 110 in the figure) is configured to add the reconstructed residual block and the prediction block generated by the prediction unit 100 to generate a reconstructed image.

[0064] The filter unit 113 is configured to perform loop filtering on the reconstructed image.

[0065] The decoded image buffer 114 is configured to store the reconstructed image after loop filtering. The intra prediction unit 126 can extract reference images of blocks adjacent to the current block from the decoded image buffer 114 to perform intra prediction. The inter prediction unit 121 can use the reference image of the previous frame cached in the decoded image buffer 114 to perform inter prediction on the PU of the current frame image.

[0066] The entropy coding unit 115 is configured to perform entropy coding operations on received data (such as syntax elements, quantized coefficient blocks, motion information, etc.) to generate a video bitstream.

[0067] In other examples, the video encoding apparatus 10 may include more, fewer, or different functional components than those in this example, for example, the transform processing unit 104 and the inverse transform processing unit 110 may be eliminated.

[0068] FIG1C is a block diagram of an exemplary video decoding device that can be used in an embodiment of the present disclosure. As shown in the figure, the video decoding device 15 includes:

[0069] Entropy decoding unit 150 is configured to perform entropy decoding on the received encoded video stream, extracting syntax elements, quantized coefficient blocks, and motion information for PUs. Prediction unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, and filter unit 159 may each perform corresponding operations based on the syntax elements extracted from the stream.

[0070] The inverse quantization unit 154 is configured to perform inverse quantization on the coefficient block associated with the quantized TU.

[0071] The inverse transform processing unit 156 is configured to apply one or more inverse transforms to the inverse quantized coefficient block to generate a reconstructed residual block for the TU.

[0072] Prediction unit 152 includes an inter-prediction unit 162 and an intra-prediction unit 164. If the current block is encoded using intra prediction, intra prediction unit 164 determines an intra prediction mode for the PU based on syntax elements decoded from the codestream, and performs intra prediction in conjunction with reconstructed reference information adjacent to the current block obtained from decoded image buffer 160. If the current block is encoded using inter prediction, inter prediction unit 162 determines a reference block for the current block based on motion information of the current block and corresponding syntax elements, and performs inter prediction on the reference block obtained from decoded image buffer 160.

[0073] The reconstruction unit 158 ​​(represented by a circle with a plus sign after the inverse transform processing unit 155 in the figure) is set to perform intra-frame prediction or inter-frame prediction on the current block based on the reconstructed residual block associated with the TU and the prediction unit 152 to obtain a reconstructed image.

[0074] The filter unit 159 is configured to perform loop filtering on the reconstructed image.

[0075] The decoded image buffer 160 is configured to store the reconstructed image after loop filtering as a reference image for subsequent motion compensation, intra-frame prediction, inter-frame prediction, etc. The filtered reconstructed image can also be output as decoded video data for presentation on a display device.

[0076] In other embodiments, the video decoding device 15 may include more, fewer, or different functional components. For example, the inverse transform processing unit 155 may be eliminated in some cases.

[0077] Based on the above-described video encoding and decoding devices, the following basic encoding and decoding process can be performed. On the encoding side, a frame of an image is divided into blocks. Intra-frame prediction, inter-frame prediction, or other algorithms are performed on the current block to generate a predicted block for the current block. The predicted block is subtracted from the original block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain quantized coefficients, and the quantized coefficients are entropy encoded to generate a bitstream. On the decoding side, intra-frame prediction or inter-frame prediction is performed on the current block to generate a predicted block for the current block. The quantized coefficients obtained from the decoded bitstream are then inversely quantized and inversely transformed to obtain a residual block. The predicted block and residual block are added together to obtain a reconstructed block. The reconstructed block is then loop-filtered on the reconstructed image based on an image or block basis to obtain a decoded image. The encoding side also performs similar operations as the decoding side to obtain a decoded image, also known as a reconstructed image after loop filtering. The reconstructed image after loop filtering can be used as a reference frame for inter-frame prediction of subsequent frames. Block division information, prediction, transform, quantization, entropy coding, loop filtering, and other mode and parameter information determined by the encoding side can be written into the bitstream. The decoding end determines the block division information, prediction, transformation, quantization, entropy coding, loop filtering and other mode information and parameter information used by the encoding end by decoding the code stream or analyzing according to the setting information, so as to ensure that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end.

[0078] Although the above is an example of a block-based hybrid coding framework, the embodiments of the present disclosure are not limited thereto. With the development of technology, one or more modules in the framework and one or more steps in the process can be replaced or optimized.

[0079] The embodiments of the present disclosure relate to, but are not limited to, the filter unit (the filter unit may also be referred to as a loop filtering module) in the above-mentioned encoding end and decoding end and the corresponding loop filtering method.

[0080] In one embodiment, the filter units at the encoding and decoding ends include tools such as a deblocking filter (DBF) 20, a sample adaptive offset filter (SAO) 22, and an adaptive loop filter (ALF) 26. Between the SAO and ALF, a neural network-based loop filter (NNLF) 26 is also included, as shown in Figure 2. The filter units perform loop filtering on the reconstructed image to compensate for distortion and provide a better reference for subsequent pixel encoding.

[0081] In an exemplary embodiment, a neural network-based loop filtering NNLF solution is provided, and the model used (also referred to as a network model) uses the filtering network shown in Figure 3A. The NNLF is denoted as NNLF1 in the text, and the filter that executes NNLF1 is referred to as the NNLF1 filter. As shown in the figure, the backbone network (backbone) of the filtering network includes the use of multiple residual blocks (ResBlock) connected in sequence, and also includes a convolution layer (represented by Conv in the figure), an activation function layer (ReLU in the figure), a merging (concat) layer (represented by Cat in the figure), and a pixel reorganization layer (represented by PixelShuffle in the figure). The structure of each residual block is shown in Figure 3B, including a convolution layer with a convolution kernel size of 1×1, a ReLU layer, a convolution layer with a convolution kernel size of 1×1, and a convolution layer with a convolution kernel size of 3×3 connected in sequence.

[0082] As shown in Figure 3A, the input of the NNLF1 filter includes the luminance information (i.e., Y component) and chrominance information (i.e., U component and V component) of the reconstructed image (rec_YUV), as well as a variety of auxiliary information, such as the luminance information and chrominance information of the predicted image (pred_YUV), QP information, and frame type information. The QP information includes the default baseline quantization parameter (BaseQP: Base Quantization Parameter) in the encoding profile and the slice quantization parameter (SliceQP: Slice Quantization Parameter) of the current slice. The frame type information includes the slice type (SliceType), that is, the type of frame to which the current slice belongs. The output of the model is the filtered image (output_YUV) after NNLF1 filtering. The filtered image output by the NNLF1 filter can also be used as the reconstructed image input to subsequent filters.

[0083] NNLF1 uses a model to filter the YUV components of the reconstructed image (rec_YUV) and output the YUV components of the filtered image (out_YUV), as shown in Figure 3C. Auxiliary input information such as the YUV components of the predicted image are omitted in this figure. The filtering network of this model has a skip connection branch between the input reconstructed image and the output filtered image, as shown in Figure 3A.

[0084] Another exemplary embodiment provides another NNLF scheme, denoted as NNLF2. NNLF2 requires training two models separately, one model for filtering the luminance component of the reconstructed image, and the other model for filtering the two chrominance components of the reconstructed image. The two models can use the same filtering network, and there is also a jump connection branch between the reconstructed image input to the NNLF2 filter and the filtered image output by the NNLF2 filter. As shown in Figure 4A, the backbone network of the filtering network includes a plurality of residual blocks (AttRes Block) with attention mechanisms connected in sequence, a convolutional layer (Conv 3×3) for implementing feature mapping, and a reorganization layer (Shuffle). The structure of each residual block with an attention mechanism is shown in Figure 4B, including a convolutional layer (Conv 3×3), an activation layer (PReLU), a convolutional layer (Conv 3×3) and an attention layer (Attintion) connected in sequence. M represents the number of feature maps, and N represents the number of samples in one dimension.

[0085] Model 1 of NNLF2 for filtering the luminance component of the reconstructed image is shown in Figure 4C. Its input information includes the luminance component of the reconstructed image (rec_Y), and the output is the luminance component of the filtered image (out_Y). Auxiliary input information such as the luminance component of the predicted image is omitted in the figure. Model 2 of NNLF2 for filtering the two chrominance components of the reconstructed image is shown in Figure 4C. Its input information includes the two chrominance components of the reconstructed image (rec_UV) and the luminance component of the reconstructed image as auxiliary input information (rec_Y). The output of model 2 is the two chrominance components of the filtered image (out_UV). Model 1 and model 2 can also include other auxiliary input information, such as QP information, block partition image, deblocking filter boundary strength information, etc.

[0086] The above-mentioned NNLF1 and NNLF2 schemes can be implemented in neural network based video coding (NNVC: Neural Network based Video Coding) using neural network based common software (NCS: Neural Network based Common Software), serving as the baseline tool in the NNVC reference software testing platform, namely baseline NNLF.

[0087] In video coding, inter-frame prediction technology allows the current frame to reference the image information of the previous frame, improving coding performance. However, the coding performance of the previous frame also affects the coding performance of subsequent frames. In the NNLF1 and NNLF2 schemes, to adapt the filter network to the influence of inter-frame prediction technology, the model training process includes an initial training phase and an iterative training phase, using multiple rounds of training. In the initial training phase, the model to be trained is not yet deployed in the encoder. The model is trained in the first round using sample data of reconstructed images, resulting in a trained model. In the iterative training phase, the model is deployed in the encoder. The trained model is first deployed in the encoder, and sample data of reconstructed images is re-collected. The trained model is then trained in the second round, resulting in a trained model. The trained model is then deployed in the encoder again, and sample data of reconstructed images is re-collected. The trained model is then trained in the third round, resulting in a trained model. The training process repeats itself. Finally, each trained model is tested on a validation set to identify the model with the best coding performance for deployment.

[0088] As mentioned above, in NNLF1` and NNLF2, for the luminance and chrominance components of the reconstructed image input to the neural network, NNLF1 adopts a joint input method, as shown in Figure 3C, and only one network model needs to be trained; in NNLF2, the luminance and chrominance components of the reconstructed image are input separately, as shown in Figure 4C, and two models need to be trained. For the two chrominance components, namely the U component and the V component, NNLF1` and NNLF2 both adopt a joint input method, that is, there is a binding relationship. As shown in Figure 5, when the stacked feature image is sent to the model for prediction and output, the three components in each frame of the image are arranged in the order of Y, U, and V, with the U component in front and the V component in the back. Currently, there is a lack of solutions to study the impact of adjusting the input order of the U component and the V component on the neural network.

[0089] The disclosed embodiment proposes a method for adjusting chrominance information, by adjusting the chrominance information input to the NNLF filter, such as swapping the order of the U component and the V component, to further optimize the encoding performance of the NNLF filter.

[0090] An embodiment of the present disclosure provides a neural network-based loop filtering (NNLF) method, which is applied to a video encoding device. As shown in FIG6 , the method includes:

[0091] S110, calculating a rate-distortion cost of performing NNLF on an input reconstructed image using a first mode, and a rate-distortion cost of performing NNLF on the reconstructed image using a second mode;

[0092] S120, determining to use a mode with the lowest rate-distortion cost between the first mode and the second mode to perform NNLF on the reconstructed image;

[0093] Among them, the first mode and the second mode are both set NNLF modes, and the second mode includes a chrominance information adjustment mode. Compared with the first mode, the chrominance information adjustment mode adds a process of performing specified adjustments on the input chrominance information before filtering.

[0094] Tests have shown that adjusting the input chrominance information can affect encoding performance. This embodiment can select an optimal mode from the second mode that adjusts the chrominance information and the first mode that does not adjust the chrominance information according to the rate-distortion cost, thereby improving encoding performance.

[0095] In an exemplary embodiment of the present disclosure, the specified adjustment of the chromaticity information includes any one or more of the following adjustment methods:

[0096] Swap the order of the two chrominance components of the reconstructed image. For example, the order of U component first and V component last is adjusted to the order of V component first and U component last;

[0097] The weighted average value and the square error value of the two chrominance components of the reconstructed image are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

[0098] In an exemplary embodiment of the present disclosure, the reconstructed image is a reconstructed image of the current frame, current slice, or current block, but may also be a reconstructed image of other coding units. The reconstructed image subjected to the NNLF filter herein may belong to coding units at different levels, such as the image level (including frame and slice) and the block level.

[0099] In an exemplary embodiment of the present disclosure, the network structures of the models used in the first mode and the second mode are the same or different.

[0100] In an exemplary embodiment of the present disclosure, the second mode further includes a chrominance information fusion mode, and the training data used in training the model used in the chrominance information fusion mode includes expanded data obtained by performing the specified adjustment on the chrominance information of the reconstructed image in the original data, or includes the original data and the expanded data;

[0101] The network model used in the first mode is trained using the original data.

[0102] In an exemplary embodiment of the present disclosure, the second mode further includes a chrominance information adjustment and fusion mode;

[0103] Compared to the first mode, the chrominance information adjustment and fusion mode adds a process of performing specified adjustments on the input chrominance information before filtering, and the training data used in training the model used in the chrominance information fusion mode includes expanded data obtained by performing the specified adjustments on the chrominance information of the reconstructed image in the original data, or includes the original data and the expanded data;

[0104] The network model used in the first mode is trained using the original data.

[0105] In an exemplary embodiment of the present disclosure, there is one first mode, and there are one or more second modes;

[0106] The calculating the rate-distortion cost cost1 of performing NNLF on the input reconstructed image using the first mode includes: obtaining a first filtered image output after performing NNLF on the reconstructed image using the first mode, and calculating the cost1 according to the difference between the first filtered image and the corresponding original image;

[0107] The calculation of the rate-distortion cost cost2 of performing NNLF on the reconstructed image using the second mode includes: for each of the second modes, obtaining a second filtered image output after performing NNLF on the reconstructed image using the second mode, and calculating the cost2 of the second mode based on the difference between the second filtered image and the original image.

[0108] In an example of this embodiment, there are multiple second modes; the mode with the lowest rate-distortion cost among the first and second modes is a mode corresponding to the minimum value of the calculated cost1 and multiple costs2.

[0109] In one example of this embodiment, the difference is represented by the sum of squared errors (SSD). In other examples, the difference can also be represented by other indicators such as mean squared error (MSE) and mean absolute error (MAE), and this disclosure is not limited to this. The same applies to other embodiments of the present disclosure.

[0110] In an exemplary embodiment of the present disclosure, the method further includes: using the filtered image obtained when the mode with the smallest rate-distortion cost is used to perform NNLF on the reconstructed image as the filtered image output after performing NNLF on the reconstructed image.

[0111] In an exemplary embodiment of the present disclosure, the NNLF filter for performing NNLF processing at the encoding end and / or the decoding end is arranged after the deblocking filter or the sample adaptive offset (SAO) filter and before the adaptive correction filter. In an example of this embodiment, the structure of the filter unit (or the loop filtering module, see FIGS. 1B and 1C) is shown in FIG. 7. In the figure, DBF represents the deblocking filter, SAO represents the sample adaptive offset filter, and ALF represents the adaptive correction filter. NNLF-A represents the NNLF filter using the first mode, and NNLF-B represents the NNLF filter using the second mode such as the chrominance information adjustment mode, and can also be referred to as the chrominance information adjustment module (Chroma Adjustment, CA). The NNLF-A filter may be the same as the NNLF filters such as NNLF1 and NNLF2 that do not perform chrominance information adjustment mentioned above.

[0112] Although it is shown as the NNLF-A filter and the NNLF-B filter in FIG. 7, in actual implementation, a single model can be used. The NNLF-B filter can be regarded as the NNLF-A filter with chrominance information adjustment added. Based on this model, without adjusting the input chrominance information (such as swapping the UV order), perform loop filtering once using this model (that is, perform NNLF on the reconstructed image using the first mode, or it can be said to perform NNLF on the reconstructed image using the NNLF-A filter), and calculate a rate-distortion cost cost1 according to the difference between the output first filtered image and the original image, such as the sum of squared differences (SSD); when adjusting the input chrominance information, perform loop filtering once using this model (that is, perform NNLF on the reconstructed image using the second mode, or it can be said to perform NNLF on the reconstructed image using the NNLF-B filter), and calculate a rate-distortion cost cost2 according to the difference between the output second filtered image and the original image, then the used NNLF mode can be determined according to the magnitudes of cost1 and cost2. The information of the used NNFL mode (that is, the selected NNFL mode) is represented by a first flag and encoded into the bitstream for the decoder to read. At the decoding end, determine the actual NNLF mode used at the encoding end by decoding the first flag, and perform NNLF on the input reconstructed image using the determined NNLF mode.

[0113] For example, when cost1 < cost2, perform NNLF on the reconstructed image using the first mode; when cost2 < cost1, perform NNLF on the reconstructed image using the second mode; when cost1 = cost2, perform NNLF on the reconstructed image using the first mode or the second mode.

[0114] During loop filtering, some or all of DBF, SAO, and ALF may be disabled. Furthermore, the deployment location of the NNLF filter is not limited to the location shown in the figure. It is readily understood that the implementation of the disclosed NNLF method is not limited to its deployment location. Furthermore, the filters in the filter unit are not limited to those shown in FIG. 7 ; there may be more or fewer filters, or other types of filters may be used.

[0115] An embodiment of the present disclosure further provides a video encoding method, which is applied to a video encoding device, including: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image, as shown in FIG8 :

[0116] Step S210: When NNLF allows chroma information adjustment, perform NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure applied to the encoder and using the chroma information adjustment mode;

[0117] Step S220 : Encode a first flag of the reconstructed image, where the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image.

[0118] Among them, the first mode and the second mode are both set NNLF modes, and the second mode includes a chrominance information adjustment mode. Compared with the first mode, the chrominance information adjustment mode adds a process of performing specified adjustments on the input chrominance information before filtering.

[0119] This embodiment can select an optimal mode from the second mode for adjusting chrominance information and the first mode for not adjusting chrominance information according to the rate-distortion cost, and encode the selected mode information into the bitstream, thereby improving encoding performance.

[0120] In an exemplary embodiment of the present disclosure, the first flag is a picture-level syntax element or a block-level syntax element.

[0121] In an exemplary embodiment of the present disclosure, it is determined that the NNLF allows chrominance information adjustment when one or more of the following conditions are met:

[0122] Decoding a chroma information adjustment permission flag at the sequence level, and determining that the NNLF allows chroma information adjustment according to the chroma information adjustment permission flag;

[0123] The chroma information adjustment permission flag of the decoded picture level is determined based on the chroma information adjustment permission flag to determine whether the NNLF allows the chroma information adjustment.

[0124] In an exemplary embodiment of the present disclosure, the method further includes: when it is determined that NNLF does not allow chroma information adjustment, performing NNLF on the reconstructed image using the first mode, and skipping encoding of the chroma information adjustment permission flag.

[0125] In an exemplary embodiment of the present disclosure, the second mode is a chrominance information adjustment mode; the first flag is used to indicate whether to perform chrominance information adjustment when performing NNLF on the reconstructed image;

[0126] The first flag for encoding the reconstructed image includes: when it is determined that the first mode is used to perform NNLF on the reconstructed image, the first flag is set to a value indicating that chromaticity information adjustment is not performed; when it is determined that the second mode is used to perform NNLF on the reconstructed image, the first flag is set to a value indicating that chromaticity information adjustment is performed.

[0127] In an exemplary embodiment of the present disclosure, there are multiple second modes; the first flag is used to indicate whether the second mode is used when performing NNLF on the reconstructed image;

[0128] The first flag of encoding the reconstructed image includes:

[0129] If it is determined that the first mode is used to perform NNLF on the reconstructed image, setting the first flag to a value indicating that the second mode is not used;

[0130] When it is determined to use the second mode to perform NNLF on the reconstructed image, the first flag is set to a value indicating the use of the second mode, and the second flag is continuously encoded, where the second flag contains index information of a second mode with the minimum rate-distortion cost.

[0131] An embodiment of the present disclosure provides a video encoding method applied to an encoding end, mainly involving NNLF processing. The second mode of this embodiment is a chroma information adjustment mode.

[0132] When loop filtering is performed on the reconstructed image, if the filter unit has multiple filters, the filters are processed in the specified order. When the data input to the NNFL filter is obtained, such as the reconstructed image (which can be the filtered image output by other filters), the following processing is performed:

[0133] Step a) determines whether chroma information adjustment is allowed in the NNLF of the current sequence based on the sequence-level chroma information adjustment enable flag ca_enable_flag; if ca_enable_flag is "1", chroma information adjustment is attempted for the current sequence, and the process jumps to step b); if ca_enable_flag is "0", chroma information adjustment is not allowed in the NNLF of the current sequence, and NNLF is performed on the reconstructed image using the first mode, and encoding of the first flag is skipped, and the process ends;

[0134] Step b) For the reconstructed image of the current frame of the current sequence, first perform NNLF using the first mode, that is, input the original input information into the NNLF model for filtering, and obtain a first filtered image from the output of the model;

[0135] Step c) NNLF is performed using the second mode (the chroma information adjustment mode in this embodiment), that is, the order of the U component and the V component of the input reconstructed image are swapped, and then the image is input into the NNLF model for filtering, and a second filtered image is obtained from the output of the model;

[0136] Step d): Calculate the rate-distortion cost C based on the difference between the first filtered image and the original image. NNLF , calculate the rate-distortion cost C based on the difference between the second filtered image and the original image CA ; Compare the two rate-distortion costs, if C CA <C NNLF , it is determined to use the second mode to perform NNLF on the reconstructed image, and the second filtered image is used as the filtered image finally output after the reconstructed image is subjected to NNFL filtering; if C CA ≥C NNLF , it is determined to use the first mode to perform NNLF on the reconstructed image, and the first filtered image is used as the filtered image finally output after the reconstructed image is subjected to NNFL filtering;

[0137] The calculation formula of the rate-distortion cost in this embodiment is:

[0138] cost=Wy*SSD(Y)+Wu*SSD(U)+Wv*SSD(V)

[0139] Where SSD(*) indicates the SSD for a certain color component; Wy, Wu, and Wv represent the weighted values ​​of the SSD of the Y component, U component, and V component, respectively, and can be 10:1:1 or 8:1:1.

[0140] The calculation formula for SSD is as follows:

[0141]

[0142] M represents the length of the reconstructed image of the current frame, N represents the width of the reconstructed image of the current frame, rec(x, y) and org(x, y) represent the pixel values ​​of the reconstructed image and the original image at the pixel point (x, y), respectively.

[0143] Step e) encoding a first flag to indicate whether chrominance information adjustment is to be performed according to the mode used by the current frame (i.e., the selected mode). In this case, the first flag may also be referred to as a chrominance information adjustment use flag, picture_ca_enable_flag, which is used to indicate that chrominance information adjustment is required when performing NNLF on the reconstructed image.

[0144] If the current frame has been processed, the next frame of reconstructed image is loaded and processed in the same way.

[0145] In this embodiment, NNLF processing is performed based on the reconstructed image of the current frame. In other embodiments, NNLF processing may also be performed based on other coding units such as blocks (such as CTU) and slices in the current frame.

[0146] As shown in Figure 12A, for a certain NNLF filter, when using the first mode, the arrangement order of its input information can be {recY, recU, recV, predY, predU, predV, baseQP, sliceQP, slicetype,…}, and the arrangement order of its output information can be {cnnY, cnnU, cnnV}, where rec represents the reconstructed image, pred represents the predicted image, and cnn represents the output filtered image.

[0147] When using the second mode for chroma information adjustment, the order of filter input information is adjusted to {recY, recV, recU, predY, predV, predU, baseQP, sliceQP, slicetype, ...}, and the order of its output network information is adjusted to {cnnY, cnnV, cnnU}, as shown in Figure 12B.

[0148] This embodiment explores the generalizability of NNLF input by adjusting the input order of the U and V components in the chrominance information, further improving the filtering performance of the neural network. Furthermore, only a few bits of flags need to be encoded as control switches, which has little impact on decoding complexity.

[0149] In this embodiment, the decision is made for the three YUV components in the image through the joint rate-distortion cost. In other embodiments, more refined processing can also be attempted. For each component, the rate-distortion cost in different modes is calculated separately, and the mode with the smallest rate-distortion cost is selected.

[0150] An embodiment of the present disclosure further provides a neural network-based loop filtering (NNLF) method, which is applied to a video decoding device. As shown in FIG9 , the method includes:

[0151] Step S310: decoding a first flag of a reconstructed image, where the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image;

[0152] Step S320, determining an NNLF mode to be used when performing NNLF on the reconstructed image according to the first flag, and performing NNLF on the reconstructed image according to the determined NNLF mode;

[0153] The NNLF mode includes a first mode and a second mode. The second mode includes a chrominance information adjustment mode. Compared with the first mode, the chrominance information adjustment mode adds a process of performing specified adjustments on the input chrominance information before filtering.

[0154] This embodiment is based on a loop filtering method of a neural network. By using a first flag, a better mode is selected from two modes of adjusting chrominance information and not adjusting chrominance information, thereby enhancing the filtering effect of NNLF and improving the quality of the decoded image.

[0155] In an exemplary embodiment of the present disclosure, the specified adjustment of the chromaticity information includes any one or more of the following adjustment methods:

[0156] Interchange the order of the two chrominance components of the reconstructed image;

[0157] The weighted average value and the square error value of the two chrominance components of the reconstructed image are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

[0158] In an exemplary embodiment of the present disclosure, the reconstructed image is a reconstructed image of a current frame, a current slice, or a current block.

[0159] In an exemplary embodiment of the present disclosure, the first flag is a picture-level syntax element or a block-level syntax element.

[0160] In an exemplary embodiment of the present disclosure, the second mode further includes a chrominance information fusion mode, and the training data used in training the model used in the chrominance information fusion mode includes expanded data obtained by performing the specified adjustment on the chrominance information of the reconstructed image in the original data, or includes the original data and the expanded data;

[0161] The model used in the first mode is trained using the original data.

[0162] In an exemplary embodiment of the present disclosure, the second mode further includes a chrominance information adjustment and fusion mode;

[0163] Compared to the first mode, the chrominance information adjustment and fusion mode adds a process of performing specified adjustments on the input chrominance information before filtering, and the training data used in training the model used in the chrominance information fusion mode includes expanded data obtained by performing the specified adjustments on the chrominance information of the reconstructed image in the original data, or includes the original data and the expanded data;

[0164] The model used in the first mode is trained using the original data.

[0165] In an exemplary embodiment of the present disclosure, there is one first mode and one or more second modes; the network structures of the models used in the first mode and the second mode are the same or different.

[0166] In an exemplary embodiment of the present disclosure, the second mode is a chrominance information adjustment mode; the first flag is used to indicate whether to perform chrominance information adjustment when performing NNLF on the reconstructed image;

[0167] Determining the NNLF mode to be used when performing NNLF on the reconstructed image according to the first flag includes: when the first flag indicates that chromaticity information adjustment is to be performed, determining that the second mode is to be used when performing NNLF on the reconstructed image; when the first flag indicates that chromaticity information adjustment is not to be performed, determining that the first mode is to be used when performing NNLF on the reconstructed image.

[0168] In an example of this embodiment, the image header is defined as follows:

[0169]

[0170] The ca_enable_flag in the table is a flag that allows sequence-level chroma information adjustment. When ca_enable_flag is 1, the following semantics are defined: picture_ca_enable_flag indicates the use of picture-level chroma information adjustment, which is the first flag mentioned above. When picture_ca_enable_flag is 1, it indicates that chroma information adjustment is performed when NNLF is performed on the reconstructed image, that is, the second mode (the chroma information adjustment mode in this embodiment) is used; when picture_ca_enable_flag is 0, it indicates that chroma information adjustment is not performed when NNLF is performed on the reconstructed image, that is, the first mode is used. When ca_enable_flag is 0, decoding and encoding of picture_ca_enable_flag are skipped.

[0171] In an exemplary embodiment of the present disclosure, there are multiple second modes; the first flag is used to indicate whether the second mode is used when performing NNLF on the reconstructed image;

[0172] The determining, according to the first flag, a NNLF mode to be used when performing NNLF on the reconstructed image includes:

[0173] When the first flag indicates that the second mode is not used, determining to use the first mode when performing NNLF on the reconstructed image;

[0174] When the first flag indicates to use the second mode, continue decoding the second flag, where the second flag includes index information of a second mode to be used; and determine to use the second mode when performing NNLF on the reconstructed image according to the second flag.

[0175] An embodiment of the present disclosure further provides a video decoding method, which is applied to a video decoding device, including: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image, as shown in FIG10 :

[0176] Step S410, determining whether NNLF allows chrominance information adjustment;

[0177] Step S420 : When NNLF allows adjustment of chrominance information, perform NNLF on the reconstructed image according to the NNLF method described in any embodiment of the present disclosure that is applied to a decoding end and can use a chrominance information adjustment mode.

[0178] The video decoding method of this embodiment selects a better mode from two modes of performing chrominance information adjustment and not performing chrominance information adjustment through the first flag, which can enhance the filtering effect of NNLF and improve the quality of the decoded image.

[0179] In an exemplary embodiment of the present disclosure, it is determined that the NNLF allows chrominance information adjustment when one or more of the following conditions are met:

[0180] Decoding a chroma information adjustment permission flag at the sequence level, and determining that the NNLF allows chroma information adjustment according to the chroma information adjustment permission flag;

[0181] The chroma information adjustment permission flag of the decoded picture level is determined based on the chroma information adjustment permission flag to determine whether the NNLF allows the chroma information adjustment.

[0182] In an example of using the sequence-level chroma information adjustment permission flag, the sequence header of the video sequence is shown in the following table:

[0183]

[0184] The ca_enable_flag in the table is the sequence-level chrominance information adjustment enable flag.

[0185] In an exemplary embodiment of the present disclosure, the method further includes: when NNLF does not allow adjustment of chrominance information, skipping decoding of the first flag, and performing NNLF on the reconstructed image using the first mode.

[0186] This embodiment selects the NNLF1 baseline tool as a comparison. On the basis of NNLF1, a mode selection process including a chrominance information adjustment mode is performed on inter-frame coded frames (i.e., non-I frames). Under the general test conditions of random access and low delay B configuration, the general sequence specified by the Joint Video Experts Group (JVET) is tested. The anchor for comparison is NNLF1. The results are shown in Tables 1 and 2.

[0187] Table 1: Performance of this example compared to the baseline NNLF1 under the Random Access configuration

[0188]

[0189]

[0190] Table 2 Performance of this embodiment compared with baseline NNLF1 under Low Delay B configuration

[0191]

[0192] The meanings of the parameters in the table are as follows:

[0193] EncT: Encoding Time, 10X% means that after the reference row sorting technology is integrated, the encoding time is 10X% compared with before integration, which means that the encoding time increases by X%.

[0194] DecT: Decoding Time, 10X% means that after the reference row sorting technology is integrated, the decoding time is 10X% compared with before integration, which means that the decoding time increases by X%.

[0195] ClassA1 and ClassA2 are test video sequences with a resolution of 3840x2160, ClassB is a test sequence with a resolution of 1920x1080, ClassC is 832x480, ClassD is 416x240, and ClassE is 1280x720; ClassF is a screen content sequence with several different resolutions.

[0196] Y, U, V are the three color components, and the columns where Y, U, V are located represent the BD-rate of the test results on Y, U, V. rate) indicator, the smaller the value, the better the encoding performance.

[0197] Analyzing the data in the two tables, we can see that by introducing the optimization method of chrominance information adjustment, the coding performance can be further improved based on NNFL1, especially in the chrominance component. The residual adjustment of this embodiment has little effect on the decoding complexity.

[0198] The method of this embodiment can also be used to select the NNLF mode for intra-coded frames (I frames).

[0199] In the aforementioned embodiments, a scheme for adjusting chromaticity information when performing NNLF on the reconstructed image is proposed. In the codec, the coding performance of NNLF is further improved by swapping the order of the U component and the V component of the chromaticity information input to the NNLF filter. This also shows that there is a certain correlation between the characteristic information of the U component and the V component. At present, there is a lack of schemes to study the impact of chromaticity information adjustments such as swapping the input order of U and V during training on the NNLF network model. After research, the embodiments of the present disclosure provide some embodiments of the training method for chromaticity information fusion. When training the model, strategies for chromaticity information adjustment such as UV swapping are introduced to expand the training data, which can further improve the performance of the NNLF model. Some embodiments of the present disclosure also use the NNLF model trained by the chromaticity information fusion training method to perform NNLF on the reconstructed image, that is, the chromaticity information fusion mode is used to perform NNLF in the above embodiments to further optimize the performance of NNLF.

[0200] The NNLF filter of a model trained using the chrominance information fusion training method (also referred to herein as the chrominance information fusion module, denoted as the CF model) can be implemented based on a baseline NNLF (such as NNLF1 and NNLF2) or based on a new NNLF implementation. When implemented based on a baseline NNLF, the CF module is equivalent to retraining the baseline NNLF model; when implemented based on a new NNLF, the CF module can be trained based on a model with a new network structure to obtain a new model.

[0201] The location of the Chroma Fusion (CF) module in the codec is shown in Figure 7. It can replace the NNLF-B filter in the figure. The use of the CF module does not depend on whether the DB, SAO, and ALF switches are enabled or not.

[0202] On the encoder side, the rate-distortion cost of performing NNLF on the reconstructed image is compared between the NNLF-A filter and the (retrained / new) NNLF-B filter trained using chroma fusion. This determines which NNLF filter (corresponding to different NNLF modes) to use. The selected NNLF mode is then encoded into the bitstream for the decoder to read. On the decoder side, after determining the NNLF mode actually used by the encoder, the decoder uses that NNLF mode to perform NNLF processing on the reconstructed image.

[0203] An embodiment of the present disclosure further provides a neural network-based loop filtering method, which is applied to a video decoding device. As shown in FIG13 , the method includes:

[0204] Step S510: decoding a first flag of a reconstructed image, where the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image;

[0205] Step S520, determining an NNLF mode to be used when performing NNLF on the reconstructed image according to the first flag, and performing NNLF on the reconstructed image according to the determined NNLF mode;

[0206] Among them, the NNLF mode includes a first mode and a second mode, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: expanded data obtained after specifying adjustments to the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

[0207] Tests have shown that using a model trained with augmented data obtained through chroma information adjustment (i.e., the model used in the chroma information fusion mode) can affect the filtering effect of the NNLF, thereby affecting encoding performance. This embodiment can select an optimal mode from the chroma information fusion mode and other modes based on the rate-distortion cost, thereby improving encoding performance.

[0208] In an exemplary embodiment of the present disclosure, the specified adjustment of the chromaticity information of the reconstructed image in the original data includes any one or more of the following adjustment methods:

[0209] Interchange the order of the two chrominance components of the reconstructed image in the original data;

[0210] The weighted average value and the square error value of the two chrominance components of the reconstructed image in the original data are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

[0211] In an exemplary embodiment of the present disclosure, the model used in the first mode is trained using the original data (ie, supplementary data obtained after chromaticity information adjustment is not used as training data).

[0212] In an exemplary embodiment of the present disclosure, the reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the first flag is a picture-level syntax element or a block-level syntax element.

[0213] In an exemplary embodiment of the present disclosure, there is one first mode and one or more second modes; the network structures of the models used in the first mode and the second mode are the same or different.

[0214] In an exemplary embodiment of the present disclosure, the second mode is the chrominance information fusion mode; the first flag is used to indicate whether the chrominance information fusion mode is used when performing NNLF on the reconstructed image;

[0215] Determining the NNLF mode to be used when performing NNLF on the reconstructed image according to the first flag includes: when the first flag indicates the use of the chrominance information fusion mode, determining that the second mode is used when performing NNLF on the reconstructed image; when the first flag indicates not to use the chrominance information fusion mode, determining that the first mode is used when performing NNLF on the reconstructed image.

[0216] In an example of this embodiment, there are multiple second modes; the method also includes: when the first flag indicates the use of the chrominance information fusion mode, continuing to decode the second flag, the second flag including index information of a second mode to be used; and, determining to use the second mode when performing NNLF on the reconstructed image based on the second flag.

[0217] An embodiment of the present disclosure further provides a video decoding method, which is applied to a video decoding device, including: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image, as shown in FIG14 :

[0218] Step S610, determining whether NNLF allows chrominance information fusion;

[0219] Step S620: When NNLF allows chrominance information fusion, perform NNLF on the reconstructed image according to the method described in any embodiment of the present disclosure that is used at the decoding end and can use the chrominance information fusion mode.

[0220] This embodiment can select an optimal mode from the chrominance information fusion mode and other modes according to the rate-distortion cost, thereby improving the encoding performance.

[0221] In an exemplary embodiment of the present disclosure, it is determined that NNLF allows chrominance information fusion when one or more of the following conditions are met:

[0222] Decode the chroma information fusion permission flag at the sequence level, and determine that the NNLF allows the chroma information fusion according to the chroma information fusion permission flag;

[0223] The chroma information fusion permission flag of the decoded image level is determined based on the chroma information fusion permission flag to determine whether the NNLF allows the chroma information fusion.

[0224] In an example of this embodiment, the sequence header in the following table may be used:

[0225]

[0226] In the table, cf_enable_flag indicates the flag for enabling chroma information fusion at the sequence level.

[0227] In an exemplary embodiment of the present disclosure, when NNLF does not allow chrominance information fusion, decoding of the chrominance information fusion permission flag is skipped, and NNLF is performed on the reconstructed image using the first mode.

[0228] In an example of this embodiment, the sequence header shown in the following table may be used:

[0229]

[0230] The cf_enable_flag in the table is the sequence-level chroma information fusion permission flag. When cf_enable_flag is 1, the following semantics are defined: picture_cf_enable_flag indicates the image-level chroma information fusion usage flag, which is the first flag mentioned above. When picture_cf_enable_flag is 1, it indicates that the chroma information fusion mode is used when performing NNLF on the reconstructed image. When there are multiple chroma information fusion modes (such as using multiple models trained with supplementary data), the index picture_cf_index indicates which chroma information fusion mode is used. When picture_cf_enable_flag is 0, it indicates that the chroma information fusion mode is not used when performing NNLF on the reconstructed image, that is, the first mode is used. When cf_enable_flag is 0, the decoding and encoding of picture_cf_enable_flag and picture_cf_index are skipped.

[0231] An embodiment of the present disclosure further provides a loop filtering method based on a neural network, which is applied to a video encoding device. As shown in FIG15 , the method includes:

[0232] Step 710, calculating the rate-distortion cost of performing NNLF on the input reconstructed image using the first mode, and the rate-distortion cost of performing NNLF on the reconstructed image using the second mode;

[0233] Step 720: Perform NNLF on the reconstructed image using a mode with the lowest rate-distortion cost between the first mode and the second mode;

[0234] Among them, the first mode and the second mode are both set NNLF modes, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: the expanded data obtained after the specified adjustment of the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

[0235] This embodiment can select an optimal mode from the chrominance information fusion mode and other modes according to the rate-distortion cost, thereby improving the encoding performance.

[0236] In an exemplary embodiment of the present disclosure, the specified adjustment of the chromaticity information includes any one or more of the following adjustment methods:

[0237] Interchange the order of the two chrominance components of the reconstructed image;

[0238] The weighted average value and the square error value of the two chrominance components of the reconstructed image are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

[0239] In an exemplary embodiment of the present disclosure, the model used in the first mode is trained using the original data;

[0240] The reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the network structures of the models used in the first mode and the second mode are the same or different.

[0241] In an exemplary embodiment of the present disclosure, there is one first mode, and there are one or more second modes;

[0242] The calculating the rate-distortion cost cost1 of performing NNLF on the input reconstructed image using the first mode includes: obtaining a first filtered image output after performing NNLF on the reconstructed image using the first mode, and calculating the cost1 according to the difference between the first filtered image and the corresponding original image;

[0243] Calculating a rate-distortion cost cost2 of performing NNLF on the reconstructed image using the second mode includes: for each second mode, obtaining a second filtered image output after performing NNLF on the reconstructed image using the second mode, and calculating the cost2 of the second mode based on a difference between the second filtered image and the original image. When there are multiple second modes, the costs2 of the multiple second modes are compared with the cost1 of the first mode to determine a minimum rate-distortion cost.

[0244] In an exemplary embodiment of the present disclosure, the method further includes: using the filtered image obtained by performing NNLF on the reconstructed image using the mode with the smallest rate-distortion cost as the filtered image output after performing NNLF on the reconstructed image.

[0245] An embodiment of the present disclosure further provides a video encoding method, which is applied to a video encoding device, including: performing the following processing when performing a neural network-based loop filtering (NNLF) on a reconstructed image, as shown in FIG16 :

[0246] Step 810: When NNLF allows chroma information fusion, perform NNLF on the reconstructed image according to the method described in any embodiment of the present disclosure applied to the encoding end and capable of using a chroma information fusion mode;

[0247] Step 820: Encode a first flag of the reconstructed image, where the first flag includes information about the NNLF mode used when performing NNLF on the reconstructed image.

[0248] This embodiment can select an optimal mode from the chrominance information fusion mode and other modes according to the rate-distortion cost, thereby improving the encoding performance.

[0249] In an exemplary embodiment of the present disclosure, the first flag is a picture-level syntax element or a block-level syntax element.

[0250] In an exemplary embodiment of the present disclosure, it is determined that NNLF allows chrominance information fusion when one or more of the following conditions are met:

[0251] Decode the chroma information fusion permission flag at the sequence level, and determine that the NNLF allows the chroma information fusion according to the chroma information fusion permission flag;

[0252] The chroma information fusion permission flag of the decoded image level is determined based on the chroma information fusion permission flag to determine whether the NNLF allows the chroma information fusion.

[0253] In an exemplary embodiment of the present disclosure, the method further includes: when it is determined that NNLF does not allow chroma information fusion, using the first mode to perform NNLF on the reconstructed image, and skipping the encoding of the chroma information fusion permission flag.

[0254] An embodiment of the present disclosure further provides a video encoding method that can be implemented in the NNLF filter of the encoding end. When the encoding end enters the NNLF process, the following steps are performed:

[0255] Step a) Determine whether NNLF in the current sequence is allowed to use the chroma information fusion mode based on the sequence-level chroma information fusion enable flag cf_enable_flag. If cf_enable_flag is "1", attempt to use the chroma information fusion mode for NNLF in the current sequence and skip to step b). If cf_enable_flag is "0", use the first mode for NNLF and skip encoding the first flag. End.

[0256] Step b) for the reconstructed image of the current frame of the current sequence, performing NNLF on the reconstructed image using a first mode, such as inputting information into a model of a baseline NNLF filter to obtain a first filtered image output by the filter;

[0257] Step c) for the reconstructed image of the current frame of the current sequence, perform NNLF on the reconstructed image using the chrominance information fusion mode, that is, input the encoded information into the (retrained / new) NNLF model in the chrominance information fusion mode, and obtain a second filtered image output by the filter;

[0258] Step d) Calculate the rate-distortion cost C based on the difference between the first filtered image and the original image NNLF , calculate the rate-distortion cost C based on the difference between the second filtered image and the original image CF ; Compare the two rate-distortion costs, if C CF <C NNLF , it is determined to use the chrominance information fusion mode, that is, the second mode, to perform NNLF on the reconstructed image, and the second filtered image is used as the filtered image finally output after the reconstructed image is subjected to NNLF filtering; if C CF ≥C NNLF , it is determined to use the first mode to perform NNLF on the reconstructed image, and the first filtered image is used as the filtered image finally output after the reconstructed image is subjected to NNLF filtering;

[0259] If there are multiple chromaticity information fusion modes (the chromaticity information adjustment and fusion mode mentioned above is also considered a chromaticity information fusion mode), there are multiple C CF Participate in the comparison.

[0260] Step e) encoding the first flag picture_cf_enable_flag of the reconstructed image of the current frame into the bitstream. If there are multiple chrominance information fusion modes, the index picture_cf_index indicating the chrominance information fusion mode to be used is also encoded into the bitstream;

[0261] If the current frame has been processed, the next frame is loaded for processing. Similarly, the above processing can also be performed on the current slice of the current frame, or the current block in the current frame.

[0262] In some embodiments of the present disclosure, the same UV swap adjustment strategy is also introduced during the training process of the NNLF model, and the NNLF model is trained by swapping the input order of U and V. As mentioned above, the NNLF model is trained by swapping the input order of U and V. It can be based on the baseline NNLF implementation or the new NNLF implementation. The main difference lies in the different NNLF network structure, which has no effect on model training. It is uniformly represented by NNLF_VAR in this article. For the NNLF_VAR model, the embodiments of the present disclosure have made improvements in both training and encoding and decoding testing.

[0263] An embodiment of the present disclosure provides a training method for a loop filter model based on a neural network, as shown in FIG17 , including:

[0264] Step 910 , performing a specified adjustment on the chrominance component of the reconstructed image in the original data used for training to obtain expanded data;

[0265] Step 920: Train the NNLF model using the expanded data, or the expanded data and the original data as training data.

[0266] In an exemplary embodiment of the present disclosure, the specified adjustment of the chromaticity information of the reconstructed image in the original data includes any one or more of the following adjustment methods:

[0267] Interchange the order of the two chrominance components of the reconstructed image in the original data;

[0268] The weighted average value and the square error value of the two chrominance components of the reconstructed image in the original data are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

[0269] In an exemplary embodiment of the present disclosure, for a certain NNLF_VAR, the order of information input to the network may be {recY, recU, recV, predY, predU, predV, baseQP, sliceQP, slicetype, ...}. Therefore, it is necessary to collect the above data through encoding and decoding testing, align it with the label data of the original image, and organize and package it into data pairs patchA{train, label} in sequence as training data for the network. The data format of patchA, i.e., the original data, is shown in the following pseudo code:

[0270] patchA:

[0271] {

[0272] train:{recY,recU,recV,predY,predU,predV,baseQP,sliceQP,slicetype,…}

[0273] label:{orgY,orgU,orgV}

[0274] }

[0275] By introducing the concept of chromaticity information fusion, the input order of U and V can be interchanged. Therefore, this embodiment expands the training data (on the basis of the original patchA, patchB, i.e., supplementary data, is added. The expanded data pair patch is shown in the following pseudo code.

[0276] patchA:

[0277] {

[0278] train:{recY,recU,recV,predY,predU,predV,baseQP,sliceQP,slicetype,…}

[0279] How to get label:{orgY,orgU,orgV}?

[0280] }

[0281] patchB:

[0282] {

[0283] train:{recY,recV,recU,predY,predV,predU,baseQP,sliceQP,slicetype,…}

[0284] label:{orgY,orgV,orgU}

[0285] }

[0286] It can be seen that for the expanded patchB, the data of U and V are swapped at the input position, and the corresponding positions of the label data also need to be swapped.

[0287] This embodiment uses both the original data and the supplementary data as training data. In other embodiments, only the supplementary data may be used as training data to train the NNLF mode in the chrominance information fusion mode.

[0288] When performing NNLF on the reconstructed image using the chroma information fusion mode (i.e., NNLF mode), the input chroma information can be adjusted, such as by swapping the order of U and V (in this case, the aforementioned chroma information adjustment and fusion mode is used). The input and output are shown in Figure 12B. However, the input chroma information can also be left unadjusted (in this case, the non-chroma information adjustment mode). The input and output are shown in Figure 12A.

[0289] In the training phase, the proposed scheme trains NNLF_VAR according to the augmented training data generated in (1). The specific training process is similar to the existing NNLF, using multiple rounds of training, which will not be detailed here.

[0290] After the NNLF model is trained, it is deployed in the codec for codec testing. As mentioned above, the network model obtained based on chrominance information fusion training has different input and output information forms in the chrominance information fusion mode with UV interchange and the chrominance information fusion mode without UV interchange. During the codec test, these two modes can be tried separately, and their respective rate-distortion costs can be calculated separately. These two rate-distortion costs are compared with the rate-distortion cost of using the first mode (equivalent to using baseline NNLF filtering) to determine whether to use the chrominance information fusion mode and which chrominance information fusion mode to use (when using the chrominance information fusion mode).

[0291] The disclosed embodiment proposes a training method for the NNLF mode with chrominance information fusion (the training data at least includes supplementary data). An interchange adjustment strategy is used for the U component and V component of the chrominance information input to the NNLF filter for NNLF training and encoding and decoding testing, thereby further optimizing the performance of the NNLF.

[0292] In the above embodiment, for the Y, U, and V components, a joint rate-distortion cost is used to make a decision, so the three components share one first flag (or first flag and index). In other embodiments, a more refined processing can be attempted, where different first flags (or first flags and indexes) are set for different components. The syntax elements of this method are shown in the following table:

[0293]

[0294] When the sequence-level chroma information fusion enable flag cf_enable_flag is 1, the following semantics are defined: picture-level chroma information fusion enable flag picture_cf_enable_flag[compIdx]

[0295] When the picture-level chroma information fusion enable flag picture_cf_enable_flag[compIdx] is 1, the following semantics are defined: the index of the picture-level chroma information fusion method picture_cf_index[compIdx]

[0296] An embodiment of the present disclosure further provides a code stream, wherein the code stream is generated by the video encoding method described in any embodiment of the present disclosure.

[0297] One embodiment of the present disclosure also provides a neural network-based loop filter, as shown in Figure 11 , comprising a processor and a memory storing a computer program. When the processor executes the computer program, it is capable of implementing the neural network-based loop filtering method described in any embodiment of the present disclosure. As shown in the figure, the processor and memory are connected via a system bus. The loop filter may also include other components such as memory and a network interface.

[0298] An embodiment of the present disclosure further provides a video decoding device, see FIG11 , comprising a processor and a memory storing a computer program, wherein the processor can implement the video decoding method as described in any embodiment of the present disclosure when executing the computer program.

[0299] An embodiment of the present disclosure further provides a video encoding device, see FIG11 , comprising a processor and a memory storing a computer program, wherein the processor can implement the video encoding method as described in any embodiment of the present disclosure when executing the computer program.

[0300] The processor of the above-mentioned embodiment of the present disclosure may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a microprocessor, etc., or other conventional processors; the processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), discrete logic or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or other equivalent integrated or discrete logic circuits, or a combination of the above devices. That is, the processor of the above-mentioned embodiment may be any processing device or device combination that implements the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. If the embodiments of the present disclosure are partially implemented in software, the instructions for the software may be stored in a suitable non-volatile computer-readable storage medium, and one or more processors may be used to execute the instructions in hardware to implement the methods of the embodiments of the present disclosure. The term "processor" used herein may refer to the above-mentioned structure or any other structure suitable for implementing the technology described herein.

[0301] An embodiment of the present disclosure further provides a video encoding and decoding system, see FIG1A , which includes the video encoding device described in any embodiment of the present disclosure and the video decoding device described in any embodiment of the present disclosure.

[0302] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, can implement the neural network-based loop filtering method as described in any embodiment of the present disclosure, or implement the video decoding method as described in any embodiment of the present disclosure, or implement the video encoding method as described in any embodiment of the present disclosure.

[0303] In one or more exemplary embodiments above, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium that facilitates the transfer of a computer program from one place to another, such as according to a communication protocol. In this way, a computer-readable medium may generally correspond to a non-transitory tangible computer-readable storage medium or a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this disclosure. A computer program product may include a computer-readable medium.

[0304] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection may also be referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient (transient) media, but rather refer to non-transient tangible storage media. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, or Blu-ray disc, among others, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

Claims

1. A neural network-based loop filtering (NNLF) method, applied to a video decoding device, comprising: Decode a first flag of a reconstructed image, where the first flag includes information about an NNLF mode used when performing NNLF on the reconstructed image; Determining, according to the first flag, an NNLF mode to be used when performing NNLF on the reconstructed image; performing NNLF on the reconstructed image according to the determined NNLF mode; Among them, the NNLF mode includes a first mode and a second mode, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: expanded data obtained after specifying adjustments to the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

2. The method according to claim 1, wherein: The specified adjustment of the chromaticity information of the reconstructed image in the original data includes any one or more of the following adjustment methods: Interchange the order of the two chrominance components of the reconstructed image in the original data; The weighted average value and the square error value of the two chrominance components of the reconstructed image in the original data are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

3. The method according to claim 1, wherein: The model used in the first mode is trained using the original data; The reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the first flag is a picture-level syntax element or a block-level syntax element.

4. The method according to claim 1, wherein: There is one first mode, and there are one or more second modes; The network structures of the models used in the first mode and the second mode are the same or different.

5. The method according to claim 1, wherein: The second mode is the chrominance information fusion mode; the first flag is used to indicate whether the chrominance information fusion mode is used when performing NNLF on the reconstructed image; Determining the NNLF mode to be used when performing NNLF on the reconstructed image according to the first flag includes: when the first flag indicates the use of the chrominance information fusion mode, determining that the second mode is used when performing NNLF on the reconstructed image; when the first flag indicates not to use the chrominance information fusion mode, determining that the first mode is used when performing NNLF on the reconstructed image.

6. The method according to claim 5, wherein: There are multiple types of the second mode; The method also includes: when the first flag indicates the use of the chrominance information fusion mode, continuing to decode a second flag, the second flag including index information of a second mode to be used; and determining, based on the second flag, to use the second mode when performing NNLF on the reconstructed image.

7. A video decoding method, applied to a video decoding device, comprising: When performing neural network-based loop filtering (NNLF) on the reconstructed image, the following processing is performed: When NNLF allows fusion of chrominance information, the reconstructed image is subjected to NNLF according to the method according to any one of claims 1 to 6.

8. The method according to claim 7, wherein: NNLF is determined to allow chrominance information fusion when one or more of the following conditions are met: Decoding a sequence-level chroma information fusion permission flag, and determining that NNLF allows chroma information fusion according to the chroma information fusion permission flag; The chroma information fusion permission flag of the decoded image level is determined based on the chroma information fusion permission flag to determine whether the NNLF allows the chroma information fusion.

9. The method according to claim 7, wherein: The method further includes: when NNLF does not allow chrominance information fusion, skipping decoding of the chrominance information fusion permission flag, and performing NNLF on the reconstructed image using the first mode.

10. A neural network-based loop filtering (NNLF) method, applied to a video encoding device, comprising: Calculating a rate-distortion cost of performing NNLF on the input reconstructed image using the first mode, and a rate-distortion cost of performing NNLF on the reconstructed image using the second mode; Determine to use a mode with the lowest rate-distortion cost between the first mode and the second mode to perform NNLF on the reconstructed image; Among them, the first mode and the second mode are both set NNLF modes, the second mode includes a chromaticity information fusion mode, and the training data used in training the model used in the chromaticity information fusion mode includes: the expanded data obtained after the specified adjustment of the chromaticity information of the reconstructed image in the original data; or includes the original data and the expanded data.

11. The method according to claim 10, wherein: The specified adjustment of the chromaticity information includes any one or more of the following adjustment methods: Interchange the order of the two chrominance components of the reconstructed image; The weighted average value and the square error value of the two chrominance components of the reconstructed image are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.

12. The method according to claim 10, wherein: The model used in the first mode is trained using the original data; The reconstructed image is a reconstructed image of a current frame, a current slice, or a current block; and the network structures of the models used in the first mode and the second mode are the same or different.

13. The method according to claim 10, wherein: There is one first mode, and there are one or more second modes; The calculating the rate-distortion cost cost1 of performing NNLF on the input reconstructed image using the first mode includes: obtaining a first filtered image output after performing NNLF on the reconstructed image using the first mode, and calculating the cost1 according to the difference between the first filtered image and the corresponding original image; The calculation of the rate-distortion cost cost2 of performing NNLF on the reconstructed image using the second mode includes: for each of the second modes, obtaining a second filtered image output after performing NNLF on the reconstructed image using the second mode, and calculating the cost2 of the second mode based on the difference between the second filtered image and the original image.

14. A video encoding method, applied to a video encoding device, comprising: When performing neural network-based loop filtering (NNLF) on the reconstructed image, the following processing is performed: In the case where NNLF allows fusion of chrominance information, performing NNLF on the reconstructed image according to the method according to any one of claims 10 to 13; A first flag of the reconstructed image is encoded, where the first flag includes information of an NNLF mode used when performing NNLF on the reconstructed image.

15. The method according to claim 14, wherein: The first flag is a picture-level syntax element or a block-level syntax element.

16. The method according to claim 14, wherein: NNLF is determined to allow chrominance information fusion when one or more of the following conditions are met: Decoding a sequence-level chroma information fusion permission flag, and determining that NNLF allows chroma information fusion according to the chroma information fusion permission flag; The chroma information fusion permission flag of the decoded image level is determined based on the chroma information fusion permission flag to determine whether the NNLF allows the chroma information fusion.

17. The method according to claim 14, wherein: The method further includes: when it is determined that NNLF does not allow chroma information fusion, using the first mode to perform NNLF on the reconstructed image and skipping encoding of the chroma information fusion permission flag.

18. The method according to claim 14, wherein: The second mode is a chroma information fusion mode; the first flag is used to indicate whether the chroma information fusion mode is used when performing NNLF on the reconstructed image; The encoding of the first flag of the reconstructed image includes: when performing NNLF on the reconstructed image using the first mode, setting the first flag to a value indicating not using the chrominance information fusion mode; When the second mode is used to perform NNLF on the reconstructed image, the first flag is set to a value indicating the use of the chrominance information fusion mode.

19. The method according to claim 18, wherein: There are multiple types of the second mode; The encoding of the first flag of the reconstructed image further includes: after setting the first flag to a value indicating the use of the chrominance information fusion mode, continuing to encode a second flag, wherein the second flag contains index information of a second mode with the minimum rate-distortion cost.

20. A code stream, wherein The code stream is generated by the video encoding method according to any one of claims 14 to 19.

21. A loop filter based on a neural network, comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, it can implement the neural network-based loop filtering method as described in any one of claims 1 to 6 and 10 to 13.

22. A video decoding device comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, it can implement the video decoding method according to any one of claims 7 to 9.

23. A video encoding device comprising a processor and a memory storing a computer program, wherein: When the processor executes the computer program, it is capable of implementing the video encoding method according to any one of claims 14 to 19.

24. A video encoding and decoding system, wherein: The method comprises the video encoding device according to claim 23 and the video decoding device according to claim 22.

25. A non-transitory computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, it can implement the neural network-based loop filtering method as described in any one of claims 1 to 6, 10 to 13, or implement the video decoding method as described in any one of claims 7 to 9, or implement the video encoding method as described in any one of claims 14 to 19.

26. A training method for a neural network-based loop filter (NNLF) model, wherein: include: Performing specified adjustments on the chrominance components of the reconstructed image in the original data used for training to obtain augmented data; The NNLF model is trained using the expanded data, or the expanded data and the original data, as training data.

27. The method of claim 26, wherein: The specified adjustment of the chromaticity information of the reconstructed image in the original data includes any one or more of the following adjustment methods: Interchange the order of the two chrominance components of the reconstructed image in the original data; The weighted average value and the square error value of the two chrominance components of the reconstructed image in the original data are calculated, and the weighted average value and the square error value are used as the chrominance information input in the NNLF mode.