A loop filtering method of deep neural network based on multi-information fusion
By introducing a multi-information fusion loop filtering method in the VVC standard, using a variety of information in the MIIN network fusion encoding process, the problems of cumbersome and unsatisfactory loop filtering calculations in the prior art are solved, and more efficient information utilization and video quality improvement are achieved.
Patent Information
- Application Number
- CN202211615020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-15
AI Technical Summary
The loop filtering technology in the existing video encoding and codec standard VVC is cumbersome in the calculation process and the filtered reconstructed frame quality is not ideal, and it fails to make full use of the intermediate information in the encoding process.
A multi-information fusion loop filtering method is proposed. By building a MIIN network, multiple information in the encoding process are fused into network input, and a loop filtering module is embedded in the VVC standard, and the network is trained using the training set to improve the filtering effect.
It effectively eliminates the block effect, ringing effect and color shift in the reconstructed frame, improves the subjective and objective quality of the video, and reduces the code rate at the same quality, simplifies network complexity.
Smart Images

Figure CN115941978B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field related to video coding technology, and in particular to a deep neural network loop filter design based on multi-information fusion and a construction method thereof. Background Art
[0002] With the continuous improvement of hardware performance and the iterative development of network technology, coupled with the continuous improvement of people's demand for video content, videos are now developing in the direction of ultra-high definition, wide color gamut, and panoramic view. Its massive video storage and stable video transmission have brought new challenges to video coding related technologies. Compared with the previous generation of video coding standard HEVC, the latest generation of video coding standard VVC reduces the bit rate by about 50% under the premise of the same perceived quality. In particular, the loop filtering technology in the video coding standard not only improves the quality of the current encoded image, but also provides a higher quality reference image for the encoding of subsequent frames.
[0003] Luma mapping with chroma scaling (LMCS) is a new technology added to the VVC video codec standard. It achieves good reconstruction results in both SDR and HDR videos by adaptively modifying the distribution of coded samples to improve coding efficiency. It mainly includes two functional components: luma mapping and chroma scaling. Luma mapping is an in-loop mapping method based on an adaptive piecewise linear model. Its basic idea is to improve compression efficiency by adjusting the dynamic range of the input signal at a given bit depth. Chroma scaling is a chroma residual scaling based on luma information. Its purpose is to compensate for the problems caused by luma adjustment by adjusting the chroma residual information value in the chroma coding block.
[0004] The deblocking filter (DBF) is mainly used to remove the blocking effect that occurs during the encoding process. The blocking effect occurs because the video encoding process is based on blocks. As processing units, blocks are independent of each other, resulting in smoother textures on both sides of the reconstructed image block boundary, but discontinuity in the pixel values on both sides. When the images on both sides of the boundary of the coding block have a strong correlation and the image texture is relatively smooth, this discontinuity in pixels forms a "block effect" in the human eye. The deblocking filter DBF first determines the boundary type of the image, and then "corrects" the pixel values near the pseudo-boundary formed by the blocking effect. This process is performed by setting the filtering parameters, which depend on the boundary strength information of the coding block.
[0005] The sample adaptive filter (SAO) is placed after the DBF because in the process of lossy compression encoding, a certain degree of ringing effect will appear after the jump signal is compressed. In order to eliminate this artifact, SAO compensates for some special positions in the reconstructed samples, such as peaks, inflection points, and valleys, to reduce the difference between the reconstructed samples and the original samples, thereby achieving the purpose of filtering. In VVC, according to the different characteristics of the reconstructed image, SAO can be divided into two types, edge compensation EO and sideband compensation BO, and different filtering types can be selected through the control parameters in the CTU.
[0006] In VVC, the geometry transformation-based ALF (GALF) is integrated to reduce its complexity and is used to remove artifacts and distortions generated in the encoding stage. Traditional ALF and GALF are collectively referred to as ALF. ALF minimizes the MSE between the reconstructed image and the original image by using Wiener filtering. It acts after the sample adaptive filter SAO. The specific process is to first classify the blocks and then use filters with different coefficients for filtering operations.
[0007] The loop filtering module can effectively improve the subjective and objective quality of the video. The smaller the compression distortion of the filtered reconstructed frame and the closer it is to the original frame, the more advantageous it is for the encoding of subsequent frames. Specifically, the filtered image frame may be used in the encoding process of subsequent image frames. The loop filtering technology in the existing video codec standard VVC can reduce the compression distortion of the reconstructed frame to a certain extent, but the calculation process is cumbersome and the quality of the filtered reconstructed frame is not ideal.
[0008] Artificial intelligence uses deep neural networks to extract and analyze data features, and has achieved remarkable results in the field of computer vision, with excellent performance in low-level visual tasks such as image super-resolution and noise reduction. At present, the research and application of the combination of filtering modules and deep learning in video coding and decoding can be divided into two categories. One is loop filtering, which specifically refers to the use of neural networks to replace the original filtering modules in VVC coding to improve coding performance. The other is out-of-loop filtering, which specifically refers to the neural network processing of the decoded video after conventional coding and decoding to achieve the filtering effect. Although the existing neural network-based loop filtering has improved the quality of reconstructed frames to a certain extent, it does not fully utilize the intermediate information generated in the encoding process, such as prediction information, residual information, and partition information in the encoding process, resulting in limited effect of the neural network-based loop filtering module and unsatisfactory quality of the restored reconstructed frames.
[0009] In view of the defects and shortcomings of the prior art, the present invention proposes a multi-information fusion loop filtering method, which fully considers the influence of the LMCS tool and the data characteristics between different components, and adopts different information as network input for different components, so that the information utilization in the encoding process is more sufficient; in addition, the method uses the partition information generated in the encoding process as input, and considers the characteristics of block-by-block encoding processing in the video compression process. Based on the above analysis, the method can efficiently utilize and fuse various information in the encoding process to achieve the goal of fully utilizing various information generated in the middle. Under the premise of roughly the same gain, the complexity of the present invention is reduced compared with other deep learning methods due to the use of richer encoding information and more efficient fusion methods. Summary of the invention
[0010] In order to solve the above problems, the present invention proposes a method and system for a loop filter module in the VVC coding standard, which relates to the field of image processing technology. The method is to fuse various information in the encoding process as the input of the network, build a network model, and then train the network using a training set. After the network model converges, it is embedded in the loop filter module in the VVC standard. Specifically, a loop filter solution that can be applied to the video coding standard VVC includes the following steps:
[0011] S1. Construct training data set and validation data set: The training data set and validation data set are generated with the help of DIV2K image data set. First, convert the images in the DIV2K data set from RGB color space to YUV color space, and then use the All Intra encoding configuration of VTM 14.0, the reference software of the video codec standard VVC, to compress each image under the conditions of QP 22, 27, 32, 37, and 42, and then save the reconstruction information and partition information that have passed the LMCS module but not the DBF filter module, and save the residual information and prediction information that have not passed the LMCS module, and obtain non-overlapping 128×128 information blocks according to the preset partitioning method, and then separate the three YUV channel components of the information block. In order to ensure the effectiveness of the network for feature learning, calculate the PSNR of each reconstructed block, and eliminate the information block groups with PSNR greater than 50 and PSNR less than 15. After the above operations, for the luminance component, the training set and the validation set corresponding to the QP of 22, 27, 32, 37, and 42 are obtained, and each set of data includes the original information block, the reconstructed information block, the partition information block, and the prediction information block of the luminance component; for the chrominance component, the training set and the validation set corresponding to the QP of 22, 27, 32, 37, and 42 are obtained, and each set of data includes the original information block, the reconstructed information block, the partition information block, and the residual information block of the chrominance component.
[0012] S2. Build a MIIN (Multi-Information Integration Network) network, use the 5 types of training sets under different QP conditions obtained in step S1 to train the networks corresponding to the luminance component and the chrominance component respectively, generate 5 QP models of the luminance component and the chrominance component respectively, and then determine the optimal hyperparameters according to the performance of each model on the corresponding validation set, and select the respective optimal models, and finally obtain the 5 QP models corresponding to the luminance component and the chrominance component.
[0013] S3. Use the LibTorch library to convert the model obtained in step S2 into a C++ available type, and use the C++API therein to embed the converted network model into the reference software VTM 14.0 provided by the video codec standard VVC. When encoding the standard video sequence provided by JVET, first turn off the DBF filter module and SAO filter module in the loop filter, and use the default configuration for the rest. For the luminance component, first use the reconstructed image after passing through the LMCS module as the main input, and save the prediction information and partition information of the intermediate process, and then select the corresponding trained and converged network model in step S2 according to the QP value, and input the pre-prepared information into this network model. The image output by the network model is the luminance component image after filtering through the network MIIN. Similarly, save the residual information and partition information in the chrominance component encoding process, and then select the corresponding trained and converged network model in step S2 according to the QP value, and input the pre-prepared information into the network model. The image output by the network model is the chrominance component image after filtering through the network MIIN. Then use the above-mentioned image as the input of the ALF filter module to complete the subsequent encoding operation.
[0014] The details of the MIIN network structure described in S2 are as follows:
[0015] (1) For the luminance component and the chrominance component, the difference lies in the different input information. Specifically, the input information used by the luminance component is reconstruction information, prediction information, and partition information, while the input information used by the chrominance component is reconstruction information, residual information, and partition information. The network structure used by the luminance component and the chrominance component is consistent. Specifically, the network structure can be divided into three parts. The following is a further explanation of the structural details in the network.
[0016] (2) The first part is the multi-information fusion module. This part can be specifically divided into two layers, one is the information fusion layer and the other is the Inception Block. Assume that the brightness reconstruction information is The chromaticity reconstruction information is Then the reconstruction information branch in the information fusion layer can be expressed as:
[0017]
[0018]
[0019] Where f2 represents the operation of the convolutional layer and ReLU activation function layer corresponding to the reconstruction information branch, D L1 represents the output of the information fusion layer of the reconstruction information branch in the brightness Luma model, D C1 Represents the output of the information fusion layer of the reconstruction information branch in the Chroma model.
[0020] The input of Inception Block is D L1 and D C1 In this process, the brightness component and the chrominance component are processed in the same way, unified as D1, assuming that the output is D IB , then the Inception Block operation process can be expressed as:
[0021]
[0022] Among them, * represents the convolution operation, represents a convolution kernel of size i×i, Represents the Concat cascade operation. Further, assuming that the prediction information, residual information and partition information are represented by x pred 、x resi and x part , then the multi-information fusion module can be expressed as:
[0023]
[0024]
[0025] Where f1 represents the convolutional layer and ReLU activation function operations corresponding to the brightness prediction information and chrominance residual information branches, f3 represents the convolutional layer and ReLU activation function operations corresponding to the partition information branch, and D L2 Denotes the output of the multi-information fusion module in the brightness Luma model, D C2 Represents the output of the multi-information fusion module in the Chroma model.
[0026] (3) The second part is the residual feature aggregation module. In the subsequent steps, the information processing method of each component is consistent, so the brightness component and the chrominance component are no longer distinguished. The input of this module is the output D2 of the first part. The main part of this module is composed of several RFA blocks cascaded together. Assuming that the RFA blocks module is represented by RFAB, the residual feature aggregation module can be expressed as:
[0027]
[0028] in Denotes a convolution kernel of size 1×1, and D3 is the output of the residual feature aggregation module. In RFAB, several RFA modules are cascaded and concat to obtain the output of the RFAB module. Specifically, it is composed of 4 RFA modules connected in order, where the output of the first three RFAs has two branches, one branch is used as the input of the next RFA, and the other branch is directly sent to the end of the RFAB module, and then connected to the output of the last RFA through Concat, and then a 1×1 convolution is used to fuse these features.
[0029] The RFA module is implemented by combining three residual blocks through the attention mechanism. Assume that the input and output of the dth RFA module are RFA d-1 and RFA d+1 , the inputs of the three residual blocks in the RFA module are represented by R1, R2 and R3 respectively, then this process can be expressed by the following formula:
[0030] R1=RFA d-1
[0031]
[0032]
[0033] in and Represent the residual module and the attention module respectively. Further, the operation process of the RFA module can be expressed as:
[0034]
[0035] (4) The third part is the output reconstruction module. This module uses a 3×3 convolution kernel to adjust the dimension. Its output is the filtered residual information, which is finally added to the original reconstruction information to obtain the output of the entire network. Assuming that the input of the output processing module is D4 and the final output of the network is Y, then:
[0036] Y=w k3 *(D4)+x rec o
[0037] The beneficial effects of the present invention are as follows: the present invention uses a filtering network MIIN to replace the DBF module and SAO module of the loop filter in the video coding standard VVC, and effectively eliminates the block effect, ringing effect, color shift and other distortions in the reconstructed frame by utilizing the information in the coding and decoding process, thereby improving the subjective quality and objective quality of the final output image; the multi-information fusion scheme used by the present invention, compared with the existing neural network-based loop filtering scheme, first considers the influence of the LMCS filtering tool and the difference between components, and uses different information as network input for different components, that is, the network corresponding to the luminance component uses prediction information as an input information, and the chrominance component uses residual information as an input information, so that the use of information is more reasonable; secondly, the block-based coding characteristics are considered, and the partition information is used as another input information. Compared with other methods based on deep neural networks, the present invention has a lower network complexity and a shorter overall filtering time when the objective quality of the filtered video is close. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to make the purpose, technical solution and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration:
[0039] Figure 1 It is a flow chart of the loop filtering module of the present invention;
[0040] Figure 2 This is a schematic diagram of the MIIN network structure;
[0041] Figure 3 This is a schematic diagram of the Inception Block structure;
[0042] Figure 4 It is a schematic diagram of the RFA structure;
[0043] Figure 5 It is the schematic diagram of RB module;
[0044] Figure 6 Schematic diagram of CA module. Specific implementation plan
[0045] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0046] The loop filter part in the video codec standard VVC includes: an adaptive loop shaper (LMCS), a deblocking filter (DBF), a sample adaptive offset filter (SAO) and an adaptive loop filter (ALF). The present invention is based on the loop filter part of the video codec standard VVC, and adopts a neural network MIIN to replace the original DBF filter module and SAO filter module. Figure 1 The technical solution of the present invention mainly includes the following steps:
[0047] A. Construction of training data set and validation data set: The training data set and validation data set are generated with the help of DIV2K image data set. First, the pictures in the DIV2K data set are converted from RGB color space to YUV color space. Then, the All Intra encoding configuration in VTM 14.0, the reference software of the video codec standard VVC, is used to encode and compress each picture under the conditions of QP 22, 27, 32, 37, and 42. Then, the reconstruction information and partition information that have passed through the LMCS module but not the DBF filter module are saved, and the residual information and prediction information that have not passed through the LMCS module are saved. According to the preset partitioning method, non-overlapping 128×128 information blocks are obtained, and the three YUV channel components of the information block are separated. In order to ensure the effectiveness of the network for feature learning, the PSNR of each reconstructed block is calculated, and the information block groups with PSNR greater than 50 and PSNR less than 15 are eliminated. After the above operations, for the luminance component, the training set and the validation set corresponding to the QP of 22, 27, 32, 37, and 42 are obtained, and each set of data includes the reconstruction information block, the partition information block, and the prediction information block of the luminance component; for the chrominance component, the training set and the validation set corresponding to the QP of 22, 27, 32, 37, and 42 are obtained, and each set of data includes the reconstruction information block, the partition information block, and the residual information block of the chrominance component.
[0048] B. Build the MIIN network, and use the five types of training sets under different QP conditions obtained in step A to train the networks corresponding to the luminance component and the chrominance component respectively, and generate corresponding models of the five QPs of the luminance component and the chrominance component respectively. Then, determine the optimal hyperparameters according to the performance of each model on the corresponding validation set, and select the respective optimal models, and finally obtain the models of 5 QPs corresponding to the luminance component and the chrominance component.
[0049] MIIN network structure is as follows Figure 2As shown in the figure, it can be divided into three parts. The first part is the input information fusion module. The three input information all pass through a convolution layer. After the reconstruction information passes through an additional Inception Block, the three branches are connected through the Concat layer. The second part is the residual feature aggregation module, which first passes through a convolution layer, then passes through the RFAB module, and then passes through a convolution layer with a convolution kernel size of 1×1. The RFAB module is composed of 4 RFA blocks connected in order. The third part is the output reconstruction module, which uses a 3×3 convolution kernel to reconstruct the residual of the image and add it to the original reconstruction information to obtain the final filtered image.
[0050] C. Use the LibTorch library to convert the model obtained in step B into a C++ usable type, and use the C++ API therein to embed the converted network model into the reference software VTM14.0 provided by the video codec standard VVC. When encoding the standard test video sequence provided by JVET, first turn off the DBF filter module and SAO filter module in the loop filter, and use the default configuration for the rest. For the luminance component, first use the reconstructed image after the LMCS module as the main input, and save the prediction information and partition information of the intermediate process, and then select the corresponding trained and converged network model in step B according to the QP value, and input the pre-prepared information into this network model, and the output image is the filtered luminance component image after the network MIIN. Similarly, save the residual information and partition information in the chrominance component encoding process, and then select the corresponding trained and converged network model in step B according to the QP value, and input the pre-prepared information into the network model, and the output image is the filtered chrominance component image after the network MIIN. Then use the above-mentioned image as the input of the ALF filter module to complete the subsequent processing.
[0051] The training process details of step B are as follows: For all network models, a two-stage training method combining L1 loss function and L2 loss function is adopted, that is, the first stage of training adopts L1 loss function, and the second stage of training adopts loss function, which are respectively:
[0052]
[0053]
[0054] Where u represents the pixel index, U represents the set of pixels, x(u) represents the processed pixel value, and y(u) represents the real pixel value. In addition, considering the impact of the quantization parameter on the image quality, the setting of the learning rate will be classified. The network model with a smaller QP, that is, when QP is 22, 27, and 32, uses a slower learning rate of 1e-5 for update; the network model with a larger QP, that is, when QP is 37 and 42, uses a faster learning rate of 1e-4 for update.
[0055] D. The input of the MIIN network varies according to the components. For the luminance component, the main input uses the luminance component information of the reconstructed image, and the auxiliary information is the partition information and prediction information of the luminance; for the chrominance component, the main input uses the chrominance component information of the reconstructed image, and the auxiliary information is the partition information and residual information of the chrominance. The output of the MIIN network is the difference image ΔF between the original frame F and the reconstructed frame F' of each component learned by the neural network. The residual information and the reconstructed frame image are added to obtain the final filtered image F L ',Right now:
[0056] F L ′=F′+ΔF=F′+MIIN(F′)
[0057] The present invention provides a neural network-based loop filtering method that uses a multi-information fusion method to achieve good filtering effects on both the luminance component and the chrominance component, and improves the overall video quality. Aiming at the problems existing in the video coding standard VVC and the existing neural network-based loop filtering method, considering the data characteristics of each component, using multi-information fusion and an effective network structure, it can not only effectively eliminate the distortion of the reconstructed frame and improve the subjective and objective quality of the video, but also make the system stable and robust.
[0058] In order to better demonstrate the technical feasibility of the solution of the present invention, a simulation example is used to illustrate:
[0059] In this experiment, VTM 14.0 is used as the experimental platform. The test sequence is the A1, A2, B, C, D, E, F test set in the JVET CTC standard test sequence. The encoding mode is All Intra mode. The comparison indicators include the BD-rates of each video sequence. Table 1 is the experimental results.
[0060] The benchmark settings for this experimental result are based on CTC, and VTM 14.0 opens all loop filter modules by default; the test group settings for the experiment are also based on CTC, except that the DBF and SAO filter modules are closed, and the MIIN neural network is added between LMCS and ALF. When BD-rates is a negative value, it indicates that the bit rate is reduced and the encoding efficiency is improved under the premise of the same reconstruction quality; when BD-rates is a positive value, it indicates that the bit rate is increased and the encoding efficiency is reduced under the premise of the same reconstruction quality.
[0061] Table 1 Experimental results
[0062] Y U V Class A1 -4.57% -7.01% -9.26% Class A2 -3.61% -8.54% -9.19% Class B -3.43% -8.90% -9.41% Class C -3.05% -10.44% -10.53% Class E -4.66% -8.35% -10.02% Overall -3.77% -8.78% -9.70% Class D -3.09% -9.26% -11.13% Class F -2.21% -7.37% -6.02%
Claims
1. A loop filtering method based on a deep neural network with multi-information fusion, which is based on the loop filtering part of the video codec standard VVC. The proposed multi-information fusion network MIIN (Multi-Information Integration Network) replaces the deblocking filter DBF and the sample adaptive compensation filter SAO in the VVC standard; the specific filtering method includes the following steps: Step 1. Construct training data set and validation data set: After converting 1000 images in the DIV2K data set to YUV, use VTM14.0 to compress each image under the conditions of QP 22, 27, 32, 37, and 42, and save the required data in the training set; for the convenience of training, all the generated frames are cropped into 128×128 information blocks, so that the samples of the training set and validation set are obtained; Step 2. The designed MIIN consists of three parts. The first part is the input information fusion module. The three input information are all passed through a convolution layer with a convolution kernel size of 3×3. After the reconstruction information is passed through an Inception Block, the three branches are connected through Concat. The second part is the residual feature aggregation module, which first passes through a convolution layer, then passes through the residual feature aggregation module RFAB, and then passes through a convolution layer. The RFAB module is composed of four residual feature aggregation blocks RFA connected in order. The third part is the output reconstruction module, which uses a 3×3 convolution kernel to reconstruct the residual of the image and add it to the original reconstruction information to obtain the final filtered image. Step 3: Use the 5 types of training sets under different QP conditions obtained in step 1 to train the networks corresponding to the luminance component and the chrominance component respectively, and generate 5 types of models with 5 QPs for the luminance component and the chrominance component respectively. Then, determine the hyperparameters according to the performance of each model on the corresponding validation set, that is, select two stages of training, L1 loss function is used in stage 1, and L2 loss function is used in stage 2; the learning rate is 1e-5 when QP is 22, 27, and 32, and the learning rate is 1e-4 when QP is 37 and 42; and select the optimal model based on this, and finally obtain the 5 QP models corresponding to the luminance component and the chrominance component; Step 4: Use the LibTorch library to convert the network model obtained in step 3 and embed it into VTM 14.
0. When encoding the standard test video sequence provided by JVET, first turn off the DBF module and SAO module in the loop filter, and use the default configuration for the rest. For the luminance component, first use the luminance information of the reconstructed image after LMCS as the main input, and save the prediction information and partition information of the intermediate process. Then, select the corresponding trained and converged network model in step 3 according to the QP value, and input the pre-prepared information into this network model. The output image is the filtered luminance component image after the network MIIN. The chrominance component processing process is similar to the luminance component, except that the prediction information of the intermediate process is replaced by residual information. At this point, the output of the network includes luminance component information and chrominance component information, and then use the above-obtained image as the input of the ALF filtering module to complete the subsequent filtering process.
Citation Information
Patent Citations
Progressive feature flow deep fusion network for monitoring video enhancement
CN112348766A
Loop filtering method of deep neural network suitable for low bit rate condition
CN114173130A