Quantization parameter adaptive convolutional neural network loop filter and its construction method
By introducing FQAM and SQAM mechanisms into the convolutional neural network loop filter, the quantization parameter QP is introduced into the model, which solves the problem of insufficient generalization ability in the existing technology, and realizes adaptive processing of different quantization parameters, significantly improves the filtering ability and reduces resource requirements.
Patent Information
- Application Number
- CN202210174226.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-24
AI Technical Summary
Existing neural network loop filters have insufficient generalization capabilities when facing different quantization parameters, resulting in poor performance under different data sets and quantization noise conditions, and excessive resources required for training and deployment.
A convolutional neural network loop filter adapted to quantization parameters is designed. By introducing FQAM and SQAM mechanisms, the quantization parameter QP is introduced into the convolutional neural network to realize the adaptive processing of different quantization parameters. FQAM performs adaptive processing in the frequency domain, while SQAM performs adaptive processing in the air domain, improving the generalization ability of the model.
This design significantly improves the filtering capability of the convolutional neural network loop filter under different quantization parameters, can effectively remove image noise under different quantization noises, and reduce the demand for resources.
Smart Images

Figure CN114596223B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network loop filters, and specifically relates to a convolutional neural network loop filter that is adaptive to quantization parameters and a construction method thereof. Background Art
[0002] Neural network-based loop filters have achieved great success in recent years. They can effectively remove common artificial imprints in the image / video encoding and decoding process, such as block effects, ringing, and Gibbs effects. However, compared with traditional post-processing modules, they often have significantly higher complexity. There is always room for optimization in the generalization ability of neural network filters, and models trained on a single data set may be difficult to use for other data sets. For encoders, using different quantization parameters means that different quantization noises will appear on the reconstructed image / video. Training a model for each quantization noise will take up a lot of training resources and storage resources in actual deployment and calling of the model, which makes this method impractical. In response to this problem, the present invention designs a new neural network filter, which, relying on the designed FQAM and SQAM mechanisms, can have significantly excellent filtering capabilities on different quantization parameters. Summary of the invention
[0003] The purpose of the present invention is to propose a convolutional neural network loop filter with strong filtering capability and adaptive to quantization parameters and a construction method thereof.
[0004] The quantization parameter adaptive convolutional neural network loop filter provided by the present invention introduces the quantization parameter QP into the convolutional neural network to improve the generalization ability of the convolutional neural network for different QPs. The specific introduction method is as follows:
[0005] (1) FQAM (Frequency QP Adaptive Mechanism), that is, frequency domain QP adaptive mechanism. Each layer of convolutional neural network features is regarded as the extracted specific frequency information. Each layer of features is multiplied by a coefficient related to QP to achieve attenuation or enhancement of the layer of features; when QP changes, the features of each layer will also be attenuated or enhanced with the change of QP, thereby achieving complete absorption of QP information;
[0006] (2) SQAM (Spatial QP Adaptive Mechanism), that is, spatial QP adaptive mechanism. Use convolution and QP information to generate attention in the spatial domain. Attention can generate different weights for different regions, and through this weight, it acts on the original features to achieve adaptive attenuation or enhancement of the original features. This enhancement depends not only on convolution but also on QP information. This enables the model to improve the QP adaptive ability in the spatial domain;
[0007] (3) Based on FQAM and SQAM, a convolutional neural network loop filter is constructed, such as Figure 1 As shown in the figure, it can effectively remove quantization noise from images with different quantization noises. Specifically, the structure contains two input tensors and one output tensor. The input tensors are the input image and the quantization parameter QP, respectively. The size and color space format of the image are not limited. For the output tensor, it represents the image obtained after the filter enhancement, and its size remains consistent with the input image. The middle of the model is composed of a convolutional network, FQAM, and SQAM. Specifically, after the image is input, a direct edge will be drawn to the output. In addition, it will also pass through the first Octave convolutional network to obtain two separate feature information, which are recorded as high-frequency and low-frequency information respectively. Then, it passes through several (for example, the number of networks is set to 24) residual network structures. Each residual network structure contains Octave convolutional network, FQAM, Octave convolutional network, and FSQAM in turn to obtain output. The residual network also contains a direct edge to help the gradient back propagation during training. Another tensor QP of the structural input is used to guide the QP information of FQAM and FSQAM here to help the model adapt to changes in different QP information. Due to the influence of Octave convolution, the convolution features continue to flow between high frequency and low frequency, so that the information is fully learned and utilized. After the stacked residual network is completed and the output is obtained, an Octave convolution network is finally included to transform the tensor back to the original image size. This tensor is then added back to the input image to obtain the final enhanced image.
[0008] The present invention provides a method for constructing a convolutional neural network loop filter that is adaptive to quantization parameters. The method introduces the quantization parameter QP into the convolutional neural network to improve the generalization ability of the convolutional neural network for different QPs. The specific steps are as follows.
[0009] (I) Building FQAM (Frequency Domain QP Adaptive Mechanism)
[0010] From the perspective of frequency domain, a modular convolutional neural network model is constructed, and the quantization parameter QP is integrated into it; first consider a simple filtering model:
[0011]
[0012] Among them, w is the filter parameter, y is the input of the filter, that is, the distorted image, is the output of the filter, i.e. the reconstructed image. Naturally, we use Fourier transform to get its equation in the frequency domain, i.e., time domain convolution is equal to frequency domain multiplication:
[0013]
[0014] F(.) represents Fourier transform; assuming that this filter has good filtering performance, the reconstructed image is approximately equal to the original image, that is, in the frequency domain:
[0015]
[0016] In order to cope with the changing QP, it is necessary to generalize it to the general case, that is, to hope that the modified filter can have better performance under a wider range of quantization noise input conditions. The modified filter parameters are recorded as w′ and the change in quantization noise is recorded as ε. At this time, the reconstructed image has changed from the original Became
[0017]
[0018] Also perform Fourier transform on equation (4) to obtain its frequency domain form:
[0019]
[0020] We hope to find such a w′ that the reconstruction of w′ The loss between the original input x and the reconstruction is the lowest. For the convenience of solving, the mean square error is used here. The mean square error between can be written as:
[0021]
[0022] According to Pasval's theorem, the distortion in the time domain is the same as the distortion in the frequency domain. Therefore, L in formula (6) can also be written in the frequency domain form, and the following expansion can be obtained:
[0023]
[0024] By taking the partial derivative of equation (7) with respect to F(w′), we can obtain that F(w′) with zero derivative can be expressed as follows:
[0025]
[0026] The first term represents the original filter in the frequency domain of formula (1), and the second term is obtained by modifying the noise, which is called the noise impact factor. The original filter can be changed by the noise impact factor here. This formula needs to be simplified:
[0027]
[0028] Considering a certain frequency domain, the strength of the filter and the strength of the original signal can actually be considered unchanged. So this formula can be approximated as k i F(n i ), and we know that F(n i ) is proportional to Qstep 2 So we can use trainable parameters θ i To express the multiple relationship here, Qstep 2 Introduced into the model:
[0029]
[0030] Considering that the complexity of doing so is too high, the present invention adopts a simplified strategy, directly approximating that the convolution layer represents the selection of the frequency domain, so that the calculation can be reduced from the square order of the feature to the feature order. At this time, the time domain form of the filter operation can be written as:
[0031]
[0032] Therefore, the FQAM algorithm is derived, where w is the original filter parameter and 1+θQstep on the denominator 2 It represents the attenuation coefficient. As QP changes, Qstep will also change. The changed Qstep is used as input to affect the filtering performance of the model. The intuitive working diagram can be referred to Figure 2 .
[0033] (II) Building SQAM (Spatial QP Adaptive Mechanism)
[0034] FQAM can only perform attenuation and enhancement related to QP at the channel level. The present invention proposes SQAM in the spatial domain as a supplement to enhance the ability of FQAM in the spatial domain. Different regions also respond differently to QP.
[0035] Derivation process of SQAM:
[0036] First, the spatial features of the input image y′ are extracted. Here, the MaxPool and AvgPool operations [1] can be used, namely, the maximum value pooling and mean pooling operations:
[0037] s(y′)={MaxPool(y′);AvgPool(y′)}, (12)
[0038] After extracting the spatial feature s(y′) of y′, this feature is passed through an FQAM to achieve the fusion of QP information. Here, the fused output uses a sigmoid activation function to achieve output limiting:
[0039]
[0040] The fusion output can be used as the intensity information in the spatial domain, so the element-level dot product is used to multiply this intensity information back to the original input y′, and finally the reconstructed pixel of SQAM is output. For a visual reference diagram, see Figure 3 .
[0041]
[0042] (III) Design of a convolutional neural network loop filter with adaptive quantization parameters based on FQAM and SQAM
[0043] like Figure 1 As shown in Figure 1, a neural network filter is obtained by stacking multiple layers of convolution. The output of the i-th layer is y i , use function f(·) to represent the change relationship between two consecutive layers, as shown in formula (15):
[0044] y i+1 =f(y i ,QP), (15)
[0045] The features of each layer are composed of high-frequency information and low frequency information It consists of two parts, and the notation of formula (16) is used to represent the characteristics of each layer:
[0046]
[0047] Each layer is transformed using octave convolution and FQAM and SQAM, as shown in equations (17) and (18):
[0048]
[0049]
[0050] Where LReLU represents the leaky relu activation function [2]; the trainable parameter θ of the model can be expressed as and A collection of:
[0051]
[0052] In the experiment, Qstep 2 is replaced by the following formula:
[0053] Qstep 2 =2 (QP-32) / 3 , (20)
[0054] Thus, the conversion between each layer in the convolutional neural network is completed, and then the entire model is completed. Therefore, it is possible to make the convolutional neural network adaptive to different quantization parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of the method of the present invention.
[0056] Figure 2 Schematic diagram of FQAM.
[0057] Figure 3 Schematic diagram of SQAM. DETAILED DESCRIPTION
[0058] This section will introduce how to use the designed quantization parameter adaptive convolutional neural network loop filter. The filter is placed at the end of the encoder as a complementary module, and it is fed with QP and a distorted image with loss, which can output a filtered image.
[0059] As for the implementation details of the model, the input of the model is QP, and what is needed in the formula is the square of Qstep, which needs to be converted in the model. The conversion formula refers to formula (20). Because Qstep^2 is multiplied by a trainable parameter, this step makes different Qsteps multiplied by the same constant, which does not affect the performance of the model. Therefore, we normalize it and use 2 (QP-32) / 3 To represent Qstep^2. Secondly, because both A and B in the formula are greater than 0, its approximate trainable parameters should also be greater than 0. Constrain the parameters in formula (19) to make them always greater than 0. There are two methods here. The first is to use the re-parameterization technique k=exp(k), and the second is to directly truncate the parameters. The second method is used here. In terms of parameter selection, the α of the intermediate convolution layer of the designed neural network filter is 0.25, and the parameter of LeakyReLU is selected as 0.2. The number of model feature layers is set to 64, so the intermediate convolution layer contains 48 original size images and 16 downsampled features to extract low-frequency information. The number of blocks is set to 24 to obtain powerful feature extraction capabilities. Considering that SQAM uses convolutional layers compared to FQAM, it requires more computing resources. Therefore, as shown in the figure, FSQAM is used at the end of a block, and FQAM is used internally.
[0060] About integrating neural network into general video codec, such as VTM. The loop filters SAO and ALF embedded in VTM will introduce additional bit rate to improve the image quality, so the CNN-filter is placed between DB and SAO, thereby providing SAO and ALF with higher quality reconstructed pixels, thereby reducing the number of bits consumed by SAO and ALF. The model designed by the present invention can be used for different QPs. We only used the luminance component of the I frame to train the model, and the neural network also has good generalization ability for the chrominance component.
[0061] References
[0062] [1]https: / / en.wikipedia.org / wiki / Convolutional_neural_network#Pooling_layer
[0063] [2]https: / / en.wikipedia.org / wiki / Rectifier_(neural_networks).
Claims
1. A method for constructing a convolutional neural network loop filter with adaptive quantization parameters, characterized in that: By introducing the quantization parameter QP into the convolutional neural network, the generalization ability of the convolutional neural network for different QPs is improved. The specific steps are: (a) Constructing FQAM; From the perspective of frequency domain, a convolutional neural network model is constructed and the quantization parameter QP is integrated into it; first consider a simple filtering model: Among them, w is the filter parameter, y is the input of the filter, that is, the distorted image, is the output of the filter, that is, the reconstructed image; use Fourier transform to get its equation in the frequency domain, that is, the time domain convolution is equal to the frequency domain multiplication: F(.) represents Fourier transform; assuming that this filter has good filtering performance, the reconstructed image is approximately equal to the original image, that is, in the frequency domain: In order to cope with the changing QP, it is extended to the general case, that is, it is hoped that the modified filter has better performance under a wider range of quantization noise input conditions; here the modified filter parameter is recorded as w ′ , the change of quantization noise is recorded as ε; at this time, the reconstructed image has changed from the original Became Also perform Fourier transform on equation (4) to obtain its frequency domain form: I hope to get such a w ′ , so that w ′ The reconstructed The loss between the original input x and the reconstruction is the lowest; for the convenience of solving, the mean square error is used here, the original input x and the reconstruction The mean square error between them can be written as: According to Pasval's theorem, the distortion in the time domain is the same as the distortion in the frequency domain. Therefore, L in formula (6) can also be written in the frequency domain form, and the following expansion is obtained: For equation (7), find the value of F(w ′ ) can be used to find the partial derivative of F(w ′ ), expressed as the following formula: The first term represents the original filter in the frequency domain of formula (1), and the second term is obtained by modifying the noise, which is called the noise impact factor. By changing the original filter with the noise impact factor here, the formula is simplified: Considering a specific frequency domain, the strength of the filter and the strength of the original signal are actually unchanged, so this formula can be approximated as k i F(n i ), and it is known that F(n i ) is proportional to Qstep 2 So we use the trainable parameter θ i To express the multiple relationship here, Qstep 2 Introduced into the model: Considering that the complexity of doing so is too high, a simplified strategy is adopted, and the convolution layer is directly approximated to represent the selection of the frequency domain, thereby reducing the calculation from the square order of the feature to the feature order; at this time, the time domain form of the filter operation is written as: Therefore, the FQAM algorithm is derived, where w is the original filter parameter and 1+θQstep on the denominator 2 It represents the attenuation coefficient. As QP changes, Qstep will also change. The changed Qstep is used as input to affect the filtering performance of the model. (II) Construction of SQAM SQAM in airspace is used as a supplement to enhance the capability of FQAM in airspace, so that different areas can respond differently to QP. The construction process of SQAM is as follows: First, the input image y ′ The spatial features are extracted, and MaxPool and AvgPool operations are used here, namely maximum pooling and mean pooling operations: s(and ′ )={MaxPool(y ′ );AvgPool(and ′ )},(12) After extracting y ′ The spatial characteristics s(y ′ ), this feature is passed through an FQAM to achieve the fusion of QP information; here the fused output uses a sigmoid activation function to achieve output limiting: The fusion output is used as the intensity information in the spatial domain. The element-wise multiplication is used to multiply this intensity information back to the original input y ′ The final output is the reconstructed pixel of SQAM (III) Constructing a convolutional neural network loop filter that is adaptive to quantization parameters based on FQAM and SQAM A neural network filter is obtained by stacking multiple layers of convolution; the output of the i-th layer is y i , use function f(·) to represent the change relationship between two consecutive layers, as shown in formula (15): y i+1 =f(y i ,QP),(15) The features of each layer are composed of high-frequency information and low frequency information It consists of two parts, and the notation of formula (16) is used to represent the characteristics of each layer: Each layer is transformed using octave convolution and FQAM and SQAM, as shown in equations (17) and (18): Among them, LReLU represents the leakyrelu activation function; the trainable parameter θ of the model is expressed as and A collection of: Qstep 2 is replaced by the following formula: Qstep 2 =2 (QP-32) / 3 ,(20) This completes the conversion between each layer in the convolutional neural network, and then completes the entire model; thus, the convolutional neural network is made adaptive to different quantization parameters.
2. A convolutional neural network loop filter with adaptive quantization parameters obtained by the construction method of claim 1; characterized in that: It contains two tensor input ports and one tensor output port; the input tensors are the input image and the quantization parameter QP respectively; for the output tensor, it represents the image obtained after the filter enhancement, and its size remains consistent with the input image; the middle of the loop filter model is composed of a convolutional network, FQAM, and SQAM; specifically, after the image is input, a direct connection is drawn to the output; in addition, it will also pass through the first Octave convolutional network to obtain two separate feature information, which are recorded as high-frequency and low-frequency information respectively, and then pass through several residual network structures, each of which contains Octave convolutional network, FQAM, Octave The e convolutional network and FSQAM are output, and the residual network also contains a direct edge to help the gradient back propagation during training; another tensor QP of the structural input is used to guide the QP information of FQAM and FSQAM here to help the model adapt to changes in different QP information; due to the influence of Octave convolution, the convolution features continue to flow between high frequency and low frequency, so that the information can be fully learned and utilized; after the stacked residual network is output, it finally contains an Octave convolutional network to transform the tensor back to the original image size, and this tensor is then added back to the input image to obtain the final enhanced image.
Citation Information
Patent Citations
Convolutional neural network medical CT image denoising method based on residual error learning
CN109978778A
Loop filtering method based on edge enhanced residual network
CN111541894A