A video coding in-loop filtering method based on dynamic convolutional neural network
By adaptively adjusting the convolution kernel parameters and combining quantization parameters through dynamic convolutional neural networks, the problems of limited model performance and poor compression artifact processing in existing technologies are solved, achieving more efficient video encoding and image quality improvement.
Patent Information
- Application Number
- CN202411930424.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing deep learning-based video coding loop filtering methods cannot adaptively adjust convolution kernel parameters according to image content during inference, and fail to fully utilize prior information of quantization parameters, resulting in limited model performance and poor processing of compression artifacts.
A dynamic convolutional neural network is used to adaptively process different image contents and quantization parameters by dynamically adjusting the convolution kernel parameters and combining quantization parameters, dynamic fusion, dynamic backbone and dynamic reconstruction modules to improve the filtering effect.
The model's adaptability to different image contents and quantization parameters is improved, the reconstructed image quality and video coding efficiency are enhanced, and BD-rate gains of 0.47%, 2.49% and 2.13% are added, while the computational complexity is kept negligible.
Smart Images

Figure CN119835415B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video coding and compression, and in particular relates to a video coding loop filtering method based on a dynamic convolutional neural network. Background Art
[0002] With the widespread adoption of high-resolution and multi-format video content, video data traffic is growing exponentially, posing significant challenges to video storage and transmission. To address this challenge, improving video compression efficiency is crucial. The latest video coding standard, Versatile Video Coding (VVC), also known as H.266, has made significant progress in this regard. Compared to its predecessor, High Efficiency Video Coding (HEVC), or H.265, VVC not only improves compression efficiency but also excels in supporting multiple video formats and providing encoding flexibility.
[0003] The VVC standard still uses a block-based hybrid coding framework, which has proven its effectiveness in video coding. Loop filtering technology, a key component of this hybrid coding framework, effectively eliminates blocking artifacts caused by parameter discontinuities between adjacent coding blocks and ringing artifacts caused by high-frequency information loss, thereby ensuring optimal video quality. Although traditional loop filters have significantly improved the quality of reconstructed images, there is still room for further improvement.
[0004] In recent years, deep learning techniques, particularly convolutional neural networks (CNNs), have become an important tool for improving video quality due to their powerful capabilities in feature extraction and representation learning. Deep learning methods can more effectively capture prior knowledge in video images and apply it to the filtering process, further enhancing filtering effectiveness. Given the potential of deep learning in computer vision, the Joint Video Experts Team (JVET) has conducted extensive research on deep learning-based loop filtering techniques, aiming to further improve video quality and coding efficiency.
[0005] The technical solution of document L. Wang, X. Xu, and S. Liu, "EE1-1.1: neural network based in-loop filter with constrained storage and low complexity," Doc. JVET-Y0078, Jan. 2022 is that the reconstructed image after luminance mapping chroma scaling (LMCS) is input into the network, and then the output of the neural network filter is processed through adaptive loop filtering (ALF) and cross-component adaptive loop filtering (CCALF), wherein LMCS, ALF and CCALF are traditional loop filtering methods. The U and V channels of the input image are up-sampled in the pre-processing stage and down-sampled in the post-processing stage. However, in the inference process of the network of this scheme, the model cannot adaptively adjust the convolution kernel parameters according to different image contents, thereby limiting the performance of the model and possibly leading to a non-optimal result.
[0006] Document 2 R. Chang, L. Wang, X. Xu, and S. Liu, "EE1-1.1: More refinements on NN based in-loop filter with a single model," Doc. JVET-AC0194, Jan. 2023 gives a network structure that currently uses a deep learning method to replace the traditional loop filter with better performance, and uses a single model to solve compressed images under different quantization parameters. However, in the input stage of the model, simply splicing the quantization parameters with the feature maps cannot well utilize the prior knowledge of the quantization parameters, and thus cannot well handle different levels of compression artifacts caused by different quantization parameters.
[0007] In summary, the defects of the prior art are:
[0008] I. Most of the existing deep learning-based schemes optimize the filter parameters through training on a large amount of data, and after training, the parameters of the model will not change. In the inference process, the parameters of the convolution kernel cannot be adaptively adjusted according to the different image contents, which limits the expression ability of the model and produces a suboptimal result.
[0009] II. The prior art cannot well utilize the prior information of the quantization parameters. In order to solve the problem of different levels of compression artifacts caused by different quantization parameters, the prior art usually fills the quantization parameters into a matrix of the same size as the input image, and inputs them into the network together with the input image to guide the network to learn the parameters compressed by different quantization parameters. However, this method actually acts on the bias term of the convolution, and cannot guide the learning process of the convolution kernel weight, and fails to fully utilize the information of the quantization parameters. Summary of the Invention
[0010] To overcome the shortcomings of the prior art, the present invention aims to provide a video coding loop filtering method based on a dynamic convolutional neural network. This method dynamically adjusts convolution kernel parameters based on image content to address compression artifacts in different images. Furthermore, by combining quantization parameters with dynamic convolution, the method better utilizes prior information about the quantization parameters.
[0011] In order to achieve the above object, the technical solution adopted by the present invention is:
[0012] A video coding loop filtering method based on a dynamic convolutional neural network comprises the following steps:
[0013] Step 1: Extract multi-dimensional shallow features from the input reconstructed image (rec) and predicted image (pred);
[0014] Step 2: Combine the shallow features with the quantization parameters and reduce the size of the combined feature map by downsampling;
[0015] Step 3: The backbone of the network extracts deep features from the fused and reduced feature map, dynamically adjusts parameters based on the characteristics of the input data, and obtains the processed feature map;
[0016] Step 4: Integrate the processed feature maps and restore them to the resolution of the original input through upsampling operations to obtain the residual map. Add the residual map to the original reconstructed image to obtain the filtered image.
[0017] In the step 1,
[0018] The reconstructed image is obtained from the inside of the loop during the video encoding process, and is obtained by adding the quantized residual to the predicted image after dequantization and inverse transformation, without passing through the loop filtering module;
[0019] The predicted image is derived from the adjacent pixels or blocks that have been encoded within the current frame, and the spatial correlation of these pixels is used to generate the predicted value. Specifically, by referring to the encoded neighboring samples, a variety of prediction modes (such as DC, planar, angular prediction, etc.) are used to construct the predicted image.
[0020] The step 1 is specifically as follows:
[0021] The reconstructed image I of the input rec , predicted image I pred Use regular convolution F of size 3×3 respectively 3x3 Perform shallow feature extraction to obtain feature map F rec and F pred , denoted as F rec =F3x3 (I rec ),F pred =F 3x3 (I pred ), and then spliced with the input QPmap, the feature map F is obtained, that is, the shallow feature after multi-dimensional splicing; F = C(F rec ,F pred ,QP), where C is the splicing operation.
[0022] The conventional convolution is used to capture different levels of information of the input image, including edges, textures, and other low-level visual elements.
[0023] In step 1, before input, the U and V components of the reconstructed image (rec) and the predicted image (pred) are first upsampled using the nearest neighbor interpolation method so that the resolution of the U and V components matches the Y component. The Y, U, and V channels of the reconstructed image (rec) and the predicted image (pred) can all be input into the network at the same resolution.
[0024] The QP map is obtained by padding the current quantization parameter (QP) value to the same spatial resolution as the reconstructed frame (rec).
[0025] The step 2 is specifically as follows:
[0026] Use dynamic convolution with parallel convolution kernels to reduce the dimension and fuse the spliced shallow features;
[0027] The dynamic convolution is expressed as:
[0028] y=σ((a1·W1+…+a k W k )*x) (1)
[0029] Where σ is the activation function, W1…W k is the weight of the convolution kernel, k is the number of parallel convolution kernels, a1…a k is the attention weight, which is obtained by performing global average pooling, fully connected (FC) layer and Sigmoid layer operations on the input feature map x (here x is the feature map F obtained in step 1), and is expressed as:
[0030] a(x)=Sigmoid(FC(GlobalAveragePool(x))) (2)
[0031] Among them, a represents the operation of generating k attention weights;
[0032] Dynamic convolution is used to reduce the dimension and fuse the concatenated feature maps. This operation can be expressed and expanded as follows:
[0033] Y = {a(X; QP) · (W1,..., W k )} * {X; QP} + b (3)
[0034] = W(X; QP) * {X; QP} + b (4)
[0035] = W'(X; QP) * X + ∑W"(X; QP) · QP + b (5)
[0036] wherein a has the same meaning as represented in formula (2), W(X; QP) is equal to the dynamic convolution kernel a(X; QP) · (W1,..., W k )}, W', W" are the expression forms of the simplified convolution kernel, and b represents the bias term of the convolution layer;
[0037] The feature extracted by the dynamic convolution layer is further subjected to a nonlinear activation function PReLU and a regular convolution with a step of 2 to reduce the calculation amount of the backbone network, so as to obtain a down-sampled feature map, which is then sent to the backbone part of the network.
[0038] The backbone part of the network of step 3 comprises two types of dynamic residual blocks stacked in sequence.
[0039] Specifically, a first type of dynamic residual block is first stacked, and then a second type of dynamic residual block is stacked thereon to form a series structure.
[0040] The first type of dynamic residual block comprises a PReLU activation function, a dynamic convolution layer before the PReLU activation function, and two regular convolution layers after the PReLU activation function, and the dynamic convolution layer adaptively extracts features according to different input feature maps.
[0041] The second type of dynamic residual block comprises a PReLU activation function, a dynamic convolution layer before the PReLU activation function, and two dynamic convolution layers after the PReLU activation function.
[0042] The first type of dynamic residual block specifically operates as follows:
[0043] The input feature map is Y, and the output of the dynamic convolution layer is Z, and the above specific steps are represented as:
[0044] Z = PReLU(DConv4(Y));
[0045] DConv4() represents a dynamic convolution operation with 4 parallel convolution kernels, two conventional convolution layers, namely 1x1 convolution layer and 3x3 convolution layer; the 1x1 convolution layer is used to reduce the dimension of the feature map, reducing the computational burden by reducing the number of channels; the 3x3 convolution layer further performs feature extraction to capture more local information. The output R1 of the residual block is expressed as:
[0046] R1=Y+Conv 3x3 (Conv 1x1 (PReLU(DConv4(Y))))
[0047] Among them, Conv 1x1 and Conv 3x3 Represent 1x1 convolution layer and 3x3 convolution layer respectively;
[0048] The feature maps after the first type of dynamic residual blocks are stacked are processed by the second type of dynamic residual blocks to deal with more complex features. The specific operations of the second type of dynamic residual blocks are:
[0049] R2=R1+DConv 3x3 (DConv 1x1 (PReLU(DConv4(R1))))
[0050] Among them, DConv 3x3 and DConv 1x1 They represent dynamic convolution layers with convolution kernel sizes of 3x3 and 1x1 respectively, and each dynamic convolution layer also has 4 parallel convolution kernels.
[0051] The step 4 is specifically as follows:
[0052] The dynamic convolution layer adaptively adjusts the parameters of the convolution kernel according to the different input feature maps; after the dynamic convolution layer, the PixelShuffle layer is used for upsampling, enlarging the low-resolution feature map to the same size as the input image to obtain the residual map.
[0053] The process of the above dynamic convolution layer is expressed as:
[0054] Y=PixelShuffle(DConv8(X));
[0055] DConv8 represents a dynamic convolution operation with 8 parallel convolution kernels, and R represents the residual. Finally, it is added to the reconstructed image (rec) of the input image to obtain the output after network filtering: Output = rec + R
[0056] Beneficial effects of the present invention:
[0057] The dynamic filtering scheme proposed in the present invention is different from the previous neural network filters in that after training, the parameters of the convolution layer cannot be adaptively adjusted according to the different contents of the input feature map. Instead, it takes into account that different image contents and different quantization parameters require different degrees of filtering and is adaptively adjusted. For example, the dynamic fusion module uses dynamic convolution to process the features after the shallow features and the quantization parameter map (QP map) are spliced together, which can more effectively cope with different quantization parameters. The dynamic backbone module can dynamically adjust the parameters of the convolution kernel according to the different contents of the input feature map. The dynamic reconstruction module can more flexibly process different types of features to achieve better reconstruction effects. Through the collaborative work of the three modules of dynamic fusion, dynamic backbone and dynamic reconstruction, the present invention not only improves the adaptability of the model to different image contents and quantization parameters, but also improves the quality of the reconstructed image and the efficiency of video encoding.
[0058] Compared to existing neural network loop filtering methods, this invention achieves additional BD-rate gains of 0.47%, 2.49%, and 2.13% for the Y, U, and V components, respectively. Furthermore, the proposed dynamic filtering method adds negligible complexity and maintains inference time comparable to existing solutions, demonstrating the effectiveness of the proposed dynamic filtering network. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic diagram of the overall network structure of the present invention.
[0060] Figure 2 Schematic diagram of dynamic convolution of the present invention.
[0061] Figure 3 Schematic diagram of the dynamic residual block (DynamicResBlock) of the present invention. DETAILED DESCRIPTION
[0062] The present invention will be described in further detail below with reference to the accompanying drawings.
[0063] like Figure 1 As shown in FIG, the video coding loop filtering method based on dynamic convolutional neural network proposed in the present invention is composed of four parts:
[0064] Feature extraction, dynamic fusion, dynamic backbone, and dynamic reconstruction.
[0065] Step 1: Feature Extraction:
[0066] Extract multi-dimensional shallow features from the input reconstructed image (rec) and predicted image (pred) to increase the number of feature maps. This stage uses convolution operations to capture different levels of information in the input image, including edges, textures, and other low-level visual elements.
[0067] Step 2: Dynamic Fusion:
[0068] The shallow features obtained in the feature extraction stage are combined with the quantization parameters, and the size of the feature map is reduced by downsampling to reduce the computational complexity of the backbone part.
[0069] Step 3: Dynamic Backbone:
[0070] Further extracting deep features, building powerful feature expressions by stacking dynamic residual blocks, and dynamically adjusting parameters according to the characteristics of the input data are the core parts of the entire network.
[0071] Step 4: Dynamic Reconstruction:
[0072] The feature maps processed by the dynamic backbone are integrated and restored to the original input resolution through upsampling to obtain the residual map. The residual map is added to the original reconstructed image to obtain the filtered image.
[0073] In step 1, the network receives inputs including: a reconstructed frame (rec) before being processed by a conventional loop filter during VVC encoding, a predicted frame (pred) generated during VVC encoding, and a QP map (qp). The QP map is obtained by padding the current quantization parameter (QP) value to the same spatial resolution as the reconstructed frame (rec).
[0074] Since both the reconstructed frame (rec) and the predicted frame (pred) are in YUV420 format, this means that the resolution of the UV channel is only one-quarter of that of the Y channel. To ensure that the data input to the network has a consistent resolution, the U and V components are first upsampled using the nearest neighbor interpolation method before entering the network so that their resolution matches the Y component and is converted into the YUV444 format. After this processing, the Y, U, and V channels of rec and pred can all be input to the network at the same resolution and enhanced in the same process, reducing the number of required models.
[0075] The step 1 is specifically as follows:
[0076] Reconstructed image I of the input rec , predicted image I pred Use regular convolution F of size 3×3 respectively 3x3 Perform shallow feature extraction to obtain feature map F rec and F pred , which can be expressed as F rec =F 3x3 (I rec ),F pred =F 3x3 (I pred), and then concatenated with the input QP map to obtain the feature map F, F = C(F rec ,F pred ,QP), where C is the splicing operation.
[0077] The step 2 is specifically as follows:
[0078] The dynamic fusion module uses a dynamic convolution consisting of 8 parallel convolution kernels to reduce the dimension and fuse the spliced shallow features. The dynamic convolution can be expressed as:
[0079] y=σ((a1·W1+…+a k W k )*x) (1)
[0080] Where σ is the activation function, W1…W k is the weight of the convolution kernel, k is the number of parallel convolution kernels, which is set to 8 here. a1…a k is the attention weight, which is obtained by performing global average pooling, fully connected (FC) layer and Sigmoid layer operations on the input feature map x, such as Figure 2 As shown in the dotted box in , this process is expressed as:
[0081] a(x)=Sigmoid(FC(GlobalAveragePool(x))) (2)
[0082] Among them, a represents the operation of generating k attention weights;
[0083] In the loop filter based on ordinary convolutional neural network, the QP map is usually directly concatenated with the feature map and then convolved to achieve QP adaptability, which can be expressed as:
[0084] Y=W*{X;QP}+b.
[0085] Where W and b are weights and biases. {;} and * represent concatenation and convolution operations, respectively. This equation can be further expanded to:
[0086] Y=W′*X+W″*QP+b=W′*X+(∑W″·QP+b);
[0087] Where W′ and W″ are the weights of the corresponding feature map and QP, respectively.
[0088] After concatenating the QP map with the features extracted from the feature extraction along the channel dimension, using a conventional convolution operation is equivalent to adding a linear variable dependent on the QP to the convolution bias. However, it is the weights of the convolution kernel, not the bias, that dominate the convolutional neural network. Therefore, we propose to use dynamic convolution to operate on the concatenated feature map. This operation can be expressed and expanded as follows:
[0089] Y={a(X;QP)·(W1,…,W k )}*{X;QP}+b (3)
[0090] =W(X;QP)*{X;QP}+b (4)
[0091] =W′(X; QP)*X+∑w″(X; QP)·QP+b (5)
[0092] Where a has the same meaning as in Equation (2), and W(X; QP) is equal to the dynamic convolution kernel a(X; QP)·(W1,…,W k )}. From the expanded formula, we can see that the QP map not only acts on the bias term of the convolution, but also on the weight of the convolution. Therefore, it is more effective for QP adaptability;
[0093] Afterwards, the features extracted by the dynamic convolution layer are passed through a nonlinear activation function PReLU and a conventional convolution with a step size of 2 to reduce the computational complexity of the backbone network, obtain the downsampled feature map, and then send it to the backbone part of the network.
[0094] The backbone of the network in step 3 is composed of two types of dynamic residual blocks stacked together, such as Figure 3 shown.
[0095] Figure 3 (a) shows the first type of dynamic residual block, which consists of a dynamic convolution layer before the PReLU activation function. The PReLU activation function is often used to introduce nonlinearity, thereby improving the network's fitting ability. However, the activation function also imposes some restrictions on specific input features, which may lead to information loss. By applying dynamic convolution before the activation function, it can more fully learn and adjust in the original space of the input features, alleviating the limitations of the activation function. In addition, due to the characteristics of dynamic convolution, its convolution kernel parameters are dynamically dependent on the input feature map, and it can adaptively extract features based on the input feature map. Here, a dynamic convolution layer with 4 parallel convolution kernels is used, and the dimension of the feature map is increased from 64 to 160 to extract richer features, thereby enhancing the model's adaptability to different inputs.
[0096] Assume that the input feature map is Y and the output of the dynamic convolution layer is Z, then the above specific steps are expressed as:
[0097] Z = PReLU(DConv4(Y));
[0098] DConv4() represents a dynamic convolution operation with four parallel convolution kernels, which is the same as the dynamic convolution layer in step 2. This residual block also includes two conventional convolution layers: a 1x1 convolution layer and a 3x3 convolution layer. The 1x1 convolution layer is mainly used to reduce the dimension of the feature map, reducing the number of channels to reduce the computational burden; while the 3x3 convolution layer further performs feature extraction and can capture more local information. The output R1 of the residual block is expressed as:
[0099] R1=Y+Conv 3x3 (Conv 1x1 (RPeLU(DConv4(Y))))
[0100] Among them, Conv 1x1 and Conv 3x3 Representing 1x1 convolutional layer and 3x3 convolutional layer respectively. Through these designs, the residual block can not only effectively extract multiple features, but also enhance the flexibility and expressiveness of the model when processing different input data.
[0101] Figure 3 (b) shows the second type of dynamic residual block, which has a similar structure to the first type of residual block. It replaces the two regular convolution layers in (a) with dynamic convolution operations. This design makes the second type of dynamic residual block more adaptable and can dynamically adjust the parameters of the convolution kernel according to the complexity of the input feature map, thereby more effectively handling different degrees of distortion caused by compression with different quantization parameters. Although this design increases the flexibility and expressiveness of the model, it also brings higher computational complexity. Considering the balance between computational complexity and model performance, the number of the first type of dynamic residual blocks (DynamicResBlock1) and the second type of dynamic residual blocks (DynamicResBlock2) are set to 28 and 4, respectively.
[0102] This configuration allows the network to significantly enhance the model's expressiveness while maintaining high computational efficiency. Specifically, 28 DynamicResBlock1 blocks handle most feature extraction tasks, while 4 DynamicResBlock2 blocks focus on processing more complex features and compression artifacts, ensuring that the model performs well in different scenarios.
[0103] By stacking the two types of dynamic residual blocks, the backbone of the network is formed, which not only enhances the expressive power of the model, but also improves the adaptability and robustness of the model to different types of video content. This design enables the network to effectively reduce compression artifacts and improve image quality when processing high-resolution videos.
[0104] The similarity between the two residual block structures is that both use PReLU activation functions and place dynamic convolution layers before the activation functions. This design allows dynamic convolution to learn and adjust more fully in the original space of input features, reducing the limitations of activation functions on information; both retain the classic residual connection, which helps to alleviate the gradient vanishing problem in deep networks and promote information transmission. In terms of differences, the second type of dynamic residual block uses a full dynamic convolution layer, which gives it stronger adaptive ability to dynamically adjust the parameters of the convolution kernel according to the complexity of the input feature map, thus more effectively processing different degrees of distortion caused by different quantization parameters compression.
[0105] In step 4, the dynamic reconstruction part uses a dynamic convolution layer with 8 parallel convolution kernels and a PixelShuffle layer. Using a dynamic convolution layer in the reconstruction part allows the parameters of the convolution kernel to be adjusted adaptively according to the different input feature maps, thus more flexibly processing different types of features to achieve better reconstruction results.
[0106] The above process is represented as:
[0107] R = PixelShuffle(DConv8(R2));
[0108] DConv8 represents a dynamic convolution operation with 8 parallel convolution kernels, and R represents the resulting residual. Finally, the reconstructed image (rec) of the input image is added to obtain the output of the network filtered Output = rec + R The specific implementation is the same as the dynamic convolution layer in step 2.
[0109] Specifically, the 8 parallel convolution kernels can capture multi-scale and multi-directional information in the feature map, enhancing the model's ability to model complex image structures. By adaptively adjusting the convolution kernel parameters, it can more accurately restore the details and edge information of the image, reducing the impact of compression artifacts.
[0110] Following the dynamic convolution layer, the PixelShuffle layer is used for upsampling, scaling the low-resolution feature map to the same size as the input image. By reorganizing the pixels in the feature map, the PixelShuffle layer effectively restores the image's high-resolution details, making the output image visually clearer and more natural. This design not only improves the quality of the reconstructed image but also maintains high computational efficiency. This dynamic reconstruction design allows the model to more effectively reduce compression artifacts and enhance the visual quality of images when processing high-resolution video, while maintaining low computational complexity and ensuring the model's feasibility and efficiency in practical applications.
[0111] Figure 3 The dynamic fusion, dynamic backbone and dynamic reconstruction modules can adopt other technologies that can dynamically adjust network parameters, such as weight prediction or dynamic structure adjustment methods.
[0112] for Figure 1 The weighted values (a1, a2, …, ak) of the k parallel convolution kernels in dynamic convolution can be obtained through various methods, such as multi-layer perceptrons (MLPs) and adaptive average pooling, and dynamic convolution can be implemented at any level of the network. In the dynamic residual block, dynamic convolution can be replaced by other methods that can adjust the convolution kernel parameters, such as deformable convolution.
[0113] This invention can adaptively adjust the parameters of the convolution kernel based on the content of the input image, enhancing the model's expressiveness and better handling complex compression artifacts. By addressing the quantization parameter adaptation issue through dynamic convolution, the invention can better handle varying degrees of artifacts caused by encoding at different quantization parameters (QP).
Claims
1. A video coding loop filtering method based on dynamic convolutional neural network, characterized in that: The following steps are included: Step 1: Extract multi-dimensional shallow features from the input reconstructed image and predicted image; Step 2: Combine the shallow features with the quantization parameters and reduce the size of the combined feature map by downsampling; Step 3: The backbone of the network extracts deep features from the fused and reduced feature map, dynamically adjusts parameters based on the characteristics of the input data, and obtains the processed feature map; Step 4: Integrate the processed feature maps and restore them to the original input resolution through upsampling to obtain the residual map. Add the residual map to the original reconstructed image to obtain the filtered image. The step 2 is specifically as follows: Use dynamic convolution with parallel convolution kernels to reduce the dimension and fuse the spliced shallow features; The dynamic convolution is expressed as: y=σ((a1·W1+…+a k ·W k )*x) (1) σ is the activation function, W1…W k is the weight of the convolution kernel, k is the number of parallel convolution kernels, a1…a k is the attention weight, which is obtained by performing global average pooling, full connection layer and Sigmoid layer operations on the input feature map x, and is expressed as: a(x)=(a1,…,a k )=Sigmoid(FC(GlobalAveragePool(x))) 2) Where a represents the operation of generating k attention weights; Dynamic convolution is used to reduce the dimension and fuse the shallow features after splicing. This operation can be expressed and expanded as follows: Y={a(X;QP)·(W1,…,W k )}*{X; QP}+b (3) =W(X;QP)*{X;QP}+b (4) =W'(X; QP)*X+ΣW”(X; QP)·QP+b (5) Where a has the same meaning as in Equation (2), and W(X; QP) is equal to the dynamic convolution kernel a(X; QP)·(W1,…,W k )}, W', W" are the representations of the simplified convolution kernel, and b represents the bias term of the convolution layer; The features extracted by the dynamic convolution layer are then passed through a nonlinear activation function PReLU and a conventional convolution to reduce the computational complexity of the backbone network, obtaining a downsampled feature map that is then fed into the backbone of the network. The backbone of the network in step 3 includes two types of dynamic residual blocks stacked in sequence, wherein the first type of dynamic residual blocks are stacked first, and then the second type of dynamic residual blocks are stacked thereon to form a series structure; The first type of dynamic residual block contains a PReLU activation function, a dynamic convolution layer before the PReLU activation function, and two regular convolution layers after the PReLU activation function. The convolution layer is used for dimensionality reduction and feature extraction of feature maps. The second type of dynamic residual block contains a PReLU activation function, a dynamic convolution layer before the PReLU activation function, and two dynamic convolution layers after the PReLU activation function. The dynamic convolution layer adaptively extracts features according to the input feature map.
2. The video coding loop filtering method based on dynamic convolutional neural network according to claim 1, characterized in that: In the step 1, The reconstructed image is obtained from the inside of the loop during the video encoding process, and is obtained by adding the quantized residual to the predicted image after dequantization and inverse transformation, without passing through the loop filtering module; The predicted image is derived from the adjacent pixels or blocks that have been encoded within the current frame. The spatial correlation of these pixels is used to generate prediction values. The predicted image is constructed by referring to the encoded adjacent samples and using multiple prediction modes.
3. The video coding loop filtering method based on dynamic convolutional neural network according to claim 2, characterized in that: The step 1 is specifically as follows: Reconstructed image I of the input rec , predicted image I pred Use regular convolution F of size 3×3 respectively 3x3 Perform shallow feature extraction to obtain feature map F rec and F pred , denoted as F rec =F 3x3 (I rec ),F pred =F 3x3 (I pred ), and then spliced with the input QPmap, the feature map F is obtained, that is, the shallow feature after multi-dimensional splicing; F = C(F rec ,F pred ,QP), where C is the splicing operation.
4. The video coding loop filtering method based on dynamic convolutional neural network according to claim 2, characterized in that: In step 1, before input, the U and V components of the reconstructed image and the predicted image are first upsampled using the nearest neighbor interpolation method so that the resolution of the U and V components matches the Y component. The Y, U, and V channels of the reconstructed image and the predicted image can all be input into the network at the same resolution. The QP map is obtained by padding the current quantization parameter value to the same spatial resolution as the reconstructed frame.
5. The video coding loop filtering method based on dynamic convolutional neural network according to claim 1, characterized in that: The specific operation of the first type of dynamic residual block is: The input downsampled feature map is Y, and the output of the dynamic convolution layer is Z. The above specific steps are expressed as: Z = PReLU(DConv4(Y)); DConv4() represents a dynamic convolution operation with 4 parallel convolution kernels, two conventional convolution layers, namely 1x1 convolution layer and 3x3 convolution layer; the 1x1 convolution layer is used to reduce the dimension of the feature map, and the 3x3 convolution layer is used to further extract features and capture more deep features. The output R1 of the residual block is expressed as follows: R1=Y+Conv 3x3 (Conv 1x1 (PReLU(DConv4(Y)))) Among them, Conv 1x1 and Conv 3x3 Represent 1x1 convolution layer and 3x3 convolution layer respectively; The feature maps after the first type of dynamic residual blocks are stacked are processed by the second type of dynamic residual blocks to deal with more complex features. The specific operations of the second type of dynamic residual blocks are: R2=R1+DConv 3x3 (DConv 1x1 (PReLU(DConv4(R1)))) Among them, DConv 3x3 and DConv 1x1 They represent dynamic convolution layers with convolution kernel sizes of 3x3 and 1x1 respectively, and each dynamic convolution layer also has 4 parallel convolution kernels.
6. The video coding loop filtering method based on dynamic convolutional neural network according to claim 5, characterized in that: The step 4 is specifically as follows: The dynamic convolution layer adaptively adjusts the parameters of the convolution kernel according to the different feature maps after input processing; after the dynamic convolution layer, the PixelShuffle layer is used for upsampling, enlarging the low-resolution feature map to the same size as the input image to obtain the residual map.
7. The video coding loop filtering method based on dynamic convolutional neural network according to claim 6, characterized in that: The process of the above dynamic convolution layer is expressed as: R=PixelShuffle(DConv8(R2)); DConv8 represents a dynamic convolution operation with 8 parallel convolution kernels, and R represents the obtained residual; finally, it is added to the reconstructed image (rec) of the input image to obtain the output after network filtering Output = rec + R.
Citation Information
Patent Citations
Method and apparatus for filtering with mode-aware deep learning
CN111194555A
Alternative quality factor learning for quality adaptive neural network-based loop filters
CN115918075A