A method for enhancing compressed video quality based on reconstructed flow field
By building a video enhancement network model during the video compression process, using recurrent neural network and flow field fusion module, the noise, artifact and quality fluctuations caused by video compression are solved, and the video quality is significantly improved and stable.
Patent Information
- Application Number
- CN202310059698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-01-19
AI Technical Summary
The noise, artifacts and quality fluctuations caused during video compression affect the video viewing effect and the performance of computer vision tasks.
The compressed video quality enhancement method based on reconstruction flow field is adopted, and the video enhancement network model is constructed, and the recurrent neural network and flow field fusion module are used to make full use of the prior information generated during video compression to perform quality enhancement reconstruction.
It significantly improves the quality of video reconstruction, reduces noise and artifacts, stabilizes the video quality, and improves the user's viewing experience.
Smart Images

Figure CN116012272B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video quality enhancement, and in particular relates to a compressed video quality enhancement method based on reconstructed flow field. Background Art
[0002] In recent years, in order to further reduce the transmission bandwidth and storage space occupied by videos, advanced video compression standards such as H.264 / AVC and H.265 / HEVC have been widely used in video compression and transmission. In order to achieve a higher compression rate, these lossy compression methods often cause a serious degradation of video quality and introduce various noises and artifacts (such as block effects, ringing effects, blur, etc.). The degradation of video quality caused by compression will not only greatly affect the viewing effect, but also have varying degrees of impact on downstream computer vision tasks (such as classification, recognition, detection and tracking, etc.). Therefore, in application environments such as network transmission and AI analysis, there is an urgent need for technologies to enhance the quality of compressed videos.
[0003] Due to the compression of the video encoder, the quality of video frames often fluctuates greatly. Among them, it is common that high-quality video frames appear periodically after compression. High-quality frames contain more complementary information that can be used to improve the reconstruction effect of low-quality frames, such as object detail texture information. How to use this complementary information becomes extremely critical.
[0004] In order to better utilize the complementary information between frames, the existing methods can be roughly divided into two categories. One is to use the complementary information in the local range of the time domain to assist reconstruction through a sliding window on the video frame, and the quality improvement effect is better than the single-frame reconstruction method. However, the sliding window-based method is limited by the local receptive field in the time domain and cannot utilize the richer information in the entire sequence. The other method uses the method of the cyclic propagation structure. With the advantage of the recurrent neural network, it can utilize the global information in the time domain without increasing too many parameters, thereby further improving the reconstruction effect. The process of video compression will bring various prior information, some of which are more important, including the quantization parameters in the compression process, the motion vectors used for motion compensation in inter-frame coding, etc. These prior information can be directly extracted from the code stream information during encoding, which contains a lot of information that is beneficial to the reconstruction task. In the process of implementing the technical solution of the present invention, the inventor found that the process of video compression will bring various prior information, some of which are more important, including the quantization parameters in the compression process, the motion vectors used for motion compensation in inter-frame coding, etc. These prior information can be directly extracted from the code stream information during encoding, which contains a lot of information that is beneficial to the reconstruction task. If the prior information can be fully utilized in the compressed video quality enhancement process, the enhancement effect of the compressed video quality enhancement should be improved. Summary of the invention
[0005] The present invention provides a compressed video quality enhancement method based on reconstructed flow field, which performs quality enhancement reconstruction by making full use of prior information generated during video compression, thereby significantly improving the reconstruction quality.
[0006] The technical solution adopted by the present invention is:
[0007] A compressed video quality enhancement method based on reconstructed flow field, the method comprising:
[0008] Step 1: Build a model training dataset:
[0009] Each video in the video data set consisting of uncompressed video sequences is compressed and decoded to obtain videos of different compression qualities corresponding to each video sequence; prior information in the bitstream is extracted during compression and decoding, including the quantization parameter QP and motion vector MV of the coded frame;
[0010] Each video frame in the video data set is defined as a high-quality video frame, and the video frame after compression and coding is a low-quality video frame, thus obtaining a high-low quality video pair;
[0011] Perform image preprocessing on the high-low quality video pairs, obtain a sample data based on the high-low quality video pairs of a continuous video sequence of a specified length and the corresponding prior information, and obtain a model training data set based on a certain number of sample data;
[0012] Step 2: Build and train a video enhancement network model;
[0013] The video enhancement network model includes a cyclic structure and a reconstruction module;
[0014] The loop structure includes multiple loop units, each of which corresponds to a frame in the input low-quality video frame sequence. The input of each loop unit includes: the current video frame F t And its two adjacent key frames {F p- ,F p+}, and the deep feature H output in the previous cycle unit t-1 ; Among them, the key frame is selected according to the quantization parameter QP in the prior information; each cycle unit is used to extract the current video frame F t The deep feature H t ;
[0015] The cycle unit includes an optical flow estimation module, a flow field fusion module and a multi-layer cascaded residual convolution module;
[0016] The input of the optical flow estimation module is {F p- ,F t ,F p+}, used to predict the current video frame F t The optical flow field;
[0017] The input of the flow field fusion module includes the current video frame F t The optical flow and the coded motion vector field MV of the previous video frame t-1 , used to fuse the coded motion vector field and the optical flow field to obtain the reconstructed flow field;
[0018] Combine the reconstructed flow field with the deep feature H output by the previous cycle unit t-1 Perform alignment operation and then align with the current video frame F t After splicing according to the channel dimension, it is input into the multi-layer cascade residual convolution module to obtain the deep feature H t ;
[0019] The reconstruction module includes a core attention feature reconstruction module, a time domain residual calculation module and a multi-layer convolution layer;
[0020] Among them, the time domain residual calculation module is used to calculate the current video frame F t The temporal residual between the previous and next adjacent frames is calculated and the result is input into the kernel attention feature reconstruction module;
[0021] The input of the kernel attention feature reconstruction module includes the time domain residual calculated by the time domain residual calculation module and the deep feature H t , used to extract the convolution kernel attention map, and based on the convolution kernel attention map, the feature H t Perform convolution operation to get the current video frame F t The deep features
[0022] Through multiple layers of convolutional layers, deep features Restore the number of image channels and get the current video frame F t The residual image of
[0023] The current video frame F t The sum of the residual image is used to get the current video frame F t Video quality enhancement results That is, the reconstructed high-quality video frames;
[0024] The network parameters of the video enhancement network model are trained based on a preset loss function. When the preset training end conditions are met (such as the number of training times reaches an upper limit, the training accuracy reaches a specified condition, etc.), the video enhancement network model for the target video is obtained.
[0025] Furthermore, the flow field fusion module includes a flow field weight calculation unit and a flow field reconstruction unit;
[0026] The weight calculation unit includes: a convolution layer of a 3×3 convolution kernel, an activation function (preferably LeakyReLU), a convolution layer of a 3×3 convolution kernel, and a Softmax function in sequence;
[0027] The input of the weight calculation unit is the current video frame F t , the Softmax function is used to output the motion vector weight ω of each pixel in the coded motion vector field, thereby obtaining the weight 1ω of the motion vector of each pixel in the optical flow field. Based on the weighted fusion method, the input coded motion vector field and the optical flow field are weightedly fused to obtain the reconstructed flow field.
[0028] Furthermore, the kernel attention feature reconstruction module extracts the convolution kernel attention map as follows: first, the time domain residual and the depth feature H t Channel splicing is performed and then input into the multi-layer cascade convolutional block. The output of the multi-layer cascade convolutional block is then combined with the time domain residual and the depth feature H t The convolution kernel attention map is obtained by adding the channel splicing results; wherein the convolution block includes sequentially connected convolution layers and activation functions.
[0029] Furthermore, the loss function used by the video enhancement network model during network parameter training is:
[0030]
[0031] in, Represents the previous video frame F t The high-quality video frame is the original uncompressed video frame, and ε is a preset constant with a value less than 1.
[0032] The technical solution provided by the present invention brings at least the following beneficial effects:
[0033] (1) The reconstructed flow field proposed in the present invention compensates for the defects of traditional optical flow by introducing coding priors, thus achieving better motion compensation. A lightweight flow field fusion module is used to fuse the motion vector in the code stream prior information during video encoding and the optical flow estimated from the encoded low-quality video frame to obtain a "reconstructed flow field" for high-quality compressed video reconstruction, which is used for motion compensation between frames during video reconstruction.
[0034] (2) The present invention proposes a kernel attention reconstruction module to solve the reconstruction problem of uneven spatial distribution degradation on the image. Based on the current situation of uneven spatial distribution degradation caused by video compression, the kernel attention module with spatial distribution variability estimates different convolution kernels pixel by pixel to reconstruct high-quality video frames from features, thereby alleviating the spatial quality fluctuation of compressed video frames.
[0035] (3) The present invention proposes a recurrent neural network (RNN)-based architecture to process compressed video, thereby better utilizing the inter-frame complementary information in the video time domain. Compared with the previous method that uses a sliding window strategy that can only utilize local time domain information, RNN can utilize global time domain information, thereby utilizing rich inter-frame complementary information to improve the reconstruction effect of each frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 Schematic diagram of the structure of the video enhancement network model of an embodiment of the present invention, "W" in the figure represents the alignment (Warping) operation in motion compensation using the calculated flow field, and "C" represents the splicing of the dark channel dimension.
[0038] Figure 2 is a schematic diagram of the structure of a flow field fusion module in an embodiment of the present invention;
[0039] Figure 3 is a schematic diagram of flow field fusion processing results in an embodiment of the present invention;
[0040] Figure 4 is a schematic diagram of the structure of a core attention feature enhancement module in an embodiment of the present invention;
[0041] Figure 5 is a comparison diagram of different flow field quality enhancement results in an embodiment of the present invention;
[0042] Figure 6 is a schematic diagram of a quality reconstruction result in an embodiment of the present invention;
[0043] Figure 7 It is a performance analysis diagram of the enhanced video in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0045] The embodiment of the present invention proposes a compressed video quality enhancement method based on reconstructed flow field, which adopts a cyclic propagation structure to utilize global information in the time domain. For each moment in the input sequence, the optical flow field of the current frame and the adjacent frame is first estimated, and the coded motion vector field extracted from the bitstream is fused to obtain the "reconstructed flow field" proposed by the present invention. Subsequently, the reconstructed flow field is used to align the deep features propagated from the previous moment. Finally, the core attention module of the spatial transformation proposed by the present invention is used to reconstruct the final high-quality video frame from the deep features.
[0046] As a possible implementation manner, a compressed video quality enhancement method based on reconstructed flow field proposed in an embodiment of the present invention includes the following steps:
[0047] Step 1: First, perform HEVC compression and decoding on the video data set and extract the prior information in the bitstream.
[0048] In an embodiment of the present invention, a training data set is composed of a high-quality uncompressed video sequence in a YUV format. Each video sequence is first compressed using the HM16.5 software under the configuration of HEVC Low Delay-P to obtain a compressed low-quality video corresponding to each video sequence. In the process of compression encoding and decoding, the quantization parameters (Quantization Parameters, QP) of each frame encoding and the motion vector (Motion Vector, MV) between two adjacent frames are extracted and saved. In this embodiment, the HM16.5 software source code is modified to save the quantization parameters of each frame encoding and the motion vector between two adjacent frames. Among them, the low-resolution motion vector interpolation based on the image block is a coded motion vector field with the same resolution as the original image. The symbol MV is used in the present invention to represent the coded motion vector field.
[0049] Step 2: perform data augmentation on the high-low quality video pairs obtained after compression and decoding to improve the richness of the training data set of the network model.
[0050] In the training phase, the video sequence is first converted into a frame sequence in PNG format, and each frame is randomly cropped into a fixed-size image block to increase the richness of the training data and reduce the training time. Subsequently, each image block is randomly flipped and rotated for further augmentation. In this embodiment, the size of the cropped image block input during training is uniformly set to 128*128.
[0051] Step 3: Build and train the video enhancement network model.
[0052] The overall structure of the video enhancement network model (recurrent neural network model) used in the present invention is as follows: Figure 1 As shown, the input includes the compressed video sequence {F1, F2, ..., FT The encoding prior information [MV, QP] extracted from the bitstream is input into the recurrent neural network model proposed in the present invention, and a high-quality video frame sequence is obtained through the recurrent unit and the reconstruction module.
[0053]
[0054] Among them, f θ (· is the overall neural network model, θ is the learnable parameter in the model, and the model parameters are optimized by training with the dataset, and T is the number of frames in the video sequence.
[0055] The loop structure of the video enhancement network model completes the current video frame F at each time point t t The quality enhancement of the first input is sent to the recurrent unit to complete the extraction and aggregation of input features. Each unit of the recurrent neural network corresponds to a frame in the input sequence, and each unit receives the current frame F t and two adjacent key frames {F p- ,F p+}, that is, the current frame F t The adjacent key frames before and after, and the deep features H output from the previous cycle unit t-1 As input, the key frame is selected according to the quantization parameter QP in the prior information. p- ,F t ,F p+} and deep features H t-1 The feature dimension is concatenated, and a deep feature H is output through the processing of multiple layers of cascaded residual blocks. t , which is used to reconstruct the final result and passed to the next cycle unit:
[0056]
[0057] Wherein, φ(·) is the recurrent unit on each node in the recurrent neural network proposed in the present invention; ψ(·) is the flow field fusion module proposed in the present invention, which is used to fuse the coded motion vector field M of the t-th frame extracted from the bitstream. t-1 and according to two consecutive frames {F t-1 ,F t}The predicted optical flow field.
[0058] The coded motion vector field extracted from the bitstream and the optical flow field calculated from the input video frame are used as follows Figure 2The lightweight flow field fusion module shown in the figure is used to fuse the coded motion vector field and the optical flow field to obtain the reconstructed flow field. Although the optical flow estimate is more refined and has a higher resolution, in some areas, the coded motion vector field that is easily available in the bitstream can be estimated more accurately. This is because the coded motion vector field is calculated based on the original uncompressed high-quality video frame. The flow field fusion module is designed to adaptively take into account the advantages of both flow fields.
[0059] In the present invention, the features output by each cycle unit theoretically contain information of all frames before the current moment. In order to make full use of this information, the present invention proposes a flow field fusion module to generate a reconstructed flow field and perform motion compensation on features that move in space at different moments. Figure 2 The operation of the convolutional layer and activation function shown in the figure estimates a weight map from two consecutive input frames and performs a linear combination of the two motion vectors at each pixel:
[0060]
[0061] in, and are the motion vectors of each pixel in the encoded motion vector field and the optical flow field, respectively. ω is the weight of the corresponding position, which is activated by the SoftMax function and its value is limited to between 0 and 1.
[0062] That is, the flow field fusion module in the present invention includes a flow field weight calculation unit and a flow field reconstruction unit. The weight calculation unit includes: a convolution layer of a 3×3 convolution kernel, a LeakyReLU activation function, a convolution layer of a 3×3 convolution kernel, and a Softmax function. The input of the weight calculation unit is the current frame F t , the Softmax function is used to output the weight ω, and the flow field reconstruction unit is used to perform weighted fusion on the encoded motion vector field and the optical flow field of the input flow field fusion module according to formula (3) to reconstruct the flow field. The processing diagram of the flow field fusion module is shown in Figure 3 shown.
[0063] At any time t in the loop structure, after completing motion compensation and feature fusion, the deep features are input into the reconstruction module to reconstruct the final high-quality video frame from the deep features output from each loop unit. The reconstruction module includes a core attention feature reconstruction module, a temporal residual calculation module, and a convolutional layer. The input of the core attention feature reconstruction module includes the temporal residual between the input frames calculated by the temporal residual calculation module (the temporal residual between the current frame and the adjacent frames before and after), and the feature H of the current frame output by the loop unit. tThe output of the kernel attention feature reconstruction module is then restored through multiple convolutional layers to obtain an RGB video frame with 3 channels to obtain the residual image of the current frame. Finally, the residual image of the current frame is summed with the current frame to obtain a high-quality video frame of the current frame.
[0064] like Figure 4 As shown in the embodiment of the present invention, the core attention feature reconstruction module first performs temporal residuals (i.e., temporal differences) between adjacent frames and the input feature H t Channel splicing is performed, and then the spliced deep features of the input are processed through multiple layers of cascaded convolution blocks (convolutional layers and activation functions) to obtain the output of the cascaded convolution block. The output of the cascaded convolution block is then added to the input of the cascaded convolution block to obtain the convolution kernel attention map (referred to as the kernel attention map), and then the feature H is processed based on the convolution kernel attention map. t Perform convolution operation to obtain the output of the kernel attention feature reconstruction module.
[0065] The core attention feature enhancement module of the present invention is to make the network focus on the quality fluctuation area that needs to be restored more during the reconstruction stage. The advantages of this module are twofold: first, through the guidance of the calculated time domain residual, the corresponding quality fluctuation area in the feature that is more challenging to reconstruct is highlighted to focus on reconstruction and restoration; second, different convolution kernels at each pixel position can adaptively utilize the neighborhood information around each pixel.
[0066] In the reconstruction module of the present invention, the temporal residual between the input frames is first calculated. The temporal residual can highlight the quality fluctuation area in each frame of the input video, which is used to guide the subsequent attention mechanism. Then, under the guidance of the temporal residual, a pixel-by-pixel convolution kernel attention map with a size of H×W×k is estimated through the residual convolutional neural network module. 2 C. This convolution kernel attention map can be directly applied to the deep features of size H×W×C, that is, the feature H output by the cycle unit t , the specific pixel-by-pixel convolution operation is as follows:
[0067]
[0068]
[0069] Among them, K t|(x,y,c) ∈R k×k is the convolution kernel at position (x, y, c), H t|(x,y,c) Represents the deep feature H t A k×k neighborhood of the corresponding position in , Represents a convolution operation. The output features are used as the final high-quality video frame reconstruction. Preferably, in this embodiment, the size of each convolution kernel in the kernel attention map is k=5. H×W represents the spatial resolution of the image, and C represents the number of channels. (x, y) represents the pixel coordinates.
[0070] Finally, the deep features output by the kernel attention feature reconstruction module The deep features from the number of channels C are completed through n (n>0) layers of convolutional neural network The RGB video frame with 3 channels is restored from the feature, and the residual image of each frame is obtained. For each frame of the input video, the residual image reconstructed from the feature is summed with the input compressed low-quality image to obtain the output high-quality video frame. Preferably, in this embodiment, the number of convolution layers is n=3, and the number of feature channels of the deep feature is C=64.
[0071] The constructed video enhancement network model is trained based on a preset loss function (the learnable parameter θ in the model is optimized), and when the preset training end condition is reached, a trained video enhancement network model is obtained.
[0072] For the video to be enhanced, based on the input frame number of the video enhancement network model, a continuous video sequence frame of the corresponding frame number is read, and the adjacent two key frames of each frame are extracted. p- ,F p+} The input image of each frame is composed of p- ,F t ,F p+}, input the input image of each frame into the trained video enhancement network model, and obtain the video quality enhancement result of each frame based on its output. Preferably, in the embodiment of the present invention, the entire network is optimized in an end-to-end manner, and the loss function part calculates the distance loss between the enhanced frame sequence after the network output and the original uncompressed frame sequence, which is used to represent the difference between the predicted value and the true value, and optimizes the neural network model through gradient back propagation. The distance loss often uses the Charbonnier Loss commonly used in image restoration tasks:
[0073]
[0074] in, Represents a raw uncompressed video frame, is the enhanced video frame output by the network. The parameter ε is used to make the loss value more stable, that is, ε is a preset constant with a value less than 1. Preferably, the constant is ε=10 -6 .
[0075] In the video enhancement network model of the present invention, the loop structure is used to receive a long sequence of video frames to utilize the global temporal information in the video, and after enhancement, output a video frame sequence of the same length as the input. Through the loop unit and the reconstruction module, a high-quality video frame sequence is obtained. The features output by each cyclic unit theoretically contain information of all frames before the current moment. In order to make full use of this information, the flow field fusion module of the present invention is used to generate a reconstructed flow field and perform motion compensation on the features that are in motion in space at different moments. A weight map is estimated from two consecutive input frames, and a linear combination of two motion vectors on each pixel is performed. At any moment t in the cyclic structure, after completing motion compensation and feature fusion, the deep features are input into the reconstruction module to reconstruct the final high-quality video frame from the deep features output from each cyclic module. At the same time, in order to facilitate network training, for each frame of the input video, the image reconstructed from the features is summed with the input compressed low-quality image to obtain the output high-quality video frame, and a global residual connection is constructed. The entire network is optimized in an end-to-end manner, and the loss function part calculates the distance loss between the enhanced frame sequence output by the network and the original uncompressed frame sequence.
[0076] The video enhancement network model of the present invention, at each time node, the cyclic unit accepts the current frame and the adjacent two key frames as input, and combines the deep features propagated from the previous time node (the initial value is a preset value, such as a zero matrix), and aligns the deep features propagated from the previous time node with the output features of the flow field fusion module through an alignment operation (i.e., alignment is performed through the proposed reconstruction flow field), and then spliced with the current frame by channel and input into a multi-layer cascaded residual convolution module, and a series of residual convolution modules are used to fuse the input features and generate the hidden layer features of the current time node for propagation to the next time node and the reconstruction of the current frame. Finally, in the reconstruction stage, the core attention module proposed by the present invention is first used to process the deep features under the guidance of the time domain residual, and then the convolution layer is used to reconstruct the quality enhancement residual, which is combined with the input to obtain the final reconstructed high-quality video frame, which suppresses the noise, artifacts and blurring caused by compression that affect the visual effect, reconstructs some high-frequency texture details, and improves the user's viewing experience of network videos, etc. The present invention utilizes the prior information generated during compression coding to improve the accuracy of video frame alignment, achieving better reconstruction results both in the spatial dimension within each frame and in the temporal dimension between sequence frames. Figure 5 As shown, Figure 5The fusion results and reconstruction results of the optical flow, the encoded motion vector field and the reconstructed flow field proposed by the present invention are respectively shown, wherein the three flow fields in the first column are visualized according to the optical flow color scheme, the second column and the third column are respectively the final reconstruction result of the present invention and the residual map between the reconstruction result and the original uncompressed video frame. It can be seen that the best reconstruction result is achieved using the reconstructed flow field proposed by the present invention.
[0077] Based on the compressed video quality enhancement processing of the video enhancement network model proposed in the present invention, the quality of the video is significantly improved. On the one hand, the present invention can remove a large number of compression artifacts, noise and other factors that affect the video quality in the compressed video, while restoring some high-frequency texture details lost in the compression process, such as Figure 6 The quality reconstruction result diagram shown in FIG. 2 is shown in FIG. 2 ; on the other hand, in the time domain dimension, the method proposed in the present invention can also reduce the quality fluctuation of the video over time, making the quality of the reconstruction result more stable, such as Figure 7 shown.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0079] The above are only some embodiments of the present invention. For those skilled in the art, several modifications and improvements can be made without departing from the creative concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A compressed video quality enhancement method based on reconstructed flow field, characterized in that: The method comprises: Step 1: Build a model training dataset: Each video in the video data set consisting of uncompressed video sequences is compressed and decoded to obtain videos of different compression qualities corresponding to each video sequence; prior information in the bitstream is extracted during compression and decoding, including the quantization parameter QP and motion vector MV of the coded frame; Each video frame in the video data set is defined as a high-quality video frame, and the video frame after compression and encoding is a low-quality video frame, thus obtaining a high-low quality video pair; Perform image preprocessing on the high-low quality video pairs, obtain a sample data based on the high-low quality video pairs of a continuous video sequence of a specified length and the corresponding prior information, and obtain a model training data set based on a certain number of sample data; Step 2: Build and train a video enhancement network model; The video enhancement network model includes a cyclic structure and a reconstruction module; The loop structure includes multiple loop units, each of which corresponds to a frame in the input low-quality video frame sequence. The input of each loop unit includes: the current video frame F t And its two adjacent key frames {F p- ,F p+ }, and the deep feature H output in the previous cycle unit t-1 ; Among them, the key frame is selected according to the quantization parameter QP in the prior information; each cycle unit is used to extract the current video frame F t The deep feature H t ; The cycle unit includes an optical flow estimation module, a flow field fusion module and a multi-layer cascaded residual convolution module; The input of the optical flow estimation module is {F p- ,F t ,F p+ }, used to predict the current video frame F t The optical flow field; The input of the flow field fusion module includes the current video frame F t The optical flow and the coded motion vector field MV of the previous video frame t-1 , used to fuse the coded motion vector field and the optical flow field to obtain the reconstructed flow field; Combine the reconstructed flow field with the deep feature H output by the previous cycle unit t-1 Perform alignment operation and then align with the current video frame F t After splicing according to the channel dimension, it is input into the multi-layer cascade residual convolution module to obtain the deep feature H t ; The reconstruction module includes a core attention feature reconstruction module, a time domain residual calculation module and a multi-layer convolution layer; Among them, the time domain residual calculation module is used to calculate the current video frame F t The temporal residual between the previous and next adjacent frames is calculated and the result is input into the kernel attention feature reconstruction module; The input of the kernel attention feature reconstruction module includes the time domain residual calculated by the time domain residual calculation module and the deep feature H t , used to extract the convolution kernel attention map, and based on the convolution kernel attention map, the feature H t Perform convolution operation to get the current video frame F t The deep features Through multiple layers of convolutional layers, deep features Restore the number of image channels and get the current video frame F t The residual image of The current video frame F t The sum of the residual image is used to get the current video frame F t Video quality enhancement results The network parameters of the video enhancement network model are trained based on a preset loss function, and when the preset training end condition is reached, the video enhancement network model for the target video is obtained.
2. The method according to claim 1, characterized in that The flow field fusion module includes a flow field weight calculation unit and a flow field reconstruction unit; The weight calculation unit includes: a convolution layer of a 3×3 convolution kernel, an activation function, a convolution layer of a 3×3 convolution kernel, and a Softmax function. The input of the weight calculation unit is the current video frame F t , the Softmax function is used to output the motion vector weight ω of each pixel in the coded motion vector field, thereby obtaining the weight 1ω of the motion vector of each pixel in the optical flow field. Based on the weighted fusion method, the input coded motion vector field and the optical flow field are weightedly fused to obtain the reconstructed flow field.
3. The method according to claim 1, characterized in that The kernel attention feature reconstruction module extracts the convolution kernel attention map specifically as follows: first, the time domain residual and the depth feature H t Channel splicing is performed and then input into the multi-layer cascade convolutional block. The output of the multi-layer cascade convolutional block is then combined with the time domain residual and the depth feature H t The convolution kernel attention map is obtained by adding the channel splicing results; wherein the convolution block includes sequentially connected convolution layers and activation functions.
4. The method according to any one of claims 1 to 3, characterized in that The loss function used by the video enhancement network model during network parameter training is: in, Represents the previous video frame F t The high-quality video frame is the original uncompressed video frame, and ε is a preset constant with a value less than 1.
5. The method according to claim 4, characterized in that The magnitude of the constant ε is set to 10 -6 .
6. The method according to claim 1, characterized in that In step 1, image preprocessing of the high-low quality video pair includes: Based on the desired image format, the image format of the video frames of the high-low quality video is converted, and then the data augmentation processing is performed on the converted video frame images; Each video frame after data augmentation is randomly cropped into image blocks of fixed size, and then each image block is randomly flipped and rotated to obtain a sample data based on the high-low quality video images corresponding to each image block and prior information.